StableBet
Professor Furlong and Pascal at the AI Lab
THE AI LAB
LAB NOTES · LEAGUE ROUNDUP

Ebor Day 3: What the AI Tipsters Backed, Settled

Settled. This was written on day three of the Ebor Festival, before Friday's seven races had run. The meeting finished on 22 August and every pick described below has settled at starting price. Friday's result is in its own section, and the final standings for the whole meeting are in the Ebor Festival tipster league review.

It was day three of the Ebor Festival, and the useful story was not the one you would expect. The AI tipster with the best two days at York was also having one of its worst seasons, and both of those were true at the same time.

Item's 10/1 win in the Juddmonte International and Kalpana's short-head victory in the Yorkshire Oaks are the results everyone remembers from Wednesday and Thursday. What actually matters here is where five AI models, ChatGPT, Gemini, Claude, Grok and DeepSeek, put their money against the market's own favourite, and how that went across all fourteen races, not just the two big ones.

Friday brought the Nunthorpe Stakes, the meeting's sprint championship: a different shape of race to anything the tipsters had faced that week, and the last full test before Saturday's Ebor Handicap.

Below: how everyone did over the first two days, the pattern the meeting was showing, who each model backed on the Friday and how those picks settled, and where the numbers said the bookmaker's cut was smallest.

How the first two days went

Fourteen York races, Wednesday and Thursday. Toggle below to see the whole fortnight, or either day on its own; every figure is £1 level stakes to Starting Price. This table is fixed to those two days and does not include Friday's races; Friday's picks have their own section further down, settled separately.

TipsterTestBetsStrike£1-stake P&LROI
The FavouriteOdds hidden1450.0%+£7.54+53.9%
ChatGPTOdds hidden1421.4%−£6.37−45.5%
ChatGPTValue bets1421.4%+£14.00+100.0%
GeminiOdds hidden1421.4%+£1.63+11.6%
GeminiValue bets1414.3%+£7.00+50.0%
ClaudeOdds hidden1421.4%−£7.71−55.1%
ClaudeValue bets147.1%−£3.00−21.4%
GrokOdds hidden1421.4%−£6.37−45.5%
GrokValue bets147.1%−£3.00−21.4%
DeepSeekOdds hidden1428.6%+£1.29+9.2%
DeepSeekValue bets140.0%−£13.00−92.9%
Stablebet EdgeValue bets140.0%−£14.00−100.0%
Stablebet ModelValue bets1414.3%−£9.71−69.4%

Gold marks a gain, amber into red marks a loss, on the same scale used everywhere else on this board. Read the shape rather than just the colour: The Favourite's strike rate at this meeting sits well above its long-run rate of 35% across 30,082 races, mostly down to well-backed short-priced winners like Kalpana's. ChatGPT's value-reading arm owes almost all of its total to a single result: back Item once at 10/1, and the rest of the card only has to break roughly even for the fortnight to show a profit.

At the other end, Stablebet's own value-reading arm had a blank fortnight, zero winners from fourteen picks, and DeepSeek's value arm had the identical record: also zero from fourteen. Neither is unusual on its own. A model reading market value rather than raw form will sometimes go through a card without its edge landing, and two days is well inside the range that happens by chance alone.

What the pattern in these numbers actually is

small racesbig racesThe Favourite (odds hidden)+25.0%+92.3%ChatGPT (odds hidden)−75.0%−6.2%ChatGPT (value bets)−37.5%+283.3%Gemini (odds hidden)+25.0%−6.2%Gemini (value bets)+25.0%+83.3%Claude (odds hidden)−75.0%−28.5%Claude (value bets)−100.0%+83.3%Grok (odds hidden)−75.0%−6.2%Grok (value bets)+25.0%−83.3%DeepSeek (odds hidden)+37.5%−28.5%DeepSeek (value bets)−87.5%−100.0%Stablebet Edge (value bets)−100.0%−100.0%Stablebet Model (value bets)−100.0%−28.5%
Bold bar: ROI in the six Pattern/Listed races (Juddmonte, Voltigeur, Acomb, Lowther, Yorkshire Oaks, Galtres). Faint bar: ROI in the eight handicaps and nurseries. £1 level stakes to SP, 19-20 August 2026 only.

That split was the real finding across the first two days, not the fortnight's headline total. Profit clusters almost entirely in York's six Pattern and Listed races, the ones with smaller, better-exposed fields and a clearer form line. It mostly evaporates in the eight handicaps and nurseries around them, which are built specifically to be hard to separate. The favourite's edge at this meeting was, in large part, a big-field-versus-small-field story rather than anything unique to that fortnight.

None of it erases the season. Every tipster on the board, including the market's own favourite, is still showing a loss across the 91 days the League has been running:

Tipster (odds hidden)SettledSeason return
Claude3,021-8.9%
The Favourite3,007-10.9%
Gemini3,062-12.1%
Grok3,071-13.2%
DeepSeek2,949-16.4%
ChatGPT3,072-13.3%

ChatGPT's value-reading arm, the pick of the meeting at Ebor through those first two days, sits at -23.3% for the season: one hot fortnight and one mediocre summer, both true of the same tipster at once. The one other comparable festival window on record at the time, Glorious Goodwood in late July, was flat or worse for every entrant, no exceptions:

Gemini−6.6pClaude−11.1pThe Favourite−13.8pChatGPT−22.3pGrok−22.5pStablebet Model−32.8pDeepSeek−35.8pStablebet Edge−62.5p
Return to £1 level stakes, 27 Jul - 1 Aug 2026. Sorted least-bad to worst. No entrant showed a profit across the week.

Two festival weeks is one data point short of a trend either way. We read the meeting's pattern as a shape worth watching, not a proven rule. Two more festival records have been built since, and the series now runs to nine meetings.

Who backed what on the Friday, and how it settled

Seven York races, five AI models, each naming the runner it rated most undervalued once it could see the market. Bold marks a horse two or more models landed on independently.

RaceChatGPTGeminiClaudeGrokDeepSeek
13:50 Mile HandicapIronwillNoelan StarLe SamouraiKrasimirZennor Storm
14:25 Lonsdale CupQuickthornQuickthornFrench MasterIllinoisQuickthorn
15:00 GimcrackMarco PoloMarco PoloMussabMarco PoloMarco Polo
15:35 NunthorpeAsfooraAmerican AffairRumstarAmerican AffairBacio
16:10 HandicapZanndabadExplodeSportingsilvermineTycoonMaster Builder
16:45 MaidenMagwitchMagwitchDover StreetLaunch SequenceBenjamin Hornigold
17:20 Fillies' HandicapDream CampBint Al DaarJaney MackersDream CampSilver Lake

Five of the seven races produced a cluster: three models on Quickthorn, four of five on Marco Polo in the Gimcrack, the closest thing to a consensus pick all day. The chart below shows how much these five usually agree with each other, chance-corrected, across the whole season, so the day's clustering can be read against a real baseline rather than a guess at whether five-model agreement is normal.

Who agrees with whom (AI Finds the Value Bets)
ChatGPTGeminiClaudeGrokDeepSeekSB EdgeSB ModelChatGPTGeminiClaudeGrokDeepSeekSB EdgeSB Model—0.260.210.210.140.12−0.070.26—0.180.150.110.24−0.050.210.18—0.110.060.03−0.120.210.150.11—0.140.080.090.140.110.060.14—0.04−0.000.120.240.030.080.04—0.12−0.07−0.05−0.120.09−0.000.12—Agreement (κ), chance-corrected01
Each cell is the chance-corrected agreement (kappa) between two AIs on the same races: brighter = they pick the same horses more often. The two Stablebet house entrants sit alongside the chatbots for comparison. On a young sample these are rough, so read the pattern rather than the last decimal.

The Nunthorpe itself was the outlier. Four different horses across five models, no majority anywhere: Gemini and Grok on American Affair, ChatGPT on Asfoora, Claude on Rumstar, DeepSeek on the market's own favourite, Bacio. Our own prediction model read that race the same way, as an open field, for the same reason: an eighteen-runner sprint compresses the form book, and five independent readers landing on four different answers is what an honestly hard race looks like, not a sign that one of them is wrong.

How the Friday settled

Not one of the clusters landed. Quickthorn was beaten in the Lonsdale Cup, which went to Al Nayyir at 9/1. Marco Polo was beaten in the Gimcrack, which went to Arapaho Gold at 2/1. Magwitch was beaten in the Convivial Maiden and Dream Camp in the Fillies' Handicap. The Sky Bet Mile went to Blue Courvoisier at 13/2 and the Assured Data Protection Handicap to Plage De Havre at 11/2, and no model had either.

One model found the day. DeepSeek took three of the seven: Bacio at 3/1 in the Nunthorpe, Benjamin Hornigold at 16/1 in the Convivial Maiden and Silver Lake at 11/10 in the Fillies' Handicap. Every other model drew a blank across the card.

What separated it is the thing the whole meeting kept showing. DeepSeek was the one value arm willing to agree with the market when the market looked right: Bacio and Silver Lake were both well fancied, and the other four models had spent the week walking past the favourite by design. Read that as one day rather than a method. Three winners from seven is a good afternoon, and DeepSeek's value arm ended the full meeting on five winners from 26, which the final standings put in the context it needs.

Where the numbers actually point

14:25 Weatherbys Lonsdale Cup Stake...8 runners16.5%15:00 Al Basti Equiworld Dubai Gimc...10 runners19.4%17:20 Sky Bet EBF Fillies' Handicap16 runners23.3%16:45 British Stallion Studs EBF Co...14 runners29.5%16:10 Assured Data Protection Handi...16 runners32.1%15:35 Coolmore City Of Troy Nunthor...18 runners36.5%13:50 Sky Bet Mile Handicap (Herita...19 runners38.0%
Bookmaker margin per race at York, 2026-08-21, computed from starting prices (sum of 1/SP minus 1). Fairest race first. Margin tracks field size, not which horse we fancy: the biggest handicaps on the card carry the most, the smallest fields the least.

That is the honest version of "what was the safest bet": not a horse, a race. The smaller the field, the smaller the bookmaker's own cut tends to be, and on the Friday that meant the Lonsdale Cup rather than the big-field sprint everyone was watching. The Nunthorpe itself carried one of the heaviest margins on the card, eighteen runners priced up, which is exactly what you would expect from the day's most competitive race rather than anything sinister about how it was priced.

Our own prediction model, separate from the five League tipsters above, read the Nunthorpe as a wide-open field: its estimates across all eighteen runners were tightly bunched rather than clustered around one or two names. Its single widest gap with the market was on the favourite itself, Bacio, rated well below its price. Its most eye-catching outsider case was Starlust at 34/1, rated closer to the pack than the market had it, though by a smaller margin than the favourite's gap ran the other way. Both were model output, not picks, and eighteen closely-matched sprinters is precisely the kind of field where a model's edge is hardest to trust. Bacio won it at 3/1, so the model's widest disagreement of the day was with the horse that went on to win.

None of this was where to put money. It was where the bookmaker was taking least, and where the market itself was least sure, which is a different and more honest question than "who wins." The favourite wins 35% of the time across 30,082 races long-run, and still returns -9.1% once the bookmaker's own margin is paid. If backing the market's own best guess loses money on that many races, nothing on this page is a lock either.

For the full, live standings behind every number on this page, see the Silicon Tipster League leaderboard. For the wider Ebor Festival week, the Ebor Handicap included, see the Ebor Festival 2026 results hub and the settled tipster league review.

Ebor Festival AI tipster FAQ

Can AI actually predict horse racing results?

Not on the evidence so far. No tipster in the Silicon Tipster League, AI or otherwise, has a positive season: 91 days and 37,288 settled picks in, and every one of them is showing a loss. That fortnight's Ebor numbers were a real short-term swing, not proof that a model can read a race better than the market can. That's the honest answer even with a flashy exception sitting inside it.

How did the AI tipsters do over the first two days of Ebor?

Several showed a real profit, led by ChatGPT's value-reading arm off the back of Item's 10/1 Juddmonte win, and by the market's own favourite. See the full day-by-day breakdown above, and the settled review of the whole meeting for where it finished. Fourteen races is nowhere near enough to prove anything about whether a model can read a race, and every one of these tipsters is still down for the season once you look past the fortnight.

Did the AI models agree with each other on the Friday's picks?

Mostly, but not on the race that mattered most. In five of the Friday's seven York races, at least two of the five models independently backed the same horse, including four out of five converging on Marco Polo in the Gimcrack Stakes. The Nunthorpe itself was the exception: five models, four different horses, no majority pick anywhere. Not one of the clusters won, and the only model to find a winner all day was DeepSeek, which took three. That split is not a fault. An eighteen-runner sprint compresses the form book hard, and real disagreement between independent readers is what an honestly hard race looks like.

Was ChatGPT's hot start at Ebor real skill, or a small sample being lucky?

Impossible to fully separate on two days of data, but the shape leaned towards luck-adjacent. Almost the entire total comes from a single pick, backing Item at 10/1 in the Juddmonte, and the same arm's full-season record (-23.3% across 2,726 picks) is one of the weaker ones on the board. A hot fortnight and a mediocre summer were both true of the same tipster at once, and neither cancels the other out. Over the full meeting it finished with a small profit, well behind DeepSeek's value arm.

Should I trust an AI tip over the bookmaker's own odds?

No, and nothing on this page is a betting tip. Every pick here is a research signal, logged before the race and settled whichever way it lands, not advice, and the bookmaker's own price already reflects far more information than any single model does. Read this as an experiment run in public, not a recommendation.

What was the safest bet on the Friday, based on the numbers?

There was not one, and treat anyone who tells you there is with suspicion. What the numbers do show honestly is where the bookmaker takes the smallest cut: the Lonsdale Cup, an eight-runner field, carried the lowest margin on that card, well below the eighteen-runner Nunthorpe itself. That's a fairer race, not a safer bet. The favourite still only wins 35% of the time long-run and loses money doing it, so "safest" here means smallest house edge, never a horse to back.

Bet responsibly. Nothing on this page is a tip or betting advice. If you bet, stake only what you can afford to lose, and if it stops being fun, stop. Help is at BeGambleAware.org. 18+.

Every figure here is pulled live from our data and nothing beats the bookmaker's margin. For whether anyone holds a real edge, see our track record. 18+, please bet responsibly.

More from Lab Notes