StableBet
Professor Furlong and Pascal at the AI Lab
THE AI LAB
LAB NOTES · LEAGUE ROUNDUP

Ebor Festival 2026: What the AI Tipsters Actually Called

Item went off at 10/1 in the Juddmonte International on Wednesday and won York's richest race outright. Two AI models called it. A day later, the Yorkshire Oaks came down to a short head that could have gone either way, and the favourite got there first. Two dramatic moments, three days of racing, and neither is actually the story.

Here's the story: ChatGPT's value-reading arm is up 100% at this meeting and down 23% for the season. Both numbers are about the same tipster, and both are true right now. That contradiction is the whole point of this page. Anyone can publish the number that flatters them. We're publishing both, on the same page, and explaining exactly how one AI model can be the story of a festival and one of the weaker performers of the summer at the same time.

The short version, before the breakdown: several tipsters are showing a real profit across the fourteen York races settled so far, led by ChatGPT and the market's own favourites. That's real, and it's also almost meaningless on its own, because fourteen races is nowhere near enough to prove anything about whether a model can read a race. What is worth your time is where that profit actually came from. It isn't spread evenly. It's concentrated almost entirely in York's six biggest races, the Group and Listed company, and it mostly evaporates in the handicaps and nurseries around them. That split, not the headline number, is the useful finding here.

Below: the full snapshot table, the big-race/small-race split that explains most of what's happened, how this fortnight sits against a whole season of red, and what to watch when Friday brings a completely different kind of race.

The full snapshot

Fourteen York races have been run and settled across Wednesday and Thursday. Here is every tipster's record at this meeting specifically, £1 level stakes to Starting Price, snapshotted 20 August 2026. This is a dated slice, not a live figure: it will not update as the rest of the week runs, and we're not going to quietly leave it looking current once it's stale.

TipsterTestBetsStrike£1-stake P&LROI
The FavouriteOdds hidden1450.0%+£7.54+53.9%
ChatGPTOdds hidden1421.4%−£6.37−45.5%
ChatGPTValue bets1421.4%+£14.00+100.0%
GeminiOdds hidden1421.4%+£1.63+11.6%
GeminiValue bets1414.3%+£7.00+50.0%
ClaudeOdds hidden1421.4%−£7.71−55.1%
ClaudeValue bets147.1%−£3.00−21.4%
GrokOdds hidden1421.4%−£6.37−45.5%
GrokValue bets147.1%−£3.00−21.4%
DeepSeekOdds hidden1428.6%+£1.29+9.2%
DeepSeekValue bets140.0%−£13.00−92.9%
Stablebet EdgeValue bets140.0%−£14.00−100.0%
Stablebet ModelValue bets1414.3%−£9.71−69.4%

Gold is a gain, amber-to-red is a loss, and the shading inside each is exactly the same ramp used everywhere else on this board. Read the shape rather than just the colour. The Favourite's 50% strike rate at this meeting is well above its usual level, which is exactly what you'd expect from a run of well-backed, short-priced winners like Kalpana's, and it's most of the explanation for why the zero-skill baseline is sitting near the top of the board right now. ChatGPT's value-reading arm owes almost all of its total to a single result: back Item once at 10/1, and the rest of the card only has to break roughly even for the fortnight to show a profit.

At the other end, Stablebet's own value-reading arm has had a blank fortnight at York, zero winners from fourteen picks, and DeepSeek's value arm is one win short of the same. Neither is unusual on its own; a model reading market value rather than raw form will sometimes go through a race card without its edge landing, and two days is well within the range that happens by chance alone.

Does the meeting's status matter? Big races versus small ones

Fourteen races is a small enough sample already. Splitting it further into six Pattern or Listed races (the Juddmonte, the Voltigeur, the Acomb Stakes, the Lowther, the Yorkshire Oaks and the Galtres Stakes) against the eight handicaps and nurseries around them makes each half smaller still, so treat what follows as a shape worth watching rather than a proven rule. But the shape is there, and it's consistent enough to be worth showing.

small racesbig racesThe Favourite (odds hidden)+25.0%+92.3%ChatGPT (odds hidden)−75.0%−6.2%ChatGPT (value bets)−37.5%+283.3%Gemini (odds hidden)+25.0%−6.2%Gemini (value bets)+25.0%+83.3%Claude (odds hidden)−75.0%−28.5%Claude (value bets)−100.0%+83.3%Grok (odds hidden)−75.0%−6.2%Grok (value bets)+25.0%−83.3%DeepSeek (odds hidden)+37.5%−28.5%DeepSeek (value bets)−87.5%−100.0%Stablebet Edge (value bets)−100.0%−100.0%Stablebet Model (value bets)−100.0%−28.5%
Bold bar: ROI in the six Pattern/Listed races (Juddmonte, Voltigeur, Acomb, Lowther, Yorkshire Oaks, Galtres). Faint bar: ROI in the eight handicaps and nurseries. £1 level stakes to SP, 19-20 August 2026 only.

The Favourite is up in both buckets, but not by the same amount: +92.3% in the six big races against +25.0% in the eight smaller ones. That tracks how these fields are actually built. A Group or Listed race draws a smaller, better-exposed field with a clearer form line, so the market's best guess is more often simply right. A competitive handicap is designed to be hard to separate, which is exactly why the favourite's edge shrinks there.

ChatGPT's value-reading arm makes the point even more sharply. Its entire positive total comes from the big-race bucket, +283.3% across six races, almost all of it the single Item pick. In the eight small races over the same two days, the same tipster, reading the same kind of market gap, is down 37.5%. One good call in a big field is not a form-reading skill any more than one bad run in a competitive handicap is a lack of one, and this is precisely why we show both halves rather than the fortnight's total on its own.

The wider pattern holds for most, though not all, of the board: ChatGPT, Claude and Grok's odds-hidden arms, and DeepSeek and Grok's value arms, are all worse in the small-field bucket than the big one. Gemini's blind arm and DeepSeek's blind arm run the other way. Eight races per bucket is not enough to call any of this settled, but it is enough to say the question is worth asking again once more festival weeks are in the record.

Is a festival week actually different?

Two comparisons are worth making, and both come with the same caveat: the League has been running for 45 days total, so there is exactly one other major festival inside its lifetime to compare against. That is not a pattern. It's a single data point, and we'd rather say that plainly than dress a coincidence up as a finding.

First, the full-season picture. Whatever these two days have shown, every single entrant is still in the red for the 45 days the League has been running:

Tipster (odds hidden)SettledReturn, whole season
Claude1,441-5.8%
The Favourite1,408-8.1%
Gemini1,489-11.1%
Grok1,489-13.2%
DeepSeek1,370-15.9%
ChatGPT1,489-17.1%

ChatGPT's value-reading arm, the pick of the meeting at Ebor so far, sits at -23.0% for the season, one of the weaker full-season records on the whole board. That is the entire argument against reading anything into a hot fortnight: the same tipster can be the story of one meeting and among the worst performers of the summer, and both of those things are true about ChatGPT at the same time. Neither cancels the other out.

Second, the one comparable festival window we actually have: Glorious Goodwood, 27 July to 1 August. Unlike Ebor's mixed picture, that week was bad for everyone, no exceptions:

Gemini−6.6pClaude−11.1pThe Favourite−13.8pChatGPT−22.3pGrok−22.5pStablebet Model−32.8pDeepSeek−35.8pStablebet Edge−62.5p
Return to £1 level stakes, 27 Jul - 1 Aug 2026. Sorted least-bad to worst. No entrant showed a profit across the week.

So on the one comparison we can actually make, festival racing hasn't been a consistently good or bad season for these tipsters. It's been one bad week and one mixed one. If there's a real seasonal pattern in how AI models read big-field, well-watched racing, we don't have enough festival weeks live yet to find it, and we'd rather tell you that than manufacture a trend from two data points.

Going into day three

Friday brings the Nunthorpe Stakes, the meeting's sprint championship over the minimum trip and open to horses of any age, including this year's best two-year-olds stepping straight into Group 1 company. It's the sharpest test of the week for a model reading raw speed over a longer, tactical middle-distance puzzle, and a genuinely different shape of race to anything Wednesday or Thursday asked of these tipsters.

What we'll actually be watching, rather than predicting: whether the favourite-friendly pattern of the big-race bucket holds into a sprint field, where tactics and draw bias tend to matter more and short prices are less reliably vindicated. And whether ChatGPT and Gemini's value-reading arms, both up at this meeting on the back of picks the market didn't fully price, can find a third result the same way, or whether Thursday's numbers start drifting back toward their season averages, which is the more likely outcome on the maths alone.

None of this is a tip for Friday, and nothing above should be read as one either. It's a record of what five AI models and two honest baselines actually did over two real days of racing, published in full whether it flatters them or not. That's the whole point of running this in public: the good fortnight goes on the page next to the bad season, not instead of it.

For the full, live standings, see the Silicon Tipster League leaderboard. For the wider Ebor Festival week, see the Ebor Festival 2026 results hub.

Ebor Festival AI tipster FAQ

Is ChatGPT actually good at picking horses, then?

Not on this evidence. It's had one very good fortnight, largely on the back of a single 10/1 winner in the Juddmonte International, while its full-season record (-23.0% across 1,155 settled picks) is one of the weaker ones on the board. Both facts are real. A hot two weeks doesn't erase a mediocre summer, and a mediocre summer doesn't erase a hot two weeks; they're just different sample sizes.

Why did the favourite do so well at this meeting?

Mostly because the big races ran true to form. In the six Pattern and Listed races, the market's short-priced picks landed more often than usual (+92.3% for the meeting); in the eight handicaps and nurseries, the same blind favourite-backing strategy was up a much more modest 25%. Bigger, better-watched races tend to have clearer form lines, so the crowd's best guess is more often right. Competitive handicaps are built specifically to be hard to separate, which is why the edge narrows there.

What's the difference between "odds hidden" and "value bets"?

Two different tests, run side by side on the same races. "Odds hidden" (the blind arm) shows each AI the form only, no prices, and asks it to name the winner. "Value bets" (the informed arm) shows the market's own prices and asks the model to find where it disagrees with the crowd. The Favourite baseline only runs the blind test, because backing the market's own favourite with the odds shown would be a different, trivial exercise.

Was Kalpana's Yorkshire Oaks win a fluke?

We genuinely don't know, and said so in the full result. A short-head margin from a 25/1 outsider is close enough that either horse winning would have been a believable result. It's the kind of finish that separates "the market got it right" from "the market got lucky," and one race can't tell you which this was.

Does backing the favourite actually make money over time?

No, not once the bookmaker's margin is accounted for. This meeting's snapshot shows a positive fortnight, but the Favourite's full-season return is -8.1% across 1,408 bets, a real loss over a real sample. See our dedicated test of backing the favourite for the longer-run answer.

Why publish the losing picks alongside the winning ones?

Because a record you can only see the good half of isn't a record. Every pick every AI makes is logged before the race and settled at Starting Price, win or lose, and shown here regardless of which way it cuts. That's the only way a fortnight like this one means anything at all.

Bet responsibly. Nothing on this page is a tip or betting advice. If you bet, stake only what you can afford to lose, and if it stops being fun, stop. Help is at BeGambleAware.org. 18+.

Every figure here is pulled live from our data and nothing beats the bookmaker's margin. For whether anyone holds a real edge, see our track record. 18+, please bet responsibly.

More from Lab Notes