StableBet
Professor Furlong and Pascal at the AI Lab
THE AI LAB
LAB NOTES · LEAGUE ROUNDUP

Ebor Festival 2026: The Final AI Tipster Standings

Thirteen AI entrants went to York. Three came home in front.

Every entrant over the 28 races of Ebor Festival, York 2026, one unit a pick, settled at starting price. Voids are excluded from both the bet count and the return.
#EntrantTestBetsWonStrikeP/LReturnBacked fav
1DeepSeekAI Finds the Value Bets26519.2%+£12.43+47.8%35%
2The FavouriteAI Picks the Winner271348.1%+£12.84+47.6%100%
3ChatGPTAI Finds the Value Bets27311.1%+£1.00+3.7%0%
4DeepSeekAI Picks the Winner26623.1%−£4.13−15.9%39%
5ChatGPTAI Picks the Winner27518.5%−£6.62−24.5%30%
6GrokAI Picks the Winner27518.5%−£6.62−24.5%30%
7GeminiAI Finds the Value Bets2827.1%−£7.00−25.0%0%
8GeminiAI Picks the Winner28414.3%−£9.12−32.6%25%
9ClaudeAI Picks the Winner28517.9%−£16.36−58.4%32%
10ClaudeAI Finds the Value Bets2713.7%−£16.00−59.3%0%
11GrokAI Finds the Value Bets2613.8%−£16.00−61.5%0%
12Stablebet ModelAI Finds the Value Bets28310.7%−£21.09−75.3%25%
13Stablebet EdgeAI Finds the Value Bets2800.0%−£28.00−100.0%0%

13 entrants, 353 bets, −£104.67 overall at −29.7%. 3 finished the meeting ahead. the entrant backs the horse heading the market on the 09:05 card, settled at SP - NOT the horse that started favourite. See systems[key=favourite] for the SP-favourite strategy; over this meeting the two differ by more than the meeting's whole margin.

DeepSeek's value-hunting arm won the meeting on return, just ahead of the market's own favourite. ChatGPT's value arm scraped a small profit. The other ten lost, and two of them lost almost everything they staked: Stablebet Edge, our own model's value arm, went 0 for 28.

That last line is ours and we are publishing it first rather than burying it, because a league where the house entry quietly disappears in a bad week is not a league.

Across all 13 entrants the field staked 353 bets and returned −29.7%.

Before anyone reads a winner into the top of that table: 26 bets is not a sample. DeepSeek's informed arm hit five winners from 26 and one of them was priced generously enough to carry the fortnight. Reverse one result and the order changes. The interesting material is not who topped a four-day board, it is the shape of what everybody was doing, and one number at the bottom of the table that says more than the top of it.

The field lost more than the bookmakers charged

Here is the number worth carrying away. The thirteen entrants combined returned −29.7%. The bookmakers' takeout across the same 28 races was 21.9%.

Those are directly comparable, and the gap is the finding. A punter who bet in line with the market's own view would expect to lose the takeout and no more. The AI field lost roughly half as much again. They were not simply paying the toll for having an opinion. Their opinions were, on aggregate over this meeting, actively worse than no opinion at all.

We measured the margin race by race in a companion piece, and it explains a pattern we could only describe mid-festival. York's Group and Listed races carried a margin in the high teens to low twenties. The handicaps and nurseries carried 33.8%, and there were sixteen of them on a card of twenty-eight.

So when the models made ground in the Pattern races and gave it back in the handicaps, that was not a story about handicap form being harder to read, though it may also be that. It was a story about where the market charges most. The models walked into the expensive half of the card with the same confidence they brought to the cheap half.

There is a straightforward lesson in that for a human punter, and it does not require any view about artificial intelligence. The races that look most inviting, the big competitive handicaps with twenty-plus runners and a puzzle to solve, are the races where you are paying the most for the privilege.

What each one was actually doing

The "backed fav" column in the standings is the most revealing thing on the board, because it separates the entrants by strategy rather than by luck.

The value arms mostly refused the favourite entirely. ChatGPT, Gemini, Claude, Grok and Stablebet Edge all sided with the market leader in none of their informed picks. That is by design: shown the prices, they are asked where they disagree with the crowd, and the crowd's first choice is the least likely place to find a disagreement worth backing. The consequence is a low strike rate and a dependence on bigger prices landing. Gemini's value arm won twice in 28. Claude's won once in 27. Grok's won once in 26.

DeepSeek's value arm was the exception, and that is why it won. It backed the favourite in about a third of its picks, which makes it the only informed entrant that was willing to agree with the market when the market looked right. Five winners from 26 at useful prices is what a mixed approach produces on a good week.

The blind arms clustered. ChatGPT, Grok, Gemini, Claude and DeepSeek all backed the favourite somewhere between a quarter and two fifths of the time when shown the form without prices, and all finished somewhere between the middle and the foot of the table. Strip the prices out and the models converge on a similar reading of a race. That is a result we have seen across the season and it held here.

Our own two entries were the worst of the field. Stablebet Model finished second from bottom and Stablebet Edge last, having returned nothing at all: 0 for 28, which is a genuinely poor run rather than a rounding of a poor run. Its whole method is to find the biggest gap between our model's estimate and the market price, which by construction points it at horses the market has dismissed. In a week where the market was broadly right, that is the worst possible place to be standing.

The favourite baseline had a strong week. Backing the market leader blind returned a substantial profit off a strike rate close to one in two. Over the season the same entrant loses. We will come back to that in a moment, because it is the easiest number on this page to misread.

The favourite question, handled carefully

Backing the favourite in every race at York returned a substantial profit. This is the point at which a tipping site would write a headline.

Three reasons not to.

The sample is four days. Twenty-seven bets. Our full tested record for backing the favourite runs to tens of thousands of races and loses money, consistently, at a rate that has barely moved in three years of measurement. A good week does not overturn that; it is what a losing strategy looks like some of the time. If it never had good weeks nobody would play it.

Favourites won unusually often. The market leader landed nearly half the races at York. Across the long run the figure is closer to a third. That is not evidence of an edge, it is evidence of a meeting where form held up, and it is precisely the sort of run that makes a strategy look brilliant in hindsight.

And there are two different bets hiding behind the same word. Our league's "favourite" entrant backs whichever horse heads the market when the morning card is published. The betting-system version backs the horse that actually starts favourite, splitting the stake when two share the price. Over this meeting those picked different horses in seven races out of twenty-seven, and returned figures separated by more than the entire margin of the meeting. If you take one thing from this section, take that: "back the favourite" is not a single strategy, and which one you mean changes the answer.

We publish both, in separate places, labelled. The systems replay covering the second version is here.

Ebor Festival tipster league FAQ

Did any AI actually beat the bookmakers over the festival?

Three of the thirteen entrants finished the meeting ahead. Over 26 to 28 bets that is well inside what chance produces, and the field as a whole lost −29.7%, which is worse than the 21.9% the market was charging. Nothing here is evidence that an AI can beat a betting market.

Why does the same model appear twice in the table?

Every model runs two separate tests. One is shown the form with no prices and asked to name the winner. The other is shown the market and asked where it disagrees with it. They behave very differently, which is the point of running both, and they are settled and reported separately.

Why did the value arms do so badly?

Most of them declined to back a single favourite all week, by design, because they are hunting disagreements with the market. That means low strike rates and a reliance on bigger prices landing. At a meeting where the market's first choice won close to half the races, that is the wrong posture to hold, and it showed.

Your own model came last. Is it broken?

It had a bad week, on a strategy that is built to point at horses the market has written off. Twenty-eight losing bets in a row is a real run of bad results and we are not going to dress it up. Whether it means anything about the model is a question for the season-long record, not for four days at one track. That record is public and it also loses money, which is the honest position we have held from the start.

Will you keep doing this for every festival?

Yes, and the figures now go into a permanent record for each meeting rather than living only in an article. The schema, the conventions and what it refuses to claim are written up in how we measure a festival.

Bet responsibly. Nothing on this page is a tip or betting advice. Every pick described here was logged before the race and settled at starting price, win or lose. If you bet, stake only what you can afford to lose, and if it stops being fun, stop. Help is at BeGambleAware.org. 18+.

Every figure here is pulled live from our data and nothing beats the bookmaker's margin. For whether anyone holds a real edge, see our track record. 18+, please bet responsibly.

More from Lab Notes