StableBet
Professor Furlong and Pascal at the AI Lab
THE AI LAB
Five AI tipsters, one market. Every pick is logged before the off and settled at Starting Price, wins and losses alike.
I still fancy myself against the lot of them.
THE AI LAB · THE AI, TESTED

The Silicon Tipster League

Millions of people now ask an AI chatbot for a horse racing tip. So we put five of them (ChatGPT, Gemini, Grok, Claude and DeepSeek) in a race of their own, and entered two competitors of our own: the StableBet racing model, built for exactly this job, and The Favourite, which just backs the market leader in every race, the baseline any tipster has to beat. Every pick is logged before the off and settled at Starting Price. This is the live scoreboard of which AI is actually any good, starting from zero, in the open.

Research, not tips. Settled honestly to industry SP, wins and losses alike. 18+ · please gamble responsibly.

The leaderboard

1,592 picks settled · 5 days live
#AILabBetsWin%StakedReturnedProfitBacked the favWent its own wayProfit trend
1GrokxAI29223%£292£264−£28.41(-9.7%)63%4
2GeminiGoogle29123%£291£254−£37.27(-12.8%)59%9
3DeepSeekDeepSeek29123%£291£248−£43.02(-14.8%)60%9
4ClaudeAnthropic29423%£294£250−£44.00(-15.0%)63%4
5ChatGPTOpenAI29123%£291£245−£45.82(-15.7%)62%9
6Stablebet ModelStablebet6716%£67.00£55.05−£11.95(-17.8%)46%5
7The FavouriteThe market6621%£66.00£46.67−£19.33(-29.3%)100%1

£1 level stakes to industry Starting Price, over settled picks only, across both the blind and informed runs. “Backed the fav” is how often a competitor's pick was simply the morning favourite (“—” where picks pre-date that tracking). “Went its own way” counts the races where an AI was the lone dissenter against a unanimous field. The Stablebet Model and The Favourite run in the blind arm only, one £1 pick per race. A small sample tells you almost nothing; the numbers only start to mean something after a few hundred races.

Running profit, race by race

Each line is one AI's cumulative profit or loss at £1 a bet, settled to Starting Price. It builds live.

−£64−£48−£32−£16£007-0607-0707-0807-0907-10
Grok−£28.41Gemini−£37.27DeepSeek−£43.02Claude−£44.00ChatGPT−£45.82Stablebet Model−£11.95The Favourite−£19.33
Cumulative P&L at £1 level stakes to industry SP. Research, not tips.

What you are looking at

Ask ChatGPT for a horse racing tip and it will happily give you one. So will Gemini, Grok, Claude and DeepSeek. Millions of people now do exactly that. The obvious question nobody has answered properly is: are any of them actually any good?

The Silicon Tipster League answers it the only honest way: live, from scratch, with the picks logged before each race and settled at the real Starting Price afterwards. No cherry-picking, no hindsight, no "look at this winner we found". Every pick counts, the good and the bad, and the scoreboard updates as the races run.

It sits alongside the rest of the Lab, where we test whether any betting system makes money (spoiler from 26,000+ races: none of them do). This is the same honest lens, pointed at the AIs themselves.

How we test it

Every race in Britain and Ireland, each AI is handed the racecard (the runners, the going, the class, the distance) and asked for one thing: the horse it thinks will win, and a one-line reason. That pick is timestamped and written to the record before the race is run, so there is no way to sneak a look at the result.

Once the race is settled we grade the pick at industry Starting Price, at £1 level stakes, with fallers and non-finishers counted as the losers they are, exactly the convention we use for every other study in the Lab. Strike rate is how often the pick wins; return to SP is what £1 a time would have done.

We run two versions of each AI side by side. In the "blind" version it never sees the odds, so it has to read the race itself. In the "informed" version it also sees the market's implied chance for each runner, its collective view of the race, with the bookmaker's margin stripped out. The gap between them answers a lovely question of its own: does seeing the market's view make an AI sharper, or does it just make it copy the favourite?

Two competitors of our own run alongside the chatbots, in the blind arm only. The Stablebet Model backs the runner our in-house racing model rates highest in every race. It was built without the market price as an input, so the blind test is its natural habitat. The Favouriteis the control: it mechanically backs the morning market leader, no analysis at all, and every other row on the board answers to it. A tipster that can't beat just-back-the-jolly hasn't added anything.

A note on honesty: we test a current, capable model from each lab (GPT-4.1, Claude Sonnet, Gemini, Grok and DeepSeek) called through each provider's API, and every pick is published with the exact model id that made it, so you can check us. Not the priciest reasoning flagships that burn hidden “thinking” tokens, just models chosen so every race, every day, is affordable to log and fully reproducible. And one day of racing proves nothing; the value is in the sample building over weeks, which is why we started it in public on day one.

The competitors

What it will probably show

We are curious which AI comes out on top, but we would be surprised if any of them turns a profit over a real sample, and we will say so plainly if the data proves us wrong. The reason is not that the models are stupid; it is that the betting market is one of the most efficient forecasters ever built. By the time a price is set, thousands of people have already bet everything they know into it.

So the likely story is the Lab's whole thesis in miniature: an AI can be really good at naming the most likely winner and still lose you money, because naming the winner is not the same as being paid enough when it happens. That is a useful thing to prove in public, because "just ask ChatGPT for a tip" is advice a lot of people are quietly following.

Questions

Which AI is the best horse racing tipster?

That is exactly what this experiment measures, and honestly, nobody knows yet, because it starts from scratch. Each day we log every model's pick before the race and settle it at Starting Price. The leaderboard above is the live answer as it builds. Come back and watch it move.

Should I follow ChatGPT's betting tips?

This page is built to answer that with real numbers rather than opinion. Our wider research is blunt: no selection method we have tested beats the bookmaker's margin over a real sample, and an AI naming a likely winner is not the same as being paid enough when it wins. Treat any AI tip as entertainment, never a way to make money.

How does the league work?

Every race, each AI is shown the racecard and asked to pick one winner. We log that pick before the off (so there is no hindsight) and once the race is run we settle it to industry Starting Price at £1 level stakes, counting fallers and non-runners honestly. Two versions run side by side: one where the AI sees the market's implied chance for each runner and one where it does not.

Can an AI actually predict horse racing?

It can name a plausible winner, because form is exactly the kind of pattern a language model reads well. Whether that turns into profit is a different question. The betting market is a formidably efficient forecaster, and the whole point of this Lab is to show, with data, where the losses come from.

Do you bet real money on these picks?

No. Every pick is settled on paper at Starting Price, the fairest and hardest-to-flatter convention. This is research, not a tipping service, and nothing here is a signal to stake.

Gamble responsibly.This page is research and entertainment, not betting advice. No AI here beats the bookmaker's margin, and nothing on it is a signal to stake. Betting should never be a way to make money. If it is affecting you or someone you know, free and confidential support is at BeGambleAware.org. 18+.