StableBet
Professor Furlong and Pascal at the AI Lab
THE AI LAB
THE LAB · REFERENCE STRATEGIES

What's the best AI for horse racing tips?

We put five leading AIs on the same racecards and scored how well each reads a race: how they pick, where they agree, and how to use them well. Research, not tips.

MixedTested on five AIs, the same live racecardsScored on: Racecraft ranking
The verdict

They read a card fast and well, but none has a proven edge over the market yet.

What this experiment settles

  • How does each of the five AIs actually read a racecard, and how do their styles differ?
  • Where do the AIs agree with each other and with the betting market, and where do they split?
  • What is the honest way to use an AI for horse racing, given none has shown a proven edge yet?

The short, honest answer

Five leading AIs, ChatGPT, Gemini, Claude, Grok and DeepSeek, can all read a racecard in seconds and explain every pick in plain English, but across our live tests none has yet shown a proven edge over the betting market, so the honest use is as research, not tips.

The rest of this page shows the working. We ran the same five AIs on the same real British and Irish racecards for weeks, scored how well each one reads a race, and set out how to actually use them. We are now 19 days into the experiment, across 5,034 settled picks, and every figure below is live, pulled straight from the boards so it can never drift from what we publish.

One racecard, five minds

Hand the identical racecard to five different AIs and you do not get one answer. You get five. Each was built by a different lab, trained on different data and nudged by different instincts, and it shows the moment they start justifying their picks. That variety is the fun of the experiment, and it is the reason a blend of the five tells you more than any single voice does.

  • ChatGPT, the crowd-reader. With the odds hidden it drifts toward the shorter prices, backing the outright favourite 30% of the time, and its blind picks most resemble the "back short-priced favourites" rule.
  • Gemini, the contrarian. Even blind it wanders furthest from the chalk, its picks mapping closest to "favourite over jumps", with a typical price around 5.74.
  • Claude, the careful reader. Form-led and choosy, it leads with the yard and recent form, its top three stated reasons being trainer, form and jockey.
  • Grok, the value-hunter. The moment it sees the market its picks swing toward "our own value strategy", chasing overpriced runners rather than the well-fancied.
  • DeepSeek, the challenger. Efficient and blunt, it backs the favourite 29% of the time blind, then flips toward "back the outsider" once the prices are shown.

None of that makes one of them the answer. It makes them five different reads on the same race, and the pages below take each one apart in turn.

How we tested, and what reading a race well means

The method is simple, and we run it the same way every day. We take live British and Irish racecards and put the same races to all 8 AIs in two conditions. In the "blind" condition the odds are hidden and the AI has to read the form itself. In the "informed" condition it sees the market and hunts for value. Every pick is settled honestly to Starting Price, and we publish all of it, wins and losses alike, across 5,034 settled picks so far.

Here is the key move for this page. We do not rank the five on a few weeks of profit. A short-run return bounces around with luck, and any table built on it would contradict the live boards the next morning. Instead we score racecraft, how well each AI reads a race, on three durable pillars.

  1. Calibration. When an AI says it is confident, is it right that often? A model whose "sure things" actually win more is reading the race, not guessing.
  2. Reasons that match the pick. Do the written justifications line up with what it actually backed? An AI that talks about the yard and the form, then backs the horse that fits that story, is reasoning, not rationalising.
  3. Agreeing with the market when it is right. On the races it gets right, does it side with what the crowd already knew? The betting market is very good, so siding with it on winners is a sign of a sound read, not a lack of nerve.

Racecraft is about reading skill, not a promise of profit. The composite weighting is fixed and shown in the methodology. It is the honest, ages-well spine of the ranking, because a model can read a race well and still hand the bookmaker its margin at the window.

As a worked example of the first pillar, here is one AI's calibration: when Claude sorts its blind picks into confidence tiers, do the high-confidence ones actually win more often?

Does Claude know when it is right? (AI Picks the Winner)
0%50%100%17%Low(n=54)25%Medium(n=418)100%High(n=1)strike rate
The AI grading its own confidence. On the blind test Claude files each pick as low, medium or high confidence, and each bar is the strike rate of the picks it filed at that level. Bars that rise left to right mean the label is informative: it wins more often when it says it is surer. Flat or falling bars mean the label carries little signal. The per-tier samples are small, so treat the ordering as a hint, not proof. Research, not tips.

The table below ranks the five on racecraft. The order here follows the cleanest single reading signal we can show live, how often each AI's blind pick actually won its race, and the fixed composite that blends all three pillars is set out in the methodology. The five sit close together, and the honest yardstick is in the "market strike" column: the crowd, backing the same races, wins more often than any of them.

RankAIBlind strikeMarket strike, same racesFavourite affinity (blind)Reasons-match note
Top of the fiveChatGPT25.3%45.1%30%Leads with form, and backs the runner that fits it.
SecondDeepSeek24.8%44.4%29%Leads with form, blunt and efficient.
ThirdClaude24.4%46.3%31%Leads with the trainer, then recent form.
FourthGrok23.9%45.1%29%Leads with the trainer, then jockey and form.
FifthGemini23.8%44.9%25%Leads with form, furthest from the chalk of the five.

Read the ranking as where each AI sits on one durable reading signal, not as a verdict on which is worth backing. That question belongs to each AI's own live scorecard, linked in the next section.

The blind herd, and the odds-flip

Here is the most interesting finding, told straight. With the odds hidden, the five AIs herd. They all drift toward the shorter-priced, better-fancied runners, arriving at roughly where the market already is without ever being shown it. That is a quiet compliment to the crowd: the AIs independently rediscover the favourites. Blind, ChatGPT, Claude and DeepSeek all map closest to short-price, odds-on behaviour, and even Gemini, the one that wanders furthest, still lands well inside the front of the book.

Then show them the prices, and several of them flip. Their picks swing out toward the longer-priced, overlooked runners as they start hunting value instead of winners. It is a real change of character, and you can watch it happen in the two distribution charts below: the blind herd bunched at the short end, the informed picks spread out across the book.

Where each AI fishes in the market (AI Picks the Winner)
backs the favouriteeven splitbacks the longest pricebacked favThe Favourite100%ChatGPT30%Claude31%Grok29%DeepSeek29%Gemini25%
Each dot is an AI's average pick position across the field, from 0 (the market favourite) to 1 (the longest price in the race). The whisker is the 95% interval, and “backed fav” is the share of its picks that were the market favourite. A whisker that straddles the centre line, or spans a wide stretch, means the sample is still small and the position is not yet pinned down. Research, not tips.
Where each AI fishes in the market (AI Finds the Value Bets)
backs the favouriteeven splitbacks the longest pricebacked favStablebet Model36%Grok1%DeepSeek6%ChatGPT0%Claude0%Gemini0%Stablebet Edge0%
Each dot is an AI's average pick position across the field, from 0 (the market favourite) to 1 (the longest price in the race). The whisker is the 95% interval, and “backed fav” is the share of its picks that were the market favourite. A whisker that straddles the centre line, or spans a wide stretch, means the sample is still small and the position is not yet pinned down. Research, not tips.

The same story shows up in how tightly the five agree with each other. Blind, they cluster on the same well-fancied runners. Informed, the agreement loosens as each one chases its own idea of value.

Who agrees with whom (AI Picks the Winner)
ChatGPTGeminiClaudeGrokDeepSeekFavouriteChatGPTGeminiClaudeGrokDeepSeekFavourite0.510.470.570.520.200.510.410.540.480.130.470.410.460.440.200.570.540.460.530.180.520.480.440.530.180.200.130.200.180.18Agreement (κ), chance-corrected01
Each cell is the chance-corrected agreement (kappa) between two AIs on the same races: brighter = they pick the same horses more often. The market-favourite baseline sits apart from the pack. On a young sample these are rough, so read the pattern rather than the last decimal.
Who agrees with whom (AI Finds the Value Bets)
ChatGPTGeminiClaudeGrokDeepSeekSB EdgeSB ModelChatGPTGeminiClaudeGrokDeepSeekSB EdgeSB Model0.280.220.280.210.18−0.090.280.210.150.270.30−0.060.220.210.200.220.07−0.150.280.150.200.200.090.050.210.270.220.200.20−0.140.180.300.070.090.200.08−0.09−0.06−0.150.05−0.140.08Agreement (κ), chance-corrected01
Each cell is the chance-corrected agreement (kappa) between two AIs on the same races: brighter = they pick the same horses more often. The two Stablebet house entrants sit alongside the chatbots for comparison. On a young sample these are rough, so read the pattern rather than the last decimal.

Two named examples make the odds-flip concrete:

  • ChatGPT. Blind, it backs the favourite 30% of the time and resembles "back short-priced favourites". Informed, it shifts to "our own value strategy".
  • DeepSeek. Blind, its favourite affinity is 29% and its closest rule is "back short-priced favourites". Informed, it moves to "back the outsider".

Keep the honesty marker close. The value arm can post a positive number over a short window. We report it honestly, and it is not yet a proven edge. On the Value Test, ChatGPT currently sits at -3.4% over just 256 settled picks, far too small a sample to call an edge, and exactly the sort of early number that hooks a punter into chasing.

The five, one by one

Here are the five, one at a time. Each profile is about character and method, how the AI reads a race, not a verdict on whether to back it. For that, follow the links to its "how it picks" write-up and its live scorecard.

ChatGPT, the crowd-reader

Blind, ChatGPT leans on form first and lands on the outright favourite 30% of the time, at a typical price around 5.26. Show it the market and it changes character: its picks map closest to "our own value strategy", which backs our model's biggest mid-priced market disagreement. On the blind card it wins 25.3% of its races against the market's 45.1% on the same cards.

How ChatGPT explains itself (AI Picks the Winner)Form72%Jockey57%Trainer45%Going42%Distance38%Class5%Course3%Weight2%Draw1%Value0%0%50%100%
Share of this AI's short written reasons that mention each theme, longest first (a reason can touch several, so the shares do not add to 100). The three it leans on most are picked out in gold. This describes how it TALKS about its picks; it is not why it wins or loses. Based on 534 reasons.

Read how ChatGPT picks · See its live record

Gemini, the contrarian

Gemini reads the form and leads with form, but even blind it sits furthest from the chalk of the five, backing the outright favourite only 25% of the time at a typical price around 5.74. Given the market it maps closest to "back the outsider", which backs the longest price in the race. Blind, it wins 23.8% of its races against the market's 44.9% on the same cards.

How Gemini explains itself (AI Picks the Winner)Form83%Jockey43%Trainer34%Going24%Distance10%Class8%Draw6%Course2%Value0%Weight0%0%50%100%
Share of this AI's short written reasons that mention each theme, longest first (a reason can touch several, so the shares do not add to 100). The three it leans on most are picked out in gold. This describes how it TALKS about its picks; it is not why it wins or loses. Based on 537 reasons.

Read how Gemini picks · See its live record

Claude, the careful reader

Claude is form-led and choosy. Blind, it leads with the trainer and recent form, landing on the favourite 31% of the time at a typical price around 5.41. Shown the odds it maps closest to "back the outsider", which backs the longest price in the race. Blind, it wins 24.4% of its races against the market's 46.3% on the same cards.

How Claude explains itself (AI Picks the Winner)Trainer74%Form57%Jockey47%Distance26%Value17%Class16%Going5%Course2%Draw1%Weight0%0%50%100%
Share of this AI's short written reasons that mention each theme, longest first (a reason can touch several, so the shares do not add to 100). The three it leans on most are picked out in gold. This describes how it TALKS about its picks; it is not why it wins or loses. Based on 492 reasons.

Read how Claude picks · See its live record

Grok, the value-hunter

Blind, Grok leads with the trainer and backs the favourite 29% of the time at a typical price around 5.46. The moment it sees the market it turns hunter, its picks mapping closest to "our own value strategy", which backs our model's biggest mid-priced market disagreement. Blind, it wins 23.9% of its races against the market's 45.1% on the same cards.

How Grok explains itself (AI Picks the Winner)Trainer60%Jockey60%Form15%Going7%Distance7%Course3%Class1%Draw1%Value0%Weight0%0%50%100%
Share of this AI's short written reasons that mention each theme, longest first (a reason can touch several, so the shares do not add to 100). The three it leans on most are picked out in gold. This describes how it TALKS about its picks; it is not why it wins or loses. Based on 535 reasons.

Read how Grok picks · See its live record

DeepSeek, the challenger

DeepSeek is efficient and blunt. Blind, it leads with form and backs the favourite 29% of the time at a typical price around 5.43. Once the prices are shown it flips, mapping closest to "back the outsider", which backs the longest price in the race. Blind, it wins 24.8% of its races against the market's 44.4% on the same cards.

How DeepSeek explains itself (AI Picks the Winner)Form85%Trainer46%Draw19%Distance18%Class16%Jockey16%Going8%Course7%Value4%Weight0%0%50%100%
Share of this AI's short written reasons that mention each theme, longest first (a reason can touch several, so the shares do not add to 100). The three it leans on most are picked out in gold. This describes how it TALKS about its picks; it is not why it wins or loses. Based on 537 reasons.

Read how DeepSeek picks · See its live record

The five side by side

AILive Form Test (blind)Live Value Test (informed)Favourite affinity (blind)Blind strategy it maps toValue strategy it maps to
ChatGPT-9.6%, strike 26%-3.4%30%back short-priced favouritesour own value strategy
Gemini-8.6%, strike 24%+39.1%25%favourite over jumpsback the outsider
Claude+1.5%, strike 25%-8.5%31%back short-priced favouritesback the outsider
Grok-13.6%, strike 24%-9.2%29%back short-priced favouritesour own value strategy
DeepSeek-9.6%, strike 25%-4.2%29%back short-priced favouritesback the outsider

These ROI figures are live and move daily over small samples. They are the record of an open experiment, not a tip sheet, and a negative or positive line is equally expected this early.

The honest verdict, and how to use AI for racing well

Pull it together and the five are genuinely good at three things: reading a full card in seconds, explaining every pick in plain words, and, blind, landing near the market's own view without being shown it. That last one is a real signal of comprehension, not a fluke. What none of them has shown is a proven, repeatable edge over the market, and the reason is structural. By the time an AI and the crowd agree on a horse, the bookmaker has already shortened its price to match, so the value is gone before you can take it.

So use them for what they are good at.

  • Use AI to speed up your own reading, not to outsource the bet. Ask it to summarise the form and surface the story, then price the race yourself.
  • Prefer the blind read for a clean opinion, and treat the informed value picks as hypotheses to check against the real price, never as tips.
  • Blend the five rather than trusting one. Where they agree blind, the market usually agrees too, which tells you something even when it does not pay.
  • Track everything and stake responsibly. An AI that reads well is a research aid, not an edge.

You can watch the experiment run on the two live boards: AI Picks the Winner for the blind read, and AI Finds the Value Bets for the informed one, with 5,034 picks settled and counting. A fuller how-to guide on using AI for racing is on the way.

Research, not tips. 18+, please gamble responsibly.

Frequently asked questions

Can AI predict horse racing?
It can read a racecard and rate the runners quickly and clearly, and blind, the five AIs we test independently drift toward the same well-fancied horses the market backs, which shows real comprehension. Predicting the result reliably is another matter: across our live tests none has shown a proven edge over the market, so treat the output as research, not tips.
Which AI is best at horse racing?
It depends what you mean by best. On our durable Racecraft score, which rewards calibration and reasons that match the picks rather than a few weeks of profit, the five separate clearly, and you can see the current live records side by side in the comparison table above. For how any single one picks, follow the how-it-picks link in its profile; for its running record, see its live scorecard.
Do AI horse racing tips work?
As tips to bet blindly, no, that is not what the data supports, and none of the five has shown a repeatable edge over a real sample. As a research aid they work well: an AI can surface the form story and the key factors in seconds, which speeds up your own reading. The honest use is to price the race yourself afterwards.
Does AI back favourites?
When the odds are hidden, mostly yes, the five herd toward the shorter prices, for example ChatGPT lands on the outright favourite 30% of the time and Claude 31%. Show them the market and several flip the other way, hunting value among longer-priced runners instead, which is the most interesting behaviour we have found.
How does AI pick horses?
Each reads the card and weighs form, the yard and the conditions, then writes out its reasons. Claude, for instance, leads with trainer and form. The striking part is how much they still agree with the crowd: Claude wins 24.4% of its blind races against the market's 46.3% on the same cards, a reminder of how good the market already is.

What this experiment doesn't cover, and what we're testing next

Related race coverage

Other Lab experiments

Browse every system tested →