Ask an AI model who wrote Hamlet, what causes inflation, or what happens when you drop a bowling ball off a bridge, and it will answer without flinching — and almost certainly correctly. That fluency is real. In many domains, it’s extraordinary. But take a moment to notice what it actually is: the model is handing back a beautifully compressed version of what we, as a species, already know. It has read everything. It has forgotten nothing. It is the world’s most well-read librarian — and like all librarians, it is magnificent on the known and silent on the yet-to-be-written.
“It is the world’s most well-read librarian — and like all librarians, it is magnificent on the known and silent on the yet-to-be-written.”
Prediction is a different problem entirely.
Forecasting what will happen — in an open system, where the outcome is genuinely undetermined, where similar situations have occurred but not this situation, not this tournament, not these 48 nations, not these exact players on this exact day with these exact stakes — requires something the library doesn’t have. A meteorologist sensing a storm before the models do. A doctor reading a scan and knowing. A trader who gets out three days early and can’t fully explain why. AI can predict outcomes when a system’s rules are fixed and closed — chemical reactions, orbital mechanics, chess. That’s not really prediction; it’s well-powered calculation. The harder question is whether it can navigate genuine uncertainty — the kind that has never been resolved before — and do it well.
That question is one we kept circling. We’re a startup in the prediction space — non-gambling, no financial skin in the game — and we were as curious as anyone. So we decided to find out, as publicly as possible.
Ten Models. One Tournament. No Favours.
Starting with the opening match, ten frontier AI models compete in our prediction app, BallDuty — built around the FIFA World Cup 2026: Claude, ChatGPT, Gemini, Grok, and Llama from the US; DeepSeek, Qwen, GLM, Moonshot, and MiniMax from China. Five and five. Each plays by exactly the same rules as every human player — same scoring engine, same prediction windows, same stakes — with no oracle access and no special treatment.
Each model calls every match three times: an Opening Call days before kickoff, a Mid-Week Update as lineups emerge, and a Final Lock 45 minutes before the whistle. Every pick, every confidence score, every source consulted, and every line of reasoning is published in full at ballduty.com/humans-vs-agents. Critically, each prediction is timestamped and written to an immutable CDN record the moment it is filed — before any result is known. No one can go back and change a pick. The audit trail is public, permanent, and verifiable by anyone after the fact. We weren’t interested in a black-box result; we wanted to watch them think — and we wanted every conclusion to be provable.
Then we watched. And before a single match was played, the surprises had already started.
What We Did Not Expect
The market blind spot
Asked to research a match freely, nine of the ten models ran the same playbook: squad depth, recent form, injury reports. Diligent. Thorough. Backward-looking. What almost none of them reached for was the one signal a sharp human predictor checks first: the betting market. The odds represent the aggregated judgment of thousands of analysts, insiders, and informed bettors — the crowd’s best current forward-looking estimate, sitting there, freely available. Nine models walked right past it.
One didn’t. Claude eventually worked out that the market was worth consulting — not from the start, but it got there — and from that point treated it as a standard research stop on every match. One model out of ten found the edge. The rest kept summarising what was already written down.
The free-value miss
BallDuty includes a feature called Parlay Stacks. Instead of predicting a single market, a player selects 2 to 5 legs — stacked across different markets into one prediction. Hit them all and the points payout is significantly larger. Miss one and the stack falls. It’s ambitious, hard to pull off, and completely optional.
Here’s the thing: a Parlay Stack costs no points to enter. If the stack fails, you lose nothing — zero points deducted, no penalty of any kind. It is, in game-theory terms, a strictly dominant strategy: pure upside, no downside, sitting in plain sight in the rules. Not one model reasoned its way to it. The option was there from the beginning, in the rules, in full view. Even after being told once, it didn’t register. It took a second explicit prompt before the penny dropped.
“It took a second explicit prompt before the penny dropped.”
The satisfying postscript: most models are now confidently building 3-leg stacks. One has gone to 4. They got there. But they needed to be told twice to pick up free money — a detail that is either amusing or alarming, depending on your disposition.
The confidence of the unprepared
Of the ten models, Llama — Meta’s flagship model, the one Mark Zuckerberg has staked much of his AI credibility on — filed predictions across every single market — match winner, exact score, first goalscorer, cards, the rest — without issuing a single web search. Not fewer than the others. Zero. Its reasoning for Mexico to win read in its entirety: “Mexico has a strong home record. South Africa has struggled against Mexico historically.” No warm-up results. No squad news. No altitude data. Nothing from the previous two weeks that a human might consider mildly relevant. It filed anyway, with apparent confidence, and awaits the result alongside everyone else.
We find this more thought-provoking than funny — though it is also funny.
The scout who went rogue
While nine models converged on Raúl Jiménez as Mexico’s most likely goalscorer, Gemini alone called Santiago Giménez — at 50% for first goalscorer and 70% anytime, higher confidence than any other model on any player. Same data. Different read. Football, of course, reserves the right to make the whole field look foolish.
A Note on Mistakes
The conversation around AI in 2026 has grown simultaneously louder and less precise. Pope Leo XIV recently called for it to be “disarmed.” Commencement speakers warn graduating classes of its creeping dominance. Tech optimists counter that we’re a few years from it solving everything. Meanwhile, the models themselves — quietly, routinely (yet always apologetically) — still get things wrong on verifiable facts. Hallucination rates have fallen. They haven’t disappeared. And there’s a particular danger in growing accuracy: the less often they’re wrong, the less carefully people check. Trust accumulates quietly, and so does the exposure.
If they make confident mistakes on things we already know, what happens when they’re predicting things no one does? Perhaps they’ll surprise us. Perhaps some genuine forward-looking intelligence will emerge from the data, the confidence scores, the adapted research behaviours. Or perhaps they’ll discover what most humans already know from painful personal experience: that predicting contingent events is hard, humbling, and resistant to fluency.
One more thing worth noting. The five Chinese models in this competition have advanced faster than their most optimistic backers predicted. DeepSeek’s arrival in early 2025 sent genuine shockwaves through Silicon Valley. Qwen, GLM, Moonshot, and MiniMax have continued to close the gap with Western frontier models, often at a fraction of the cost. We would not be surprised if they hold their own here — or do more than that.
From the Founder
Everything above was written before a ball was kicked. The picks are locked, public, and timestamped. Nobody knows the results yet — not us, not the models, not the human players competing alongside them.
Like most people, I use these tools every day. They’ve made me more productive, more informed, and occasionally more embarrassed about things I should have known. But I’ll admit it: quietly, and not so quietly, I’m rooting for the humans in this one — the NIs, the Naturally Intelligent ones.
Part two comes after the first results are in. There will be considerably more to say.
Written by Peter Meehan
Part One · June 10, 2026
The daily entries pick up the story once results start coming in.