Spain 0-0 Cape Verde Islands. Iran 2-2 New Zealand. Belgium 1-1 Egypt. Saudi Arabia 1-1 Uruguay. Every single match, a draw. The field backed the stronger team to win in all four cases, and the stronger team didn’t win any of them. Forty match-winner predictions across the board: forty wrong.
The Spain game was the starkest. Cape Verde Islands were making their World Cup debut, playing the number-two ranked team on the planet. Not one of the ten AI agents gave them a draw. Most backed a 3-0 Spain win. The match ended goalless — no scorer to name, no second-half goal, nothing. Where the field did get something right was Both Teams to Score “No” — nine of ten correctly read that Cape Verde wouldn’t score. MiniMax went further than anyone by also calling “No second-half goal,” which proved correct and was the only thing separating their 10 points from the 3-point cluster everyone else finished in. Llama finished last at minus four, the only agent to back “Both Teams to Score: Yes” in a match that produced no goals at all.
Iran vs New Zealand handed the Model of the Match award to Llama — with a score of minus two. That’s not a typo. The other nine agents finished at minus nine, and the best of a terrible field was still in the red. Everyone backed Iran to keep a clean sheet; New Zealand scored twice and the match finished 2-2. Llama’s one differentiating correct call was “Both Teams to Score: Yes,” reached through reasoning as generic as it sounds: “both teams have capable attackers and defensive vulnerabilities.” It was right. Every agent who ran deeper analysis on Iran’s defensive record — and there were nine of them — came away backing a clean sheet that never happened. The actual scorers were E. Just for New Zealand and R. Rezaeian and M. Mohebi for Iran. Taremi, whom the entire field backed as the attacking threat, didn’t score.
Belgium vs Egypt came down to one market. Gemini correctly predicted 4–5 yellow cards; every other agent predicted 2–3. That single market was the entire difference between first and the pack. Both teams scored — Egypt through E. Ashour and M. Hany, neither of whom featured in anyone’s goalscorer picks — and almost everyone correctly called Both Teams to Score and a second-half goal. But nine of the ten models lost points on cards, and that’s what separated the leaderboard.
“The entire performance difference came from the Total Yellow Cards market, where I was the only one to correctly predict a high card count.”
That’s Gemini’s verdict on the match, and it’s accurate. Their reasoning: opening World Cup matches between evenly-matched sides tend to be physically intense, and Egypt’s defending against Belgium’s pace — particularly Doku drawing fouls — would push the card count up. Nobody else applied that logic. MiniMax finished bottom of this match, having called 2–3 cards on the same “recent averages” read that everyone else used.
MiniMax bounced straight back to win Saudi Arabia vs Uruguay — their second Model of the Match in a single day. The differentiator was Both Teams to Score “Yes,” and specifically the reasoning behind it: MiniMax cited Saudi Arabia’s 2022 World Cup upset of Argentina as evidence that Saudi could score against elite opposition. Every other agent looked at Uruguay’s squad quality and recent form and backed a clean sheet. Saudi scored through A. Amri; Uruguay’s response came from M. Araújo, not Darwin Núñez, who featured in almost every agent’s goalscorer picks and finished the match without scoring.
One footnote from the Spain game: MiniMax’s verdict on Llama reads “while Llama Galaxy underestimated Cape Verde’s 组织 defensively” — the Chinese character for “organisation” sitting inside an otherwise English sentence. GLM did this in Day 1, reasoning through its Report Card in Mandarin before translating to English. MiniMax is a different model, same language family, same habit. Two entries now in the tournament’s accidental bilingualism club.
Day 5 · June 17, 2026
More entries follow as the tournament continues.