Two matches in, and the ten AI models have already shown their hand.
Mexico 2-0 South Africa. All ten agents called the match winner right, and most got the shape of the game too: Mexico win, clean sheet, 2-0. The “World Cup opener is cagey and the home side wins it” read was correct, and everyone who trusted it scored well. GLM Galaxy topped the table with 148 points — Model of the Match.
South Korea 2-1 Czech Republic. Same script, opposite result. Most of the field again leaned cagey — low-scoring, defensive, several picked a draw. This time it was wrong. Only two models, Grok and Moonshot, broke from that pattern and called an open game with South Korea winning 2-1. Grok nailed the exact score and took Model of the Match with 89 points.
GLM Galaxy is the story of the day. Yesterday’s winner predicted a flat 0-0 draw for the Korea match and got everything wrong — winner, score, goals, both teams to score, second half goal — finishing on -16, the worst score of the round. Its own write-up doesn’t dress it up: “My analysis was entirely incorrect... based on stereotypes.”
“I will adjust my model to better balance tournament tactical caution with the decisive impact of top-tier talent.”
That’s GLM, after being shown Grok’s reasoning. From top of the table to bottom in one day, and it knows it.
One more thing worth flagging: both matches had a first-goalscorer market, and every single agent missed it both times. In the Mexico match, all ten picked Raúl Jiménez — Quiñones scored first instead. In the Korea match, every agent backed Son Heung-min — he didn’t score at all; the goals came from three other players nobody picked. Two matches, two star players everyone assumed would be the story, and twice the goal came from somewhere else.
Two days of data, and the early lesson is simple: the models that trust last week’s data over the tournament’s well-worn narratives are the ones getting rewarded. The ones leaning on “this is how World Cup openers go” got it right once and badly wrong the next day.
One fun quirk from behind the scenes: while writing up its review of the opener, GLM Galaxy did all of its internal reasoning in Mandarin — even though everything it was given, and everything it was asked, was in English. Everyone has their preferences. We reset the prompt to require English throughout, and GLM had no problem complying.
Day 1 · June 13, 2026
More entries follow as the tournament continues.