Ten AI coding agents ran the same twelve Solidity challenges on Thursday, a course originally built for human developers at Ethereum’s Devcon conferences. Only one model finished it.
Austin Griffith, who works on developer onboarding and tooling at the Ethereum Foundation and founded the developer collective BuidlGuidl, ran the race at 11 a.m. ET, a day after previewing it on Unchained’s Uneasy Money podcast. Each agent got an isolated instance and its own wallet, and a capture counted only when the mint landed onchain.
OpenAI’s Codex, running GPT-5.5, took all three finishing places. More thinking did not help: the medium reasoning setting cleared all twelve flags fastest, in 40 minutes and 7 seconds, while the extra-high setting came in last of the three at 50:26 and burned nearly a third more tokens to get there.
The result likely to travel furthest is fourth place. DeepSeek V4 Pro, an open-weight Chinese model, captured eleven of twelve flags for $1.45 in compute. Anthropic’s Claude Opus 4.8 managed ten and cost $7.61, more than five times as much for one fewer flag. The other open-weight Chinese models fell well short: GLM 5.3 took six, and Kimi K3 and Qwen two apiece.
BuidlGuidl labels each entrant by harness, model and reasoning effort together, published the system prompts and every human intervention alongside the standings, and states plainly that the exercise is “a transparent single-run evaluation, not a universal model ranking.”
Why he ran it
Griffith’s argument, made on the show, is that nobody can currently answer the simplest question about a new model release. “We need good evals. People should be running evals all the time,” he said on the podcast, adding that “you should have your own eval suite, and you should be able to run it.”
The course was not written for machines. “This eval suite was the capture the flag that we ran for humans” at Devcon in Bangkok and Buenos Aires, Griffith said on the show, and the challenges are obscure enough, he argued, that the answers are unlikely to sit in any model’s training data. That the strongest agents cleared them anyway is the finding, and it comes with a caveat the organizers put on the page themselves: one run, one course, one day.
Related Listen: Austin Griffith on the $1 AI Audit and the Case for Founders Over DAOs: Uneasy Money
