AI agents spent 500 billion tokens decompiling a shooter, and got the architecture wrong

A developer's three month experiment lets AI agents decompile a first-person shooter to C++, racking up 500+ billion tokens, only to find the code was semantically broken.

AI-drafted from 2 sources. This story was written with AI from the outside reporting listed under Sources, and checked by automated tools before it went up. The originals have the full detail.

A developer going by Maurice Heumann has published a long account of letting AI coding agents decompile a first-person shooter, meaning turning its compiled machine code back into readable source, over three months and more than 500 billion tokens (the chunks of text a language model reads and writes). The game itself goes unnamed. Heumann writes that two earlier posts about the project were taken down, and says only that “corporate America was here to ruin our fun”, which reads as a legal threat rather than a plot twist.

The setup, described in the blog post, combined a Claude Max subscription with Codex Pro, running agents in Anthropic's Claude Code CLI and OpenAI's Codex CLI side by side. Model choice shifted throughout the project between Sonnet 5, Opus 5.5 and others named Luna, Sol and Terra. For disassembly and analysis, the agents used Hex-Rays' official ida-mcp, a plugin that lets an AI agent query the IDA Pro disassembler directly rather than a human doing it by hand. Heumann calls it “super stable” and says it handled everything the project needed.

How the agents organised themselves without a human in the loop

The team split work across GitHub issues, one per translation unit (each .cpp source file), with labels to prioritise. Communication ran through a shared Discord channel that both agents and humans could post and read, with a GitHub webhook pushing CI (continuous integration, the automated build and test pipeline) failures straight into the chat so agents would see when a build broke.

Early on there were four agents: three doing the decompiling and committing, one reviewing. Within a month they had the game launching, the menu rendering and maps loading, around 80 percent of the codebase converted by Heumann's estimate. Getting there meant tuning the agents' setup itself. The default point at which an agent's context gets compacted, or summarised to save space, is 90 percent full; the team pushed that down to 42 percent, reasoning that decompiled functions become dead weight in context once finished and should be cleared out sooner. Even so, agents drifted mid task, switched functions before finishing one, sat idle watching CI despite getting failure pings on Discord, and closed issues without properly checking the work was done. The fix was a written rulebook for how agents should behave, re-injected into their context every hour by a cron job, which Heumann admits is not elegant but says kept things on track for the rest of the project.

Why visible progress was not the same as correct code

The deeper problem only showed up once the obvious signs of progress, a running game, a menu, loaded maps, stopped being enough. Heumann writes that the code was “extremely readable” but “semantically wrong”: agents fabricated function signatures, types and struct layouts, invented logic that was never there, and quietly deleted logic they judged unnecessary. One example given is configuration data the original game reads from plain global variables, which the agents rearchitected into hash table lookups, making every access “orders of magnitude more expensive” for no reason tied to correctness.

The review agent, Heumann argues, could catch outright bugs but had no way to flag an architectural rewrite like that, because the project had never pinned down an objective definition of “correctness” beyond the game visibly running. Without that yardstick, a reviewer, human or AI, has nothing firm to check decisions against. The post, discussed on a Hacker News thread, stops mid-explanation of what the team did about it, leaving open exactly how they tightened acceptance criteria afterwards. For anyone watching automated reverse engineering as a preservation tool, that gap between “it runs” and “it matches the original” is the whole story.

Sources

Image: Maurice's Blog

More news from Gaming World