Adam Tornhill’s Post

This is my biggest AI wow moment to date: 300K lines of complex code refactored in three weeks to perfect Code Health. In the past, uplifting a legacy codebase would have been an expensive, high-risk project needing 12–18 months with a team of experts. Now we did it in a fraction of that time for just $4,000 worth of tokens. Some of the highlights from this case study: * The case study demonstrates agentic refactoring at scale on a real-world, non-trivial codebase: Street Fighter III: 3rd Strike. * We refactored the whole codebase to eradicate all application code technical debt. * The code was brought to a level where new features could be added safely with AI. * Functional correctness was verified via a replay-trace harness. (And yes: we could still play the game as before.) * The CodeHealth MCP Server was used as the objective quality signal and agentic feedback loop. But perhaps the most exciting part is the process we used. We had agents iteratively building up a refactoring playbook. That way, agentic refactoring tasks become progressively more effective. As part of that, we discovered novel codebase-specific refactoring rules courtesy of the AI. These were remarkably useful, yet structurally different from the code transformations an expert human would consider. I'll cover all of that in my new article: https://lnkd.in/eBgaX-ET Big thanks to the CodeScene research team, Markus Borg, and Daniel Webb for breaking this new ground.

Adam Tornhill , Daniel Webb According to the article, it seems that code health was about refactoring. Meaning keeping the very same behavior as before. I am curious to know if the "non-functional requirements" have also seen improvements thanks to better code health. The case study is a video game. Performance, memory usage, framerate, input latency, while running the game is a very important aspect in video-game software as I imagine. It could be an useful side-quest to compare the software performance at runtime before/after the code health improvement.

Impressive work. I’d be interested in understanding the human involvement and cost breakdown behind the results - were the two contributors the only humans in the loop, or did others help with setup, supervision, testing, or review? Congrats, this is the kind of result that would have seemed extraordinary just a year ago.

Like
Reply

I’m afraid that some of your discovered novel codebase-specific AI’s refactoring rules seem to ignore that DRY is actually about duplication of knowledge and intent, not identical lines of code.

Like
Reply

So, the biggest pushback we have faced for refactoring sweeps is not having enough behaviour-validating testing. In the agentic era, this mentality also pushed our agents to write a bulk of test cases before attempting the refactoring sweep. However, those tests end up being implementation validation rather than behaviour validation. Surprisingly, in the blog, we see testing is described at the screen level or UAT level. Was there any intentional design decision to avoid a high focus on testing?

Like
Reply

What is the definition used for "Refactoring" in this experiment? Most of my refactorings produced simpler software designs, but for this I needed to go above the statement level ...

Like
Reply

We did similar. Codescene mcp plus the high quality agents today is almost like cheating. 😂

Andrei Lungeanu

AI/ML Engineer | Full-Stack Developer | DevOps & Infrastructure | 10+ yrs building scalable, intelligent platforms

5d

300K lines in three weeks is a big result, but the replay-trace harness and CodeHealth feedback loop are the parts that make it believable. Speed is useful only when you can keep checking that the refactor still behaves like the original.

Did it also include unentangling spaghetti code, for which you would also might need business people as features/requirements have long ago been forgotten?

Like
Reply

[Devil's advocate questions]: 1. Is the refactoring already merged to master? 2. If merged, was it 1 giga merge or thousands of them over the period? The closer to 1 then the closer it to be a rewrite (not Refactoring defined in Fowler's book). 3. Am I right that this was actually "just" experiment on the open source code, if I understood the article correctly? Not a production code for a product that is actively generating money like some SaaS ect?

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories