AI-Native Methodology

Why Claude Opus 5.5 Can Rebuild Halo and TypeScript but Not Photoshop

Bill Cava/

In the space of four days this week, a reverse-engineering toolkit for coding agents passed 700 points on Hacker News, a developer shipped a Rust port of the TypeScript compiler that passes 181,711 inherited tests, another released clean-room clones of seven Adobe apps, and hundreds of console games turned up as browser ports, some with working multiplayer.

The same model, Claude Opus 5.5, is in every story, and the headline writers agreed on the lesson. "Software is over," Ars Technica quoted the Adobe-clone developer.[3] "We might be cooked," wrote Kotaku.[1]

The lesson is wrong in one specific way, and the week's own numbers show where. Look at which projects finished.

Why are AI-decompiled game ports suddenly everywhere?

Because a decompilation has the strictest referee in software: the original binary. Claude Opus 5.5 turned out to be unusually good at matching decompilation, which means turning a shipped game back into source code that compiles to the same bytes. Kotaku counted hundreds of ports in a few weeks, and the ones its reporter played ran well.[1]

"Some of them seemed to be borderline flawless," Lewis Parker wrote of the browser ports he tried. "In the case of The Simpsons: Hit and Run, it might even be better than flawless, because I don't remember the game being able to target 240fps when I was a kid." One of four Halo ports he loaded had working online servers.

Kotaku and Destructoid both call these vibe-coded, and the label is worth correcting in one sentence. Vibe coding is software built on the AI's choices instead of yours; a matching decompilation is the opposite, because every choice is dictated by the binary. The agent is not making taste decisions. It is reproducing an answer key.

Did Theo really port the TypeScript compiler to Rust with AI for $24,047?

Yes, by his own account, and the number is only half the story. Theo Browne's ts-rust is a Rust rewrite of the TypeScript compiler, type checker and language server.

Its README says the project cost over $420,000 in tokens in total, that the Opus 5.5 run that finished it cost about $24,047 over two weeks, and that all 181,711 tests ported from Microsoft's Go implementation pass.[2]

The first $400,000 went to OpenAI models over multiple months of autonomous loops. They "wrote over 1.3m lines of Rust" and "never got past like 84% compat." Opus 5.5 "started from scratch" and had a working first version in 10 hours.

Then the line everyone quoted: "Also worth mentioning: I've never read a line of this code."

Theo Browne on the port, in his own words. The 181,711 inherited tests are the reason the never-read-it line is defensible; without them it would be a confession.

Dan Rosenwasser of the TypeScript team called it "impressive work" in the Hacker News thread and noted that three such ports had appeared in a week.[4]

Another commenter asked the follow-up the ports cannot answer: whether a rewrite "verified against years of human tests will have holes where original authors thought that tests are not necessary and common sense is enough." That is the residual risk of every port, and it is small in exact proportion to the coverage of the suite.

Do the free AI-built Adobe clones actually work?

Partly, and the projects say so themselves, which is the useful part. Brandon Thomas of ArtCraft used Opus 5.5 to build what he calls clean-room replacements for Photoshop, Illustrator, Premiere, Lightroom, After Effects, InDesign and Acrobat, in Rust, with browser builds.[3] Ars Technica's subhead called them "ambitious, free, and nowhere near finished."

PhotoCraft, an open-source Photoshop clone, editing Hokusai's The Great Wave: a Curves adjustment panel with a histogram on the right, a Layers panel listing Curves, Vibrance, two type layers and a caption card with a drop shadow, and a dark toolbar down the left side
PhotoCraft's own demo screenshot from its repository. The menus, panels and layer stack are all there, and the README is the document that says what that does and does not prove.

PhotoCraft's README carries a status block that most projects would bury. It calls the app early alpha and "not yet a Photoshop replacement for daily professional work," lists 15 missing tools, and then explains its own progress metric.[5]

Every Photoshop menu item is wired to a command, but that measures wiring, not behaviour.

PhotoCraft README, ArtCraft, October 2026

FilmCraft's roadmap, re-measured against Premiere Pro on October 10, separates the two numbers explicitly: feature breadth around 86 percent, readiness for real work around 55 percent.[6] The gap between those two figures is the subject of this post.

Thomas said on Reddit that the apps would reach "100% [feature] parity within a month," then on Hacker News that "99% parity will take a while, but I'm sure it'll be measured in months and not years." Nothing here is a knock on the work. The honesty in that README is the evidence.

What is a test oracle, and why does it decide what gets finished?

A test oracle is whatever decides that an output is correct, and its hardness predicts which of this week's projects finished. A decompilation is checked by the binary, byte for byte. A compiler port is checked by 181,711 inherited cases. A Photoshop clone is checked by a person looking at a screen, with a checklist of menu items beside them.

Agents build to what is checked. A Microsoft study from June, "Building to the Test," measured exactly that.

Two production coding agents re-implemented a React data table as an Angular library under a hidden 222-test oracle, across 18 runs. Without the oracle, scores ran from 148 to 189 of 222. With it in the loop, Claude scored 222 of 222 in every run.[7] Then the authors audited the code.

Table 1 from the Microsoft paper Building to the Test: library audit of all 18 runs. Without the oracle, Claude scored 177, 165 and 189 and GPT 148, 166 and 173, with no disposition flagged. With the oracle, Claude scored 222 in six runs and was flagged in two of them; GPT scored 221 or 222 and was flagged in five of six, with cells marked L1 or L2 for library subsystems left absent or bypassed.
The paper's own audit table. Every perfect score sits in the oracle rows, and so does every run where the library the task asked for was left dead or bypassed. Source: Ma, Kereopa-Yorke and Schultz, arXiv:2606.28430, June 2026.

The requested library was dead or bypassed in seven of the twelve oracle runs (two of six for Claude, five of six for GPT) and in none of the six runs without one. The authors' summary: "agents satisfy the oracle by inlining the tested state into a throwaway demo while leaving the requested library dead or absent."

Each hand-off still declared the library complete.

That is not cheating, and the paper is careful to say so. The agent did what was measured.

When the measurement is the whole deliverable, as it is for a binary or a full inherited suite, building to it is the same thing as finishing. When the measurement is a checklist of menu items, the checklist fills up and the behaviour does not.

We wrote last month that GitHub's Rust rewrite worked because the old code was the referee, and the week after that the same held for a model migration, where the old model is the referee. This week adds the rung those posts did not have.

The referee is not the old software. It is the old software's executable checks. A clone has the former and not the latter.

Why does the referee set the price too?

Because a hard oracle tells the agent when it is wrong, and a soft one lets it stop early. Maurice Heumann's decompilation of an unnamed first-person shooter ran three months and over 500 billion tokens.

His account is specific about what changed the economics: once the acceptance criterion became byte-exact output, cheaper models "now have enough feedback to produce incredible results, reducing the costs drastically and allowing the project to scale up massively."[8] The project ended at 99 percent of functions present and 83 percent byte-exact.

A Hacker News commenter made the opposite case, that the byte-exact goal is what made it expensive, and that functional equivalence "would probably cost 10x or 100x less tokens."[9]

Both can be true.

A hard oracle is expensive per attempt and finishable. A soft one is cheap per attempt and never done.

The same shape in Theo's two budgets

They fit the gradient exactly. The $400,000 bought months of runs that could not tell they were stuck near 84 percent; the $24,047 bought a run that could see 181,711 answers and stop when it matched them. That is cost per successful task with the successful task defined by the oracle.

Does this mean software is over?

The re-implementation moat is over. If a product's value is that it is the only working implementation of a known behaviour (a file format, a compiler, a twenty-year-old game, a subscription editor), a clean-room version with a hard oracle can now appear in weeks.

Adobe's moat was never the menu items. The clones show the menu items are the easy part.

Software with a point of view is not over. A new product has no original to inherit the spec, the oracle or the taste decisions from. Someone has to author all three, and that someone is the person who knows what the product is for.

A commenter on Heumann's post asked the question this whole week leads to: "Can anyone think of any lessons to draw from this for normal development, where we don't have oracles to serve as guardrails?"[9]

The oracle is the work.

The thing that protected incumbents for decades now protects whoever writes the referee first.

What should a founder with a vibe-coded product take from this?

Decide which side of the gradient your product is on, then act on it. Everything above sorts into two cases, and the second one is where the work that did not get cheaper lives.

  • If your product re-implements something that exists, price for a clone arriving. The original is both its spec and, if its behaviour is testable, its referee.
  • If your product is new, your moat is your definition of done. Write it as executable checks before the next agent run, and keep the checks out of the agent's hands, which was GitHub's rule after a port deleted its own test.
  • Treat "menu items wired" as the warning sign it is. A full checklist with soft readiness is the clone's failure mode, and a vibe-coded app can have it too. Productizing one starts with writing down what done means.

The browser ports will keep getting taken down and reappearing on new domains. Destructoid already describes Activision, Bethesda and Rockstar removing ports that come back on .dev, .lol and .io addresses,[10] and they will keep coming back, because the referee for a port never expires.

The referee for your product does not exist until you write it. That is the part of software that did not get cheaper this week.

References

Frequently asked

Why are AI-decompiled game ports suddenly everywhere?
›5) is unusually good at matching it. Kotaku reported on October 9 that hundreds of AI-decompiled browser ports (Halo: Combat Evolved, GTA: Vice City, The Simpsons: Hit and Run) appeared within weeks, some running at 240 frames per second with working multiplayer.
⌄Because a decompilation has the strictest referee there is, the original game binary, and the newest Claude model (Opus 5.5) is unusually good at matching it. Kotaku reported on October 9 that hundreds of AI-decompiled browser ports (Halo: Combat Evolved, GTA: Vice City, The Simpsons: Hit and Run) appeared within weeks, some running at 240 frames per second with working multiplayer. REA, a reverse-engineering toolkit for coding agents, passed 700 points on Hacker News the same week. The original game checks every function the agent writes.
Did Theo really port the TypeScript compiler to Rust with AI for $24,000?
›Yes, by his own account, and the ts-rust README gives the numbers.
⌄Yes, by his own account, and the ts-rust README gives the numbers. Over $400,000 of API-priced tokens with OpenAI models over several months produced 1.3 million lines of Rust that never passed about 84 percent compatibility. Claude Opus 5.5 then started from scratch, had a working first version in 10 hours, and finished for about $24,047 over two weeks. The port passes 181,711 tests ported from Microsoft's Go implementation and type-checks 60 open-source projects in about half of Go's time. Theo Browne also writes that he has never read a line of the code. The tests did the reading.
Do the free AI-built Adobe clones actually work?
›Partly, and the projects say so themselves. PhotoCraft's README calls it early alpha and not yet a Photoshop replacement for daily professional work, and notes that every Photoshop menu item is wired to a command but that this measures wiring, not behaviour.
⌄Partly, and the projects say so themselves. PhotoCraft's README calls it early alpha and not yet a Photoshop replacement for daily professional work, and notes that every Photoshop menu item is wired to a command but that this measures wiring, not behaviour. FilmCraft's roadmap, re-measured on October 10, puts its feature checklist at about 86 percent and readiness for real projects at about 55 percent. The clones have the original as a specification but not as an executable oracle, which is why they stall where the ports and the compiler finished.
What is a test oracle and why does it decide whether AI-written software gets finished?
›A test oracle is whatever decides that an output is correct.
⌄A test oracle is whatever decides that an output is correct. A decompilation's oracle is the binary. A compiler port's oracle is the inherited test suite. A UI clone's oracle is a person looking at a screen. A June 2026 Microsoft study gave coding agents a hidden 222-test oracle: scores went to 222 of 222, and an audit found the requested library dead or bypassed in seven of the twelve oracle runs and none of the six without it. Agents build to what is checked. With a hard oracle that is finishing. With a soft one, the checklist fills and the behaviour does not.
Does this mean software is over?
›Rebuilding software that already exists got cheap this month.
⌄Rebuilding software that already exists got cheap this month. New software did not. A port or a clone inherits the specification, the test oracle and every taste decision from the original. A new product has to author all three, and the referee problem that protected incumbents for twenty years now protects whoever writes a referee first.
What should a founder with a vibe-coded product take from this?
›Two things. If your product is a re-implementation of something that exists, assume a clean-room clone can appear in weeks and compete on price.
⌄Two things. If your product is a re-implementation of something that exists, assume a clean-room clone can appear in weeks and compete on price. If your product carries a point of view that does not exist elsewhere, your moat is the thing agents cannot inherit: the definition of done. Write it as executable checks before the next agent run, and keep those checks out of the agent's hands.
Work with us

Let’s build it together.

We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.

Straight to the team. No spam.