Why AI Projects Still Fail After Twenty Years of Better Models
I have been working in AI since my graduate research in 2005. In 2015 I co-founded Orbita, a voice AI company for healthcare, and ran product and engineering there. Over those twenty years the field replaced its main technique twice. The things that decided whether a project shipped did not change at all.
What did AI look like in 2005?
In 2005, AI meant search, planning and a lot of knowledge typed in by hand. Working systems combined path-finding algorithms, constraint solvers, rule-based expert systems and planners, with statistical learning as a junior partner. Getting a demo was hard, and getting a system into production was harder.
The toolbox had names nobody uses at dinner parties: A-star search, constraint satisfaction, STRIPS-style planners, support vector machines, conditional random fields. A useful system needed hand-built features and small labeled datasets that cost real money to produce. Most of the intelligence was written down by people who understood the domain.
Healthcare added one more layer. Every system that touched a patient's words touched protected health information, and that was true from the first prototype.
What actually changed?
Nearly everything about the technique changed, and it is a real revolution. Deep learning replaced hand-built features with learned ones, and large pretrained models replaced most of the hand-typed knowledge. Work that took a team a year of rules and features now starts with a prompt. Anyone who says otherwise is not paying attention.
The interface changed too. People ask in plain language instead of filling in a form. Agents call tools, read documents and take multi-step actions. A demo that would have taken months in 2005 takes an afternoon.
For a practitioner from 2005, the right reaction to that is delight.
Are search and planning really coming back?
Yes, inside the neural models rather than beside them. Modern agents spend compute at answer time exploring options, checking candidates and deciding whether to plan before they act. That is search in a new form, and there is now active research on exactly when an agent should stop and plan.
The person who predicted this shape is Rich Sutton. In 2019 he looked back over 70 years of AI research and named the two methods that keep winning as computers get faster.[1]

The two methods that seem to scale arbitrarily in this way are search and learning.
The last twenty years ran that experiment in both directions. In 2005 we leaned on search. From roughly 2012, learning did almost everything and search looked like a museum piece. Now the two are recombining.
What the 2026 research says about planning
A 2026 study of agents that learn when to plan found that planning before every action "is computationally expensive and degrades performance on long-horizon tasks, while never planning further limits performance."[2] Knowing when to search is the skill.

Anyone who understood why planners were expensive in 2005 already understands why agent loops are expensive now.
I would not claim we were right all along. Classical planning is not coming back intact. The instincts transfer, and that is enough to make the old skill set useful again.
What stayed the same for twenty years?
Three things never moved: testing against real cases, a person checking the output where mistakes are costly, and the law governing the data. The model architecture was replaced twice while those stayed fixed, and they are the three that decide whether an AI project ships or stalls.
Testing
A demo is not evidence. In 2005 and today you need real cases, held-out data, and a way to know when the system got worse. The failure looks the same as it did then, only faster and better dressed.
A person in the loop
Wherever a wrong answer is expensive, someone still checks it. The threshold moved as models improved. The requirement did not, and the same 2026 study found its agents did better when "steered by human-written plans, surpassing their independent capabilities."[2] We have written about why that review has to be designed into the system, not left to vigilance.
Data law
This is the one people underestimate, and the dates show why.
- FERPA, which governs student records, dates from 1974.[5]
- HIPAA, which governs health data, dates from 1996.[3]
- GLBA, which governs financial customer data, dates from 1999.[4]
- European data protection ran on a 1995 directive until GDPR replaced it in 2018.[6][7]
HIPAA governed the health data a 2005 system touched, and it governs the health data a 2026 agent touches.
In voice AI for healthcare, that stops being abstract. A patient describing a symptom out loud creates health information the moment the audio is captured, whatever model transcribes it. The models that turn speech into text have changed more than once since 2015, and the rules for the recording have not.
Two full turnovers in technique, and the only change in the law column is Europe replacing one data protection regime with a newer one.
Why do AI projects still fail if models got so much better?
AI projects still fail mostly on data and ownership, not on the model. Unclear provenance, no rule for how long data is kept, no agreement on what the output may be used for, and no named person accountable for a wrong answer. Better models made the demo easier and fixed none of those.
The specific failures are rarely dramatic:
- Nobody can say where a training or reference dataset came from, so nobody can say whether it may be used.
- There is no retention rule, so data from a pilot lives forever in a vendor's logs.
- The output goes to customers with no one assigned to answer for it when it is wrong.
- The evaluation set was the demo, so nobody notices when a model update makes things worse.
None of those is a model problem, which is why swapping models does not fix them. On the software side, the same pattern holds: AI changed how we build, not what makes software work.
The work that decides whether an AI project ships was never in the model.
Where should a team put its effort now?
Put effort where the constant is. Your model choice will look dated within a couple of years. Your evaluation set, your review process and your data policy will outlive several generations of models, and they are what make the next model cheap to adopt instead of a rebuild.
Teams that built those three things early swapped in new models almost for free. They changed one component, ran their tests, and kept their review and data rules. Teams that did not are rebuilding from scratch each time the technique turns over, and in this field it turns over often.
That is the lesson of twenty years, and it costs nothing to act on today. Write down what your system is tested against, who checks its output, and which law governs its data. Those answers will still be true after the next model arrives.
References
- ^
- ^2.Paglieri, Cupiał, Cook, Piterbarg, Tuyls, Grefenstette, Foerster, Parker-Holder, Rocktäschel (arXiv), “Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents (v3)” (February 17, 2026)
- ^3.U.S. Department of Health and Human Services (ASPE), “Health Insurance Portability and Accountability Act of 1996”
- ^4.U.S. Government Publishing Office, “Public Law 106-102 (Gramm-Leach-Bliley Act), November 12, 1999”
- ^5.Legal Information Institute, Cornell Law School, “20 U.S. Code § 1232g (FERPA), as added August 21, 1974”
- ^
- ^
Frequently asked
How is AI different from twenty years ago?›The technique changed almost completely. In 2005 the working tools were search, planning, constraint solvers, expert systems and early statistical learning, and getting a system to work meant encoding a lot of human knowledge by hand.
What has stayed the same in AI over twenty years?›Three things, and they are the three that decide whether a project ships.
Are symbolic AI and classical planning obsolete?›No, and planning is coming back inside modern agents. Agents spend compute at answer time exploring options and deciding when to plan before acting, which is search in a new form.
Why do AI projects still fail if the models got so much better?›Because the failure was rarely the model. It is usually the data and the ownership: where the data came from, what you are allowed to do with it, how long you keep it, whether anyone checked the output, and who answers for a wrong one.
What should a team learn from twenty years of AI history?›Put your effort where the constant is. Model choice goes stale within a couple of years.
Let’s build it together.
We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.