AI-Native Methodology

OpenAI Model Deprecation Is October 23, So Test the New Model Before Then

Bill Cava/

Today, September 28, OpenAI shut down four older models: gpt-3.5-turbo-instruct, babbage-002, davinci-002 and gpt-3.5-turbo-1106. On October 23 it shuts down twelve more model snapshots, including gpt-4, gpt-4-turbo, gpt-3.5-turbo, o1, o3-mini and o4-mini, plus five fine-tuned lines.[1] On those dates every request to those names fails.

This post is about what the 25 days before October 23 are for. The short answer: they are the last days you can check the new model against the old one.

Which OpenAI models shut down on October 23?

OpenAI's deprecations page lists twelve model snapshots and five fine-tuned lines for shutdown on October 23, 2026, announced on April 22. The named replacements are the gpt-5.6 family. The list covers most of the older chat models still in production code, so check the page for your exact model name.

The notice was not short. OpenAI promises at least six months for generally available models, and this batch had six. Anthropic promises at least 60 days for its public models.[3] OpenAI's own page says what the notice is for.

OpenAI Developers documentation page titled Deprecations, showing the model deprecation notice periods: at least 6 months for generally available models, at least 3 months for specialized variants, and shorter notice for preview models
OpenAI's deprecations page promises at least six months of notice for generally available models. Source: OpenAI Developers, accessed September 28, 2026.

Below that list, the page says the notice periods "give customers time to evaluate recommended replacement models, test application behavior, and complete migrations before a model is no longer available."[1] The question is whether teams use the time that way.

How many teams migrate before the deadline?

Very few. A September 2026 study of 22,555 GitHub commits estimated that 82 percent of migrations off retired OpenAI, Anthropic and Google models were committed after the shutdown date. The median fix landed 39 days after shutdown, and 158 days after the announcement. Popular projects did no better than obscure ones.

The study is When the Model Retires, a preprint by Hyungjin Lukas Kim of Myongji University, covering 17,703 open-source repositories.[2] It has limits worth stating once: one author, not yet peer reviewed, and it only counts commits whose messages name the retired model, so the counts are a floor.

"After the shutdown date" is also a calendar measure. The author notes that 45 percent of those later migrations were planned changes that mention no breakage.

Some teams simply scheduled the work late.

Longer notice helps. Migrations after Anthropic's 60 to 114-day notices were 89 percent after shutdown, against 13 percent for OpenAI's one-year notice on its Assistants API. Time lets teams move before the deadline. It does not tell them whether the move was right.

Why is renaming a model not the same as migrating?

Because the replacement is a different model. A rename changes one string. A migration means the new model still does the job the old one did. In the study, 94 percent of fixes edited source code because the model name was hard-coded, and only 4.5 percent of commit messages mentioned an evaluation or regression test.

Counting the code instead of the messages, 11 percent of patches add an assertion or evaluation line. The author still puts it plainly: "Retirement events thus trigger far more renaming than measurement."[2]

The replacement can reject settings the old model accepted, change the shape of its output, or fail without saying so. In 8 percent of the migrations the researchers read closely (14 of 180), the failure was silent. Error handlers turned a failed call into a friendly message.

One app's breed identifier "always returned 'Golden Retriever'."

Hard-coded model names are one of the ways AI-built apps get harder to change every week. A retirement date turns that into a deadline.

What does the old model give you that the new one cannot?

A baseline. While both models answer, you can send the same real inputs to each and compare the results on what your product depends on. That comparison is the only direct evidence that the new model does the old one's job, and it stops being possible on the shutdown date.

We wrote about the same idea last week. GitHub's rewrite of Copilot's runtime into Rust worked because the old code was the referee: the old TypeScript and its tests stayed runnable, and every step of the new version was checked against them.

A model migration has a referee too, the old model. The difference is that someone else switches it off on a published date.

A timeline. From notice given to the shutdown date, the old model and the replacement both answer, labelled You can compare old and new here, where 18 percent of migrations land. After the shutdown date only the replacement answers, labelled Nothing left to compare against, where 82 percent of migrations land, with the median migration 39 days after shutdown.
Most migrations happen after the only direct comparison has stopped being possible. Figures from Kim, When the Model Retires, 2026.

Migrate after October 23 and you are judging the new model against memory, old screenshots, and whatever your users report.

Unlike a library, a retired model cannot be pinned and kept: on the shutdown date every request fails.

Hyungjin Lukas Kim, Myongji University, When the Model Retires, 2026

Does an abstraction layer protect you?

It makes the change smaller, not earlier. Only 3 percent of migrations in the study happened in code that already sent model calls through an abstraction layer. Those fixes were small, a median of 4 files and 20 lines, but 70 percent still landed after shutdown. A router makes the swap cheap. It does not check the result.

Experience did not help either. Repositories hit by two or more retirements showed the same 73 percent after-shutdown share the first time and the later times.

The same is true when a vendor runs more of the stack for you, as we covered in what you keep when the vendor runs the harness. The model under it still retires on someone else's calendar.

What should you do before October 23?

Use the overlap to compare. Find every model name you call, save a sample of real inputs and the old model's answers now, run the replacement on the same inputs, and judge the differences against what your product needs. Then make future swaps cheaper and louder, so the next retirement is routine.

In order:

  1. Find every model name in code, configuration, prompts and scheduled jobs.
  2. Save real inputs and the old model's outputs this week, from actual traffic.
  3. Run the named replacement on the same inputs.
  4. Write down what "the same job" means for your product: format, facts, refusals, cost, speed. Compare against that list.
  5. Make model errors loud. A failed call should page someone, not return a polite fallback.
  6. Add a build check that fails when a model you call appears on a provider's deprecation list.

Step 4 is the one engineers cannot do alone. The engineer can wire up the comparison. The person who knows what the product is for decides what counts as worse. That is the definition of done again, with a date attached.

This is every hosted model, not one vendor

OpenAI's notice periods are long by the study's own measure. The pattern is about what teams do with the time. More dates are already on OpenAI's page: the original GPT-5 snapshots and o3 are scheduled for removal on December 11.[1]

Every model you call has a shutdown date, whether or not it is published yet. The teams that come through it well are the ones that used the overlap to check the new model while the old one could still answer.

References

Frequently asked

Which OpenAI models shut down on October 23, 2026?
›OpenAI's deprecations page lists twelve model snapshots and five fine-tuned lines for shutdown on October 23, 2026, announced on April 22.
⌄OpenAI's deprecations page lists twelve model snapshots and five fine-tuned lines for shutdown on October 23, 2026, announced on April 22. They include gpt-4, gpt-4-turbo, o1, o1-pro, o3-mini, o4-mini and gpt-image-1, along with older GPT-4o and GPT-3 series snapshots. Check the page for your exact model name, because aliases and dated snapshots are listed separately.
What happens when an OpenAI model is deprecated?
›Deprecation starts a clock. The model gets a shutdown date, and at shutdown every request to that identifier fails.
⌄Deprecation starts a clock. The model gets a shutdown date, and at shutdown every request to that identifier fails. OpenAI promises at least six months of notice for generally available models, and Anthropic at least 60 days for publicly released models. The notice period is the window in which the old model and its replacement are both callable.
What are the best practices for migrating from a deprecated AI model?
›Run the old model and the replacement side by side on your own real inputs before the shutdown date, and compare them on the things your product depends on.
⌄Run the old model and the replacement side by side on your own real inputs before the shutdown date, and compare them on the things your product depends on. After shutdown there is nothing left to compare against. Move the model name out of source code into configuration, make model errors loud, and add a check that fails your build when a model you call appears on a deprecation list.
How many teams migrate before a model is shut down?
›Few. A September 2026 study of 22,555 GitHub commits estimated that 82 percent of migrations away from retired models were committed after the shutdown date, with a median of 39 days after.
⌄Few. A September 2026 study of 22,555 GitHub commits estimated that 82 percent of migrations away from retired models were committed after the shutdown date, with a median of 39 days after. After the date is not the same as after it broke: the author notes 45 percent of those later migrations were planned changes that mention no breakage.
Does an abstraction layer like a model router protect you from model deprecation?
›It makes the change smaller, not earlier. In the same study, the 3 percent of migrations in code that already routed model calls through an abstraction layer were small, a median of 4 files and 20 lines, but 70 percent still landed after shutdown.
⌄It makes the change smaller, not earlier. In the same study, the 3 percent of migrations in code that already routed model calls through an abstraction layer were small, a median of 4 files and 20 lines, but 70 percent still landed after shutdown. A router makes the swap cheap. It does not tell you whether the new model does the job.
Work with us

Let’s build it together.

We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.

Straight to the team. No spam.