A prompt can work perfectly on Tuesday and fail on Wednesday, with no change on your side. The cause is that the model underneath moved. This piece is about writing prompts that hold up when the model changes, and about the practice that catches the break early.
The first thing to accept is that the problem is real and structural. OpenAI’s prompting guide is direct about it: content generated from a model is non-deterministic, and even different snapshots of models within the same family can produce different results. Your prompt does not live in a stable room. It lives on top of a moving floor.
Given that, OpenAI’s guide gives two defenses, and both are practices you carry out. Anthropic’s overview supports the second.
The first defense is to pin the model. OpenAI’s guide recommends pinning production applications to specific model snapshots, with a concrete example of the form. The logic is simple: if the model cannot move without your say-so, then your prompt cannot silently break when someone else upgrades the default. Pinning does not make the prompt last forever. It makes the moment of change yours to choose and to test.
The second defense is to test the prompt against real evaluations. OpenAI recommends building tests and evaluation suites that measure prompt behavior, so you can monitor performance as you iterate, or when you change and upgrade model versions. Anthropic frames the same idea from the other direction: before you start prompting at all, you need a clear definition of success and some way to test against it empirically. A prompt with no test is a claim with no check.
These two together are the whole answer. Pin the model so nothing moves under you. Test the behavior so that when you do choose to move, you know whether the prompt survived.
There is a third, quieter practice in the same direction: keep the prompt in code. OpenAI’s guide says to store production prompts in your application code, with typed inputs, code review, tests, and your normal deployment process. The reason is the same as the reason to pin. A prompt that lives in your repository, versioned and reviewed like the rest of the code, changes only through a process you control. It also means the prompt and the test sit next to each other, which is where they belong.
What about the prompt itself? The vendors differ in detail, but they converge on a few stable shapes that outlast any single model.
Use explicit roles and structure. OpenAI’s guide describes message roles with differing levels of authority: developer messages carry the application’s rules and business logic, user messages carry the end user’s input, and the model’s own turns sit in the assistant role. This structure belongs to the interface, so I expect it to outlast any single model, though neither guide promises that. Anthropic’s overview lists role prompting, examples and XML structuring among its techniques and points to a separate best-practices page for them.
Mark your boundaries. OpenAI suggests Markdown headers and lists to mark sections and communicate hierarchy, and XML tags to mark where one piece of content begins and ends. This is a stable convention, not a model-specific idiom.
Provide examples of the desired output. OpenAI’s guide covers few-shot examples, and Anthropic’s overview lists examples among its techniques. An example is a specification. It says what you want more reliably than adjectives do, and it survives a model change better because it is concrete.
The deeper point sits under all of this. A prompt that survives an upgrade is a prompt that leans on the stable parts of the interface and the durable parts of the task, and leans away from lucky phrasing. The lucky phrasing is the part that breaks. The explicit instruction, the pinned model, the versioned prompt, and the test suite are the parts that last.
One honest caveat from Anthropic belongs here: prompt engineering is not always the fix. Its guide says plainly that not every failing evaluation is best solved by prompt engineering, and that you can sometimes improve latency and cost more easily by selecting a different model. When a prompt starts failing after an upgrade, the right move is sometimes to change the prompt, sometimes to change the model, and always to let the test tell you which.
So the practice is short. Define what success means and how to measure it before you write a word. Pin the model in production. Keep the prompt in code, versioned and reviewed. Write tests that assert the behavior, and run them every time the model or the prompt changes. And when the upgrade comes, as it will, the test will show whether your prompt still passes the checks you wrote.
A prompt is a small piece of software with a model as its runtime, and like any software it needs to be pinned, versioned, and tested if it is going to outlive the next release.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.