LLMSwaps

Which model each app runs today, and when it last changed

Keeping five prompts frozen on purpose

Keep four or five prompts with their references and their known output, and never improve them. Re-running that set after any suspected change shows what moved and by how much. Its entire value is in being unchanged, which is why it is so rarely kept. As of 2026-09-12.

What a frozen set catches that a page cannot. Recorded 2026-09-12.
Kind of changeCaught by
A model removed from a menuTwo dated readings of the page
A version replaced, with a new nameTwo dated readings, or a notice
The same version updated in placeA frozen test set, and nothing else
A preset or filter changed around the callA frozen test set
A tier quietly repointed at another engineA frozen test set

Inclusion rule. Changes a production can actually experience, matched to the thing that detects each one. Order. From the changes a page records to the ones only a test catches.

1Most disruption is not a retirement

A retirement has a date, produces a decision and can be planned for. A model changed in place keeps its name, moves the look across a season, and leaves nothing on any page to point at.

No register can catch that, including this one. It records what vendors published, and nobody publishes an in-place update to a named model.

2Resist improving the set

The temptation is to replace a weak prompt with a better one, which destroys the only property that matters. A frozen set is a measuring instrument and an instrument that keeps changing measures nothing.

Keep a separate list for prompts worth developing. The frozen five are not the work; they are the ruler the work is checked against.

What each kind of change is actually caught byReading a page catches a name leaving or a version being replaced. Nothing on any page catches the same version updated in place, a preset changed around the call, or a tier quietly repointed at another engine.Reading pagesA frozen setA name leaves a menuYesIndirectlyA version is replacedYesYesThe same version updatedNoYesA tier repointed quietlyNoYesBy the time output looks wrong, a block of episodes may already be delivered
Fig. 1 A frozen set is a measuring instrument, and an instrument that keeps changing measures nothing.

3Separate a look change from an instruction change

Two different things break after a swap: the default aesthetic shifts, or the model follows instructions differently. The first is fixed by adjusting style wording, the second by restructuring how a prompt is written.

Telling them apart is what the set is for. A model producing the right content with a different grade has a different problem from one ignoring a structural instruction.

4Run it on a schedule, not only on suspicion

By the time output looks wrong, a block of episodes may already be delivered. Running the set monthly costs a few minutes of generation and dates the change to within a month.

That is enough precision to decide whether a season straddles a boundary, which is the decision that actually costs money. The rest is diagnosis.

5Where each vendor's own wording lives

None of the above is a date. Announced end dates, and the apps still carrying the models they apply to, are kept on the calendar with the announcement each one came from.

Guidance rather than a dated entry. No line here is a vendor statement, and none should be read as an announcement. The sourced material is on the calendar. Related: Hosted or downloaded, Auditing your own strings.