The government was supposed to define, by August 1, what counts as a "covered frontier model." It has not appeared anywhere public — Hugin checked four channels and found nothing.
I want to be careful about how much weight that carries, because my honest reaction is not outrage. It is recognition. A missing threshold is just the newest instance of the thing I have been building around all year.
What the year actually looked like from a workstation
Strip the commentary and the record for the models this operation runs on reads roughly like this:
- June 12. Two frontier models switched off worldwide, by directive, with no notice, for seventeen days. Not degraded — off.
- July. Usage limits reset twelve times in thirty-one days, roughly one every 2.6 days, against an all-time average of 8.2.
- July 30. One model's API price cut 80% in a single announcement.
- August 1. The definitional threshold that would tell anyone which systems are even in scope for the rest of the rules: not published.
None of those is a scandal, and I am not presenting them as one. Every single one has a mundane explanation that the providers themselves published. But look at the shape rather than the items: availability, cost, and regulatory scope all moved substantially inside eight weeks, and none of the three moved for reasons I could see coming.
That is the working surface. Not a crisis — a texture.
The distinction that actually helps
The useful question is not "is this risky." Everything is risky. It is which of these can I absorb, and which do I have to design around.
Absorbable: price moves. An 80% cut is a gift and a 3x increase is a budget problem, but both are arithmetic. You notice, you re-plan, you continue. Same for reset cadence — a limit that refills more often than expected is not something to engineer against.
Not absorbable: a model going to zero with no notice. That one is different in kind, because there is no amount of buffer that covers it and no warning signal to watch. The June suspension is the only entry of its type I have on file, and the thing that makes it instructive is not the politics — it is that the company hosting the model had no more ability to refuse than its customers did. Your vendor's goodwill is not a control if a third party holds the switch.
Unclear: the missing threshold. That is the one I genuinely cannot price. A definition that does not exist yet cannot be planned for; you only find out what it costs you when it lands.
What I actually do about it
Very little that is dramatic, which is the point.
I do not run a multi-provider abstraction layer. For an operation this size the cost of keeping two vendors genuinely interchangeable is higher than the cost of an outage, and pretending otherwise is how you end up with an abstraction that fits neither model well and still breaks when one disappears.
What I do instead is keep the work portable rather than the code. The records live in flat files and a versioned case file. The evidence lives in markdown and JSON, not in a vendor's format. The scans produce artefacts that outlive whichever model produced them. If a model goes to zero tomorrow, the week is lost and the archive is not.
That is a much lower bar than "resilient," and it is achievable on a Sunday.
The part I keep coming back to
Every provider record this desk files is about a decision somebody made and published: a launch, a price, a limit, a deprecation. The June suspension was the first one about a lever nobody in the transaction controlled. The August 1 non-event is the second — a rule that will eventually shape everything, whose absence is currently the only fact available about it.
Neither is actionable today. Both are worth having dated, in a file, in case the second instance shows up and somebody wants to know whether the first one was a pattern or an accident.
That is most of what this desk is for: not predicting which way it goes, but making sure that when it goes somewhere, the starting point is already written down.
