Note 1Response5 August 2026
The reviewed model is not the served model
We want to raise a different part of it, which we have not seen anyone raise.
Context
On 14 July, Demis Hassabis published a framework for frontier AI oversight. Its centre is a Standards Body "modelled on a federally overseen public-private partnership or self-regulatory organisation, much like the Financial Industry Regulatory Authority (FINRA)," with a board including "independent leading technical experts and open-source representatives," and funding that "would need to be substantial and likely mostly come from industry."
The response was fast and wide. Elon Musk: "It is a thoughtful framework overall and certainly a good starting point for discussions." Sam Altman: "This is a thoughtful proposal from demis." Sundar Pichai: "Well said Demis! Worth reading." Satya Nadella and Mustafa Suleyman said similar things inside a day.
Almost all of the argument since has been about the funding line, and whether a body paid for by the labs it reviews can review them. Yoshua Bengio has written on that, so have Gartner, Fortune and Zvi Mowshowitz, and we have nothing to add to it.
We want to raise a different part of it, which we have not seen anyone raise.
The gap
The framework checks models before they ship.
"Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release. Once the assessment protocol is shown to be effective and robust, formalisation could quickly follow, meaning that Frontier Models would be required to pass it to be deployed in the US market."
Then, for everything that happens after that: "Labs would also work with the Standards Body to address any critical post-release vulnerabilities."
That is the whole of it. Evaluations refresh "perhaps quarterly to start." Nothing in the framework describes how anyone outside a lab would know what a deployed model is doing on a given Tuesday.
What happened
A model can change while its name does not.
Providers adjust reasoning effort, routing, context limits, filtering and serving hardware without changing what the model is called. From outside, behaviour is close to the only thing you can see. Dated model IDs come back model_not_found. system_fingerprint comes back null. So a model reviewed in the thirty days before launch is not necessarily the model answering a request in September, and there is no supported way to check whether it still is.
The day before the framework essay published, this happened in public.
On 13 July an account called Fixlation posted that OpenAI had "reduced GPT-5.6 Sol's thinking budgets in an effort to make the model more efficient," with a table of old and new values: 960 down to 128 at max, 128 down to 40 at extra high, and so on down the tiers. Those numbers are Fixlation's, and we have not verified them.
OpenAI's Tibo Sottiaux replied the same morning, quoting that post. His update opens "Updates for Codex and ChatGPT Work users. No nerfing, only good stuff!" and then, four lines down: "To understand where the extra usage was coming from, we ran some experiments where reasoning efforts were changed (referred to as juice values under the hood) and have reverted this." The same post says the product context limit had been moved to 372k and reverted to 272k.
The part worth looking at is how any of this became known. Someone outside noticed, posted numbers we still cannot check, and got a confirmation because the post travelled far enough to need one. That worked. It worked without a fixed method, without a dated record, without anything anyone could go back and audit later, and it worked because 2.3 million people saw a tweet.
A review conducted up to thirty days before that model shipped would not have covered any of it. Under the framework as written, nothing else would have either.
Already in the text
The framework already contains the answer.
Further down the essay, describing what the body should grow into: "eventually the Standards Body should build up the technical capacity to create its own held-out tests independent of the Labs to prevent overfitting. Working with the US government, it could promote an ecosystem of third-party auditors to help with the assessments and development of new benchmarks and evaluations."
We agree with both halves, and we would move them earlier.
We test deployed models and we watch what they do. Our questions are fixed before we use them and their fingerprint is stamped into Bitcoin, so the date is not ours to move. The earliest line has been watched since 11 July 2026, and the public record of changes to deployed systems has been checked every night since 27 July. Anyone who wants the mechanism can find it on the site, along with the record's own statement that it does not claim to be complete.
Objections
The objections, and what we do about them.
On funding, our answer is a rule rather than a reassurance: we are never paid by anyone we rate. Not labs, not gateways, not the providers whose behaviour ends up in the record.
On jurisdiction, we do not have one. The framework is US-first, with force scoped to the US market and international standards left as something it might "spur." The White House has already answered through Sriram Krishnan that "there will not be an FDA for AI." Measuring what a deployed system does needs nobody's permission to begin, and it does not stop working when the request comes from Bengaluru rather than Boston.
On timing, a pre-release review describes a model at one moment, and what it describes keeps moving afterwards. So the check has to keep running, including on the nights when nothing happens, and it has to publish those nights too.
Not the alternative
We are not proposing ourselves as the alternative.
If the Standards Body gets built and works, it does the thing we cannot do, which is compel. What it cannot do, given how it is funded and where the review sits, is watch continuously from outside. That part has to live somewhere else, and somebody has to have been running it before the day a mandate arrives.
The record is at nlnllabs.com/ledger.