Independent reliability evidence

The Meter · Method

How the Meter measures

A model's name can stay the same while the model behind it changes, and nobody has to tell you. The same goes for the route you buy it through. Catching the two takes two different tricks.

Model

Watching a model

You cannot freeze a model. It sits on the company's computers and they can change it whenever they like, without telling anyone.

So we freeze everything else.

The model is then the only thing left that can move, so when the answers move it was not us.

Some companies send our settings back with the answer and some do not, which is why a few checks cover only some models.

Checks we run on a model

CheckWhat it looks forStatusApplies to
The daily examA fixed set of 63 questions, asked every day. This is the one check that publishes a score.publishingEvery model
Is it still the same modelA fixed forced-choice test that each model tends to answer in its own way. Run up to twice a day.private, still collectingEvery model
What the provider says it servedOn every call the provider sends our settings back with the answer. We check them against what we sent.recorded on every callGPT-5.6-Luna, GPT-5.6 Sol
Does the hour matterThe same slice of questions, fired at a different hour each day.private, still collectingEvery model
Does it hold if we ask againThe same question, ten times, on the same day.private, still collectingEvery model
What does it refuseRefusals measured on their own prompts, kept separate from ability.private, still collectingEvery model
Does it remember a long documentRecall from a long input, every third day.private, still collectingEvery model
Does text come back unchangedText sent through the model and checked against what went in. Weekly.private, still collectingEvery model
What the response shows about servingThis provider does not send our settings back. We record what the response does show, and never say a setting was applied.recorded on every callClaude Fable 5, Claude Opus 5, Gemini 3.1 Pro
Google's own description of the modelGoogle publishes a description of this model. We check it daily for edits.recorded dailyGemini 3.1 Pro

Route

Watching a route

Most apps don't reach the AI company directly. They go through a middleman. We call that path a route, and it can change the answer while the model has not changed at all.

So we ask the same question both ways, at the same moment.

Both paths share the model, the questions, the settings and the second, so all of that cancels. What is left over is the route. If the two disagree we check the company on its own before we blame the route.

Checks we run on a route

CheckWhat it looks forStatusApplies to
Does the route answer like the providerThe same 36 questions are sent to the company and the route twice: once for a complete answer and once as the answer arrives, for 72 comparisons.publishingEvery route
Does our request arrive unchangedDoes the route pass on the settings we sent, or drop some without saying so?recorded dailyEvery route
What the route says it offersHas the route changed the list of models it says it can reach?recorded dailyEvery route
Which provider actually served itDid the answer come from the company we asked for?recorded dailyEvery route
The same question down both routes, comparedIf a result changes, we check the company directly before saying the route caused it.recorded dailyEvery route

Sealed

Sealed, and held back

Every set of questions is fingerprinted and stamped to Bitcoin before we ask it once, including the ones whose results you cannot see. Change a question later and the fingerprint stops matching.

A check stays private until it can tell a real signal from our own bugs, finds something you could act on and can be explained to a sharp twenty-year-old.

SEE THE FINGERPRINTS