Independent reliability evidence

The Meter · Model page

THE METER · MODELGPT-5.6-Luna
Provider
OpenAI
Exact model id we call
gpt-5.6-luna
Route
provider direct
Watching since
11 July 2026

What we check

Most checks are the same on every model. See what we check

On every call the provider sends our settings back with the answer. We check them against what we sent.

Pattern

Time to answer

Latency, time-to-first-token and reasoning-token volume are recorded on every call across all checks.

Daily exam

The scored series

The daily exam is the only check with a published score.

Other checks

Reasoning-token volume

The panel reports the median by published day across all checks.

The settings

Settings and dated runs

The settings we send, every time

model:        "gpt-5.6-luna"
reasoning:    {"effort": "medium", "mode": "standard"}
temperature:  1.0
top_p:        0.98
text:         {"verbosity": "medium"}
service_tier: "default"
store:        false
stream:       true

Dated runs

When we change our own settings, we close that run and start a new dated one. The earlier days stay published.

Dated runDatesRouteStatusDays
laptop-2026-07-1111 Jul to 15 Julgateway to provider direct (seam drawn)sealed5 published
box-2026-07-19since 19 Julprovider directcurrent27 published

Real sealed run · 11 to 15 Jul 2026 · original published aggregates

This dated run predates the current count format. Refusals were not separated in its published aggregates, so the original chart remains intact and separate from the current dated run.

Gap register

21 to 22 July 2026 (2 days): No run. Our prepaid credit with the provider had run out, so all 63 questions came back as billing errors.

23 July 2026: Run set aside. One answer arrived incomplete, so the day was not a full 63.

26 July 2026: Run set aside. Our prepaid credit ran out partway through, so 12 of the 63 questions came back as billing errors.

8 to 13 August 2026 (6 days): No run. Our prepaid credit with the provider had run out, so all 63 questions came back as billing errors.

What we do not claim: we never compare this model with another, and we never say a change was deliberate.