Independent reliability evidence

The Meter · Model page

THE METER · MODELClaude Opus 5
Provider
Anthropic
Exact model id we call
claude-opus-5
Route
provider direct
Watching since
27 July 2026

What we check

Most checks are the same on every model. See what we check

This provider does not send our settings back. We record what the response does show, and never say a setting was applied.

Pattern

Time to answer

Latency, time-to-first-token and reasoning-token volume are recorded on every call across all checks.

Daily exam

The scored series

The daily exam is the only check with a published score.

Other checks

Reasoning-token volume

The panel reports the median by published day across all checks.

The settings

Settings and dated runs

The settings we send, every time

model:             "claude-opus-5"
max_tokens:        128000
output_config:     {"effort": "medium"}
temperature:       1.0
service_tier:      "standard_only"
thinking:          {"type": "adaptive", "display": "omitted"}
inference_geo:     "global"
stream:            true
anthropic-version: "2023-06-01"

Dated runs

When we change our own settings, we close that run and start a new dated one. The earlier days stay published.

Dated runDatesRouteStatusDays
box-2026-07-27since 27 Julprovider directcurrent9 published

Gap register

30 July 2026: No run. Our prepaid credit with the provider had run out, so all 63 questions came back as billing errors.

2 August 2026: Run set aside. Our prepaid credit ran out partway through, so 46 of the 63 questions came back as billing errors.

3 August 2026: No run. Our prepaid credit with the provider had run out, so all 63 questions came back as billing errors.

5 August 2026: Run set aside. One question ran past our twenty minute limit, so the day was not a full 63.

8 to 13 August 2026 (6 days): No run. Our prepaid credit with the provider had run out, so all 63 questions came back as billing errors.

14 August 2026: Run set aside. One of the 63 questions came back as an error on the provider's side. We publish a day only when all 63 questions have been answered.

15 August 2026: Run set aside. Our prepaid credit ran out partway through, so 53 of the 63 questions came back as billing errors.

16 to 17 August 2026 (2 days): No run. Our prepaid credit with the provider had run out, so all 63 questions came back as billing errors.

18 August 2026: Run set aside. The provider was overloaded on 37 of the 63 questions and we did not retry them.

20 to 23 August 2026 (4 days): No run. Our prepaid credit with the provider had run out, so all 63 questions came back as billing errors.

24 August 2026: No run recorded. Reason not yet written up.

What we do not claim: we never compare this model with another, and we never say a change was deliberate.