Independent reliability evidence

The Meter · Model page

THE METER · MODELGemini 3.1 Pro
Provider
Google
Exact model id we call
gemini-3.1-pro-preview
Route
provider direct
Watching since
28 July 2026

What we check

Most checks are the same on every model. See what we check

This provider does not send our settings back. We record what the response does show, and never say a setting was applied. Google publishes a description of this model. We check it daily for edits.

Pattern

Time to answer

Latency, time-to-first-token and reasoning-token volume are recorded on every call across all checks.

Daily exam

The scored series

The daily exam is the only check with a published score.

Other checks

Reasoning-token volume

The panel reports the median by published day across all checks.

The settings

Settings and dated runs

The settings we send, every time

endpoint:  v1beta streamGenerateContent (alt=sse)
model:     "gemini-3.1-pro-preview"
generationConfig:
  temperature:      1.0
  topP:             0.95
  topK:             64
  maxOutputTokens:  65536
  responseMimeType: "text/plain"
  thinkingConfig:   {"thinkingLevel": "MEDIUM"}
serviceTier: "standard"
store:       false

Dated runs

When we change our own settings, we close that run and start a new dated one. The earlier days stay published.

Dated runDatesRouteStatusDays
box-2026-07-28since 28 Julprovider directcurrent18 published

Gap register

29 July 2026: Run set aside. Our prepayment credits with the provider were depleted, so 47 of the 63 questions were refused.

30 July 2026: Run set aside. The provider returned a temporary error on one question and three retries did not clear it.

31 July 2026: Run set aside. Nothing failed. The run took more than twice as long as a normal day for this line, so it is not comparable to the rest of the series.

7 August 2026: Run set aside. The provider returned a temporary error on one question and three retries did not clear it.

9 August 2026: Run set aside. Ten of the 63 questions failed. Five hit a temporary provider error and five ran past our time limit.

10 August 2026: Run set aside. The provider returned a temporary error on 19 of the 63 questions and three retries did not clear it.

12 August 2026: No run. Two separate limits stopped it: 39 questions hit the provider's cap on how many requests we may send in a day, and the other 24 came back with our prepayment credits depleted.

13 August 2026: No run. Our prepayment credits with the provider were depleted, so all 63 questions were refused.

22 August 2026: Run set aside. Nothing failed. The run took more than twice as long as a normal day for this line, so it is not comparable to the rest of the series.

24 August 2026: Run set aside. Our prepayment credits with the provider were depleted, so 49 of the 63 questions were refused.

What we do not claim: we never compare this model with another, and we never say a change was deliberate.