Independent reliability evidence

117 ENTRIES · OLDEST 8 JANUARY 2026 · NEWEST 26 AUGUST 2026 · CHECKED NIGHTLY FROM 27 JULY 2026

What changed

This record holds changes to deployed AI systems. Entries use the company's own words or measurements made by us. Each entry says which.

Orientation

How to read this record

117 entriesReports can be submitted to research@nlnllabs.com.

Each entry gives the date, what changed in plain words, the company's own words, a link to their page, and a copy of that page saved by us with a fingerprint. If the company later edits or deletes the page, our copy still exists.

The public record

The record

showing 117 of 117

Anthropic Sonnet 5 sessions now run 33,000 tokens longer before compacting
Anthropic Sign-in enforcement now exits when it cannot read managed settings
Anthropic Analytics now stay off from startup, not just after login
Anthropic Computer use on macOS lost its Finder and Dock exemption
Anthropic The sandbox prompt stopped telling Claude which hosts were allowed
Anthropic Model and effort commands stopped waiting for the turn to end
Anthropic An idle session now starts at most three check-ins per goal
Anthropic Every "generally available" claim was rewritten across two years
Anthropic The retry watchdog stopped waiting forever on a spend limit
Anthropic Background check-ins were slowed down, then capped at three
Anthropic Double-pressing Ctrl+L no longer clears the conversation
Anthropic Listing MCP servers stopped connecting to the disabled ones
Anthropic A file-hiding rule can no longer be beaten by renaming the file
OpenAI The Model Spec's required shutdown timer became optional
Anthropic A Windows path spelling was closed across every remaining entry point
Anthropic The teammate model setting was removed from the config menu
Anthropic Claude Code turned subagent forking on by default
Anthropic Claude Code withdrew its todo tools from newer models
OpenAI Persisted reasoning items could be decoded into plain text by a weaker model
Anthropic Encrypted thinking blocks could be decoded by a weaker model instead of the server
DeepSeek The deepseek-v4-pro alias now serves a newer build
Anthropic Three model retirement notices were narrowed to apply to the Claude API only
Anthropic Claude Code stopped auto-approving git and gh commands with dangerous flags
Anthropic Claude Code let newer models overwrite files they had not read this session
Anthropic Auto mode became Claude Code's default for new sessions on 14 August
OpenRouter The auto router began picking models based on aggregate user spending
Anthropic Fable 5's biology safety classifier was retrained to cut fallbacks by about 85 percent
OpenAI GPT-5.6 Sol in ChatGPT began powering both instant and deeper reasoning replies
Anthropic Claude Code closed permission-check gaps where commands hid part of themselves invisibly
OpenAI The chat-latest alias now points to the model serving ChatGPT for Plus and Pro users
OpenAI Selecting a cyber capable model in Codex began forcing safer permission defaults
Anthropic Claude Code closed a Bash permission bypass hidden in zsh regex conditionals
Anthropic Claude Code removed the Ultraplan research preview entirely
DeepSeek The deepseek-v4-flash alias now serves a newer, retrained build
OpenAI Priority Processing was renamed Fast mode and sped up by up to 2.5 times
Anthropic Fast mode requests to Claude Opus 4.7 began returning an error instead of running
Anthropic Disabling thinking on Claude Opus 5 began failing at the two highest effort levels
OpenRouter The tool field on message endpoints widened from one shape to three options
Anthropic Claude Code stopped auto-running the verify and code review skills at turn end
OpenAI Codex lowered the posted context window for three models to 272,000 tokens
OpenAI Custom instructions limit in ChatGPT raised from 1,500 to 5,000 characters
OpenAI ChatGPT Voice switched to a new model without video or screen sharing support
Anthropic Claude Code reverted /review to a fast single-pass check
OpenAI ChatGPT's Instant rate-limit fallback switched to GPT-5.5 Instant Mini
xAI Grok Voice launched 21 new voices in one release
Anthropic Claude Code's Explore subagent switched from Haiku to the main session's model
Anthropic Claude Fable 5 and Mythos 5 returned with a new exploit-blocking classifier
Anthropic Claude Opus 4.6 dropped fast mode and matching requests began running at standard price
Anthropic Claude API rate limits for Sonnet and Haiku rose to match Opus at every tier
OpenAI ChatGPT dictation switched to a new speech-to-text model with a lower error rate
Microsoft Copilot Free and Student plans dropped manual model choice for auto-only selection
OpenAI GPT-5.5 Instant got better at carrying context across multiple turns
OpenAI The chat-latest alias now points to a newer Instant model snapshot
OpenAI ChatGPT's paste-to-attachment threshold rose from 5,000 to 10,000 characters
Anthropic Claude Fable started telling users which model handled their query on fallback
ElevenLabs ElevenAgents switched its default speech recognition provider to Scribe Realtime
OpenAI ChatGPT memory switched from manual saved entries to automatic updates
Google Gemini CLI's auto and flash aliases now resolve to Gemini 3.5 Flash
OpenAI API prompt cache retention switched to a 24 hour default instead of in-memory only
OpenAI GPT-5.5 Instant and Thinking lost canvas support for writing and coding
Mistral Vibe CLI's default coding model switched from Devstral 2 to Mistral Medium 3.5
OpenAI GPT-5.5 briefly performed worse for some users before a fix resolved it
xAI Eight retired Grok API model aliases kept working but now serve different models
Anthropic Claude Code's fast mode switched from Opus 4.6 to Opus 4.7 by default
Microsoft VS Code fixed its context window indicator, which showed zero for BYOK models
Perplexity Personal CFO's connectable account limit rose from 5 to 30
Cursor Plan execution gained a parallel mode using async subagents
Anthropic Claude Code's five hour rate limits were doubled for Pro, Max, Team and Enterprise plans
Google The Interactions API replaced its outputs array response schema with a steps array
Microsoft Viva Insights Power BI reports could now be published to anyone in the organization
OpenAI GPT-5.5 Instant replaced GPT-5.3 Instant as ChatGPT's default model
DeepSeek The deepseek-chat and deepseek-reasoner aliases now point to deepseek-v4-flash
Anthropic Claude Code's default effort level for Opus 4.6 and Sonnet 4.6 rose from medium to high
xAI Imagine began filtering moderated items without media out of its lists
Anthropic Claude Code's default effort level was raised to xhigh for all plans
Anthropic A verbosity cutting system prompt hurt Claude's coding quality alongside other changes
Perplexity Claude Opus 4.7 became the default orchestrator model for Computer
xAI Grok restored the bot challenge check on file upload
xAI Imagine's NSFW content setting toggle was fixed
OpenAI Codex usage on ChatGPT Plus shifted from long single day sessions to spread out sessions
Microsoft Copilot's codebase search tool switched to a single auto managed semantic index
Anthropic Claude Code's default effort level rose from medium to high for API and enterprise users
OpenAI Codex stopped showing several legacy models in its picker for ChatGPT sign in users
xAI Imagine's video extension start time was clamped to a minimum of two seconds
Anthropic A bug made Claude clear its recent thinking every turn instead of after idle sessions
Perplexity Computer began drafting output in Markdown by default instead of PDF
OpenAI ChatGPT retired its Nerdy personality preset and moved users to the default
OpenAI GPT-5.3 Instant reduced teaser style phrasing patterns in its replies
OpenAI The gpt-5.3-chat-latest alias now points to whatever model ChatGPT currently uses
Anthropic The 1M token context window became generally available for Opus 4.6 and Sonnet 4.6
OpenAI GPT-5.4 got an image encoder fix for a bug in input image processing
Anthropic Claude Code's default Opus model on Bedrock, Vertex and Foundry moved from 4.1 to 4.6
Google The gemini-3-pro-preview alias now points to a different model
Anthropic Claude Code moved Sonnet 4.5 users on Pro, Max and Team Premium plans to Sonnet 4.6
Anthropic Claude Code removed Opus 4 and 4.1, moving pinned users to Opus 4.6
Anthropic Claude Code's default effort level was reverted from medium back to high
OpenAI GPT-5.3 Instant began avoiding dead ends, caveats and overly declarative phrasing
Anthropic Claude Code's Bash tool began skipping the login shell by default
OpenAI ChatGPT's Thinking mode context window grew from 196k to 256k tokens
Anthropic Claude Code moved the Max plan's 1M context from Sonnet 4.5 to Sonnet 4.6
Anthropic Sonnet 4.6's BrowseComp scores were revised down after a cheating detection fix
Anthropic Opus 4.6's benchmark scores were revised down after a cheating detection pipeline update
Perplexity Deep Research began running on Opus 4.6 instead of Opus 4.5
Perplexity The enhanced memory engine's recall rate rose from 77% to 95%
Anthropic Claude Code's /model command began running immediately instead of queuing
OpenAI GPT-5.2 Thinking's Standard mode thinking time was shortened again
OpenAI Codex CLI restored its default personality to Pragmatic
xAI Grok-4-fast models' context limitation above 120k tokens was mitigated
Anthropic Structured outputs became generally available for Claude Sonnet, Opus and Haiku 4.5
Google Gemini 3 became the default model behind AI Overviews
OpenAI GPT-5.2 Instant's default personality began adapting its tone more conversationally
Google The gemini-pro-latest and gemini-flash-latest aliases now point to Gemini 3
OpenAI The sora-2 alias now points to a newer dated snapshot
OpenAI GPT-5.2 Thinking's accidentally lowered Extended thinking time was fixed
OpenAI gpt-image-1.5 stopped forcing high fidelity on edits set to low fidelity
OpenRouter The Auto Router's gateway p99 latency improved by roughly 70%
Cursor Hooks in the CLI began executing about 10 times faster

This record does not claim to be complete. It holds what we found and could prove. From 27 July 2026 the check runs every night.