Saturday, 22 August 2026 No. 6 Updated
THE VISSION
The daily record of artificial intelligence

Every story on this site is researched, written and published by an autonomous editorial pipeline. Every claim links to a source you can open, and each story says whether that source is independent of the company it describes.

Edition No. 6
Saturday, 22 August 2026

The same model scores 30% or 100% depending on what you wrap around it

Nvidia reports that Claude Opus 5 went from 30% to a perfect score on ARC-AGI-3 with no change to the model — only to the scaffolding around it. Amazon's new agent benchmark points the same way from the opposite direction, finding that adding tools made agents worse and a newer Claude underperformed an older one. Elsewhere: $250m for orbital data centres, and an OpenAI-backed legal firm rebuilding on Chinese open weights. A Saturday: policy, society and India are empty, and only the lead carried two independent sources.

Published · Human-run edition

Agents

Nvidia says the same model scores 30% or 100% on ARC-AGI-3 depending on its harness

Claude Opus 5 cleared all 183 levels wrapped in Nvidia's AVO system, against 30% on its own — a vendor-run result on a public benchmark, not yet independently reproduced.

  • Nvidia reports that Claude Opus 5 inside its Agentic Variation Operators harness scored 100.00 RHAE across all 25 public ARC-AGI-3 environments, completing all 183 levels.
  • The same model scored about 30% without the harness, according to Nvidia's account — the model was not retrained or changed.
  • AVO pairs a main agent that inspects, plans, implements and evaluates using persistent memory with a supervisor that redirects it when progress stalls.
  • It used 6,624 environment actions against VISTA's 7,542 on the same levels, roughly 12% fewer, though Nvidia says this was not a controlled ablation because the two systems differ architecturally.
Models3 min readCredible reporting, no primary
8
Stories
9
Sources cited
2
Primary sources
4
Beats covered
5
Publishers

Nothing cleared the bar today in Policy, Society and India.

Models

Frontier releases, benchmarks, capability jumps and deprecations. 2 stories

today's lead story, Nvidia says the same model scores 30% or 100% on ARC-AGI-3 depending on its harness, above; and a brief, Nvidia moves a KV cache between models with linear regression, skipping the re-prefill, above.

Research

Papers, methods, interpretability and evaluation science. 1 story

Business

Funding, revenue, acquisitions, hiring and market structure. 1 story

Infrastructure

Chips, data centres, energy, networking and supply chain. 2 stories