Methodology V1.0 · Frozen 31 August 2026
AI Progress Index methodology.
Every score, source and rule.
A public, reproducible account of what CAPI measures, what it excludes and how six verified observations become one conservative score.
01 · Architecture
What the Index measures.
CAPI measures demonstrated frontier capability on difficult, externally evaluated digital work. It is not a forecast, a model popularity ranking, a product leaderboard or a claim about general intelligence.
Reasoning & Adaptation
Strong on static novel puzzles; still weak when the rules must be discovered through interaction.
50.1 / 100Real-World Work
The weakest pillar: accepted professional deliverables remain much harder than polished demos suggest.
22.5 / 100Autonomy & Agency
AI can reach into longer expert tasks, but strict end-to-end computer-workflow completion stays low.
46.2 / 10002 · Deterministic calculation
From six scores to 40.2.
0.50(92.50) + 0.50(7.78) = 50.140.55(15.80) + 0.45(30.60) = 22.460.40(84.51) + 0.60(20.60) = 46.1640.30(50.14) + 0.30(22.46) + 0.40(46.164) = 40.2456 → 40.2The source file is parsed into a typed schema, then the same calculator supplies the webpage, JSON, CSV and snapshot share card. Display rounding happens only after the exact composite is calculated.
03 · Evidence policy
Claims need a surface, not a slogan.
Headline eligible
Exact benchmark surface, metric, system boundary, evaluator and primary source are resolved.
Visible, not scored
A relevant result exists, but at least one required identity or provenance field remains unresolved.
Research queue
A question or claim worth tracking without enough evidence for a public capability statement.
Vendor claims do not enter CAPI unless they match the frozen evaluation surface. Living benchmarks require an exact snapshot and a series-break review when their task corpus, grader or scoring definition changes.
04 · Frozen metric register
Every input, in the open.
ARC-AGI-2
- Scored metric
- Semi-Private pass@2 accuracy
- Frontier system
- GPT-5.6 Sol · Max
- Frozen surface
- Semi-Private Evaluation Set · 120 tasks
- System boundary
- model_direct
- Evaluator
- ARC Prize Foundation
- Headline weight
- 15%
The Verified score and the Public Eval diagnostic table are separate surfaces; the ARC2 provenance addendum closes this distinction.
Open primary source ↗ARC-AGI-3
- Scored metric
- Semi-Private interactive adaptation score
- Frontier system
- GPT-5.6 Sol · Max
- Frozen surface
- Semi-Private · 55 environments
- System boundary
- model_direct
- Evaluator
- ARC Prize Foundation
- Headline weight
- 15%
Benchmark-native interaction only; the public demo result is excluded.
Open primary source ↗Remote Labor Index
- Scored metric
- Automation Rate
- Frontier system
- Fable-5
- Frozen surface
- 230 private professional projects
- System boundary
- agent_system
- Evaluator
- Scale Labs / RLI operator
- Headline weight
- 16.5%
A project counts only when independent reviewers judge the deliverable client-acceptable.
Open primary source ↗Agents’ Last Exam
- Scored metric
- Overall Full Pass Rate
- Frontier system
- Codex + GPT-5.6 Sol · XHigh
- Frozen surface
- ALE-V1-FULL-2026-08-30 · 152 tasks
- System boundary
- agent_system
- Evaluator
- Agents’ Last Exam
- Headline weight
- 13.5%
Full pass rate only; partial-credit score and best-per-task composites are excluded.
Open primary source ↗METR Time Horizon 1.1
- Scored metric
- 50% success horizon, normalized to a 40-hour target
- Frontier system
- Claude Opus 4.6
- Frozen surface
- TH1.1 · approximately 11h 59m
- System boundary
- agent_system
- Evaluator
- METR
- Headline weight
- 16%
This is task difficulty, not literal unattended runtime; uncertainty is substantial near the suite’s edge.
Open primary source ↗OSWorld 2.0
- Scored metric
- Binary Accuracy
- Frontier system
- Claude Opus 4.8 · Max · Batched tool
- Frozen surface
- v2026.06.24 · 108 tasks · 500-step budget
- System boundary
- agent_system
- Evaluator
- OSWorld 2.0 benchmark team
- Headline weight
- 24%
Strict binary completion is scored; the 54.8% partial score is display-only.
Open primary source ↗05 · Governance
A frozen baseline, with explicit change rules.
Freeze record
Source lock, system boundaries, Grade-A evidence, the ALE manifest, ARC2 provenance, independent calculation, basis-point totals and concentration guard all passed.
ARC2 provenance
The Verified 92.5 score is Semi-Private. Public task details shown on the same result page are diagnostic and are not used to infer the scored surface.
Revision rule
Corrections are recorded; benchmark version changes are adjudicated; incompatible surfaces create a series break instead of a silent splice.
No synthetic history
This is the first canonical snapshot. CAPI history begins here; earlier capability events appear only as sourced milestones.
Stable source register
Snapshot: CAPI-V1-2026-08-31 · source lock: 2026-08-31 · methodology: V1.0