Methodology V1.0 · Frozen 31 August 2026

AI Progress Index methodology.
Every score, source and rule.

A public, reproducible account of what CAPI measures, what it excludes and how six verified observations become one conservative score.

01 · Architecture

What the Index measures.

CAPI measures demonstrated frontier capability on difficult, externally evaluated digital work. It is not a forecast, a model popularity ranking, a product leaderboard or a claim about general intelligence.

30%

Reasoning & Adaptation

Strong on static novel puzzles; still weak when the rules must be discovered through interaction.

50.1 / 100
30%

Real-World Work

The weakest pillar: accepted professional deliverables remain much harder than polished demos suggest.

22.5 / 100
40%

Autonomy & Agency

AI can reach into longer expert tasks, but strict end-to-end computer-workflow completion stays low.

46.2 / 100

02 · Deterministic calculation

From six scores to 40.2.

Reasoning & Adaptation0.50(92.50) + 0.50(7.78) = 50.14
Real-World Work0.55(15.80) + 0.45(30.60) = 22.46
Autonomy & Agency0.40(84.51) + 0.60(20.60) = 46.164
CAPI0.30(50.14) + 0.30(22.46) + 0.40(46.164) = 40.245640.2

The source file is parsed into a typed schema, then the same calculator supplies the webpage, JSON, CSV and snapshot share card. Display rounding happens only after the exact composite is calculated.

03 · Evidence policy

Claims need a surface, not a slogan.

Verified

Headline eligible

Exact benchmark surface, metric, system boundary, evaluator and primary source are resolved.

Under verification

Visible, not scored

A relevant result exists, but at least one required identity or provenance field remains unresolved.

Open

Research queue

A question or claim worth tracking without enough evidence for a public capability statement.

Vendor claims do not enter CAPI unless they match the frozen evaluation surface. Living benchmarks require an exact snapshot and a series-break review when their task corpus, grader or scoring definition changes.

04 · Frozen metric register

Every input, in the open.

Reasoning & Adaptation

ARC-AGI-2

92.5
Scored metric
Semi-Private pass@2 accuracy
Frontier system
GPT-5.6 Sol · Max
Frozen surface
Semi-Private Evaluation Set · 120 tasks
System boundary
model_direct
Evaluator
ARC Prize Foundation
Headline weight
15%

The Verified score and the Public Eval diagnostic table are separate surfaces; the ARC2 provenance addendum closes this distinction.

Open primary source ↗
Reasoning & Adaptation

ARC-AGI-3

7.78
Scored metric
Semi-Private interactive adaptation score
Frontier system
GPT-5.6 Sol · Max
Frozen surface
Semi-Private · 55 environments
System boundary
model_direct
Evaluator
ARC Prize Foundation
Headline weight
15%

Benchmark-native interaction only; the public demo result is excluded.

Open primary source ↗
Real-World Work

Remote Labor Index

15.8
Scored metric
Automation Rate
Frontier system
Fable-5
Frozen surface
230 private professional projects
System boundary
agent_system
Evaluator
Scale Labs / RLI operator
Headline weight
16.5%

A project counts only when independent reviewers judge the deliverable client-acceptable.

Open primary source ↗
Real-World Work

Agents’ Last Exam

30.6
Scored metric
Overall Full Pass Rate
Frontier system
Codex + GPT-5.6 Sol · XHigh
Frozen surface
ALE-V1-FULL-2026-08-30 · 152 tasks
System boundary
agent_system
Evaluator
Agents’ Last Exam
Headline weight
13.5%

Full pass rate only; partial-credit score and best-per-task composites are excluded.

Open primary source ↗
Autonomy & Agency

METR Time Horizon 1.1

84.51
Scored metric
50% success horizon, normalized to a 40-hour target
Frontier system
Claude Opus 4.6
Frozen surface
TH1.1 · approximately 11h 59m
System boundary
agent_system
Evaluator
METR
Headline weight
16%

This is task difficulty, not literal unattended runtime; uncertainty is substantial near the suite’s edge.

Open primary source ↗
Autonomy & Agency

OSWorld 2.0

20.6
Scored metric
Binary Accuracy
Frontier system
Claude Opus 4.8 · Max · Batched tool
Frozen surface
v2026.06.24 · 108 tasks · 500-step budget
System boundary
agent_system
Evaluator
OSWorld 2.0 benchmark team
Headline weight
24%

Strict binary completion is scored; the 54.8% partial score is display-only.

Open primary source ↗

05 · Governance

A frozen baseline, with explicit change rules.

Freeze record

Source lock, system boundaries, Grade-A evidence, the ALE manifest, ARC2 provenance, independent calculation, basis-point totals and concentration guard all passed.

ARC2 provenance

The Verified 92.5 score is Semi-Private. Public task details shown on the same result page are diagnostic and are not used to infer the scored surface.

Revision rule

Corrections are recorded; benchmark version changes are adjudicated; incompatible surfaces create a series break instead of a silent splice.

No synthetic history

This is the first canonical snapshot. CAPI history begins here; earlier capability events appear only as sourced milestones.

Stable source register

Snapshot: CAPI-V1-2026-08-31 · source lock: 2026-08-31 · methodology: V1.0