Independent capability measurement

Choosely AI Progress Index

How close AI is to doing complex digital work on its own.

A conservative, evidence-backed measure of what today’s best AI can actually do — not what companies say it can do.

As of 31 August 2026, the Index places frontier AI at 40.2 out of 100 across reasoning and adaptation, real-world work, and autonomy and agency.

Methodology V1.0First canonical snapshot
Current frontier/ 100

Verified capability—not marketing claims.

What is driving the score?

Three ways to understand where AI is now.

30% of index
50.1

Reasoning & Adaptation

Can AI figure out difficult problems and adapt when things change?

Strong on static novel puzzles; still weak when the rules must be discovered through interaction.
30% of index
22.5

Real-World Work

Can AI reliably finish useful work from start to finish?

The weakest pillar: accepted professional deliverables remain much harder than polished demos suggest.
40% of index
46.2

Autonomy & Agency

Can AI keep working on its own without constant human help?

AI can reach into longer expert tasks, but strict end-to-end computer-workflow completion stays low.

What has actually been counted?

The Live Edge

Choosely separates proven capability from promising claims and unanswered questions.

VerifiedProven and counted.
Novel reasoning

Static novel reasoning is the only scored measure above 90.

ARC-AGI-2 · Verified evidence
Longer autonomous work

AI can attempt work that would take human experts almost 12 hours, but uncertainty rises as tasks get longer.

METR · Verified evidence
Broad agent tasks

Fewer than one in three broad agent tasks in the frozen test set pass completely.

ALE · Verified evidence
Being checkedInteresting—not counted yet.
New benchmark results

Checked before they can change the Index

CHECKING
Company performance claims

Not counted without independent evidence

EXCLUDED
Still openImportant capabilities not reliably solved.
A reliable history of progress

This first snapshot is where the series begins

OPEN
World Ahead comparisons

Possible after a second approved edition

V1.1
View all evidence →

What is holding AI back?

What AI still can’t do.

Impressive demos are not the same as dependable autonomous work. These are the clearest gaps in the frozen evidence.

01

Adapt when the rules are hidden

AI performs poorly when it must discover the rules through interaction instead of seeing them upfront.

ARC-AGI-3 · 7.78%
02

Deliver client-acceptable work

Most professional projects still need human work before independent reviewers call the result client-ready.

RLI · 15.8%
03

Finish computer workflows cleanly

AI often makes useful progress on computer tasks but fails to complete every required step.

OSWorld 2.0 · 20.6% binary accuracy
04

Pass broad agent tasks in full

Broad agent tasks only count when the whole job is finished correctly; partial progress does not pass.

ALE · 30.6% full pass rate

How did capability reach this point?

How AI got here.

A few capability shifts changed what AI systems could plausibly do next. This is context—not fabricated Index history.

  1. GPT-4 launches

    Broad multimodal reasoning advances

  2. Function calling

    Models begin using external tools

  3. Claude 3 Opus

    Long context and robust analysis

  4. GPT-4o

    Real-time multimodal interaction

  5. Agentic systems

    Autonomous step-by-step workflows

  6. AI Progress Index V1.0

    First canonical measurement

  7. Reliable autonomy

    The work ahead

Why trust the measurement?

What the Index is based on.

Six independent tests. Exact results are frozen before they can influence the headline score.

Novel reasoningVerified

92.5%

AI can solve many difficult unfamiliar puzzles when the rules are visible in the task.

ARC-AGI-2
Adapting to hidden rulesVerified

7.78%

AI still struggles when the rules must be discovered through interaction.

ARC-AGI-3
Professional workVerified

15.8%

Most professional projects still do not reach a client-acceptable finish without human help.

RLI
Broad agent tasksVerified

30.6% full pass rate

AI completes fewer than one in three broad agent tasks in full on the frozen test set.

ALE
Longer autonomous workVerified

~11h 59m at 50% success

AI can now tackle tasks that take human experts hours, but reliability falls as work gets longer.

METR
Computer useVerified

20.6% binary accuracy

AI often makes partial progress across computer workflows but rarely finishes them cleanly.

OSWorld 2.0

Measured capability → plausible consequence

The World Ahead

An evolving glimpse of the future made plausible by today’s verified AI capabilities.

The Index tells you how far capability has moved. The World Ahead lets you feel why that might matter.
Current EditionHistory — Coming soon2026 · First canonical edition
Choosely World Ahead 2026: a plausible coastal city shaped by today’s verified AI capabilities
MeasuredAI Progress IndexPlausible consequenceThe World Ahead
Why this world?

Plate I reflects today’s capability imbalance: AI reasoning has advanced meaningfully, but dependable real-world execution still lags. The image is not a forecast or a score visualization. It is a plausible consequence layer grounded in today’s verified evidence.

The Change Brief

Follow the frontier as it evolves.

Weekly intelligence on the AI changes that actually matter.

By submitting, you’re asking to receive The Change Brief by email. Unsubscribe anytime.
Choosely Chimp