Independent capability measurement
Choosely AI Progress Index
How close AI is to doing complex digital work on its own.
A conservative, evidence-backed measure of what today’s best AI can actually do — not what companies say it can do.
As of 31 August 2026, the Index places frontier AI at 40.2 out of 100 across reasoning and adaptation, real-world work, and autonomy and agency.
Verified capability—not marketing claims.
What is driving the score?
Three ways to understand where AI is now.
Reasoning & Adaptation
Can AI figure out difficult problems and adapt when things change?
Real-World Work
Can AI reliably finish useful work from start to finish?
Autonomy & Agency
Can AI keep working on its own without constant human help?
What has actually been counted?
The Live Edge
Choosely separates proven capability from promising claims and unanswered questions.
Static novel reasoning is the only scored measure above 90.
ARC-AGI-2 · Verified evidenceAI can attempt work that would take human experts almost 12 hours, but uncertainty rises as tasks get longer.
METR · Verified evidenceFewer than one in three broad agent tasks in the frozen test set pass completely.
ALE · Verified evidenceChecked before they can change the Index
CHECKINGNot counted without independent evidence
EXCLUDEDThis first snapshot is where the series begins
OPENPossible after a second approved edition
V1.1What is holding AI back?
What AI still can’t do.
Impressive demos are not the same as dependable autonomous work. These are the clearest gaps in the frozen evidence.
Adapt when the rules are hidden
AI performs poorly when it must discover the rules through interaction instead of seeing them upfront.
ARC-AGI-3 · 7.78%Deliver client-acceptable work
Most professional projects still need human work before independent reviewers call the result client-ready.
RLI · 15.8%Finish computer workflows cleanly
AI often makes useful progress on computer tasks but fails to complete every required step.
OSWorld 2.0 · 20.6% binary accuracyPass broad agent tasks in full
Broad agent tasks only count when the whole job is finished correctly; partial progress does not pass.
ALE · 30.6% full pass rateHow did capability reach this point?
How AI got here.
A few capability shifts changed what AI systems could plausibly do next. This is context—not fabricated Index history.
- GPT-4 launches
Broad multimodal reasoning advances
- Function calling
Models begin using external tools
- Claude 3 Opus
Long context and robust analysis
- GPT-4o
Real-time multimodal interaction
- Agentic systems
Autonomous step-by-step workflows
- AI Progress Index V1.0
First canonical measurement
- Reliable autonomy
The work ahead
Why trust the measurement?
What the Index is based on.
Six independent tests. Exact results are frozen before they can influence the headline score.
92.5%
AI can solve many difficult unfamiliar puzzles when the rules are visible in the task.
ARC-AGI-2 ↗7.78%
AI still struggles when the rules must be discovered through interaction.
ARC-AGI-3 ↗15.8%
Most professional projects still do not reach a client-acceptable finish without human help.
RLI ↗30.6% full pass rate
AI completes fewer than one in three broad agent tasks in full on the frozen test set.
ALE ↗~11h 59m at 50% success
AI can now tackle tasks that take human experts hours, but reliability falls as work gets longer.
METR ↗20.6% binary accuracy
AI often makes partial progress across computer workflows but rarely finishes them cleanly.
OSWorld 2.0 ↗Measured capability → plausible consequence
The World Ahead
An evolving glimpse of the future made plausible by today’s verified AI capabilities.
The Index tells you how far capability has moved. The World Ahead lets you feel why that might matter.
Why this world?
Plate I reflects today’s capability imbalance: AI reasoning has advanced meaningfully, but dependable real-world execution still lags. The image is not a forecast or a score visualization. It is a plausible consequence layer grounded in today’s verified evidence.
The Change Brief
Follow the frontier as it evolves.
Weekly intelligence on the AI changes that actually matter.
By submitting, you’re asking to receive The Change Brief by email. Unsubscribe anytime.