{"schemaVersion":"choosely.capi.snapshot.v1","snapshot":{"id":"CAPI-V1-2026-08-31","date":"2026-08-31","status":"frozen","methodologyVersion":"1.0","verificationLabel":"Verified launch snapshot","sourceLockDate":"2026-08-31"},"index":{"name":"Choosely AI Progress Index","shortName":"CAPI","descriptor":"How close AI is to doing complex digital work on its own.","summary":"A conservative, evidence-backed measure of what today’s best AI can actually do — not what companies say it can do."},"pillars":[{"id":"reasoning","name":"Reasoning & Adaptation","question":"Can AI figure out difficult problems and adapt when things change?","weight":0.3,"summary":"Strong on static novel puzzles; still weak when the rules must be discovered through interaction.","metrics":[{"id":"arc2","name":"ARC-AGI-2","shortName":"ARC-AGI-2","publicLabel":"Novel reasoning","publicFinding":"AI can solve many difficult unfamiliar puzzles when the rules are visible in the task.","displayResult":"92.5%","metric":"Semi-Private pass@2 accuracy","rawValue":92.5,"normalizedScore":92.5,"pillarWeight":0.5,"headlineWeight":0.15,"system":"GPT-5.6 Sol · Max","systemBoundary":"model_direct","surface":"Semi-Private Evaluation Set · 120 tasks","evaluator":"ARC Prize Foundation","evidenceStatus":"VERIFIED","verifiedAt":"2026-08-30","sourceUrl":"https://arcprize.org/results/openai-gpt-5-6-sol","methodologyAnchor":"arc-agi-2","note":"The Verified score and the Public Eval diagnostic table are separate surfaces; the ARC2 provenance addendum closes this distinction."},{"id":"arc3","name":"ARC-AGI-3","shortName":"ARC-AGI-3","publicLabel":"Adapting to hidden rules","publicFinding":"AI still struggles when the rules must be discovered through interaction.","displayResult":"7.78%","metric":"Semi-Private interactive adaptation score","rawValue":7.78,"normalizedScore":7.78,"pillarWeight":0.5,"headlineWeight":0.15,"system":"GPT-5.6 Sol · Max","systemBoundary":"model_direct","surface":"Semi-Private · 55 environments","evaluator":"ARC Prize Foundation","evidenceStatus":"VERIFIED","verifiedAt":"2026-08-30","sourceUrl":"https://arcprize.org/results/openai-gpt-5-6-sol","methodologyAnchor":"arc-agi-3","note":"Benchmark-native interaction only; the public demo result is excluded."}]},{"id":"work","name":"Real-World Work","question":"Can AI reliably finish useful work from start to finish?","weight":0.3,"summary":"The weakest pillar: accepted professional deliverables remain much harder than polished demos suggest.","metrics":[{"id":"rli","name":"Remote Labor Index","shortName":"RLI","publicLabel":"Professional work","publicFinding":"Most professional projects still do not reach a client-acceptable finish without human help.","displayResult":"15.8%","metric":"Automation Rate","rawValue":15.8,"normalizedScore":15.8,"pillarWeight":0.55,"headlineWeight":0.165,"system":"Fable-5","systemBoundary":"agent_system","surface":"230 private professional projects","evaluator":"Scale Labs / RLI operator","evidenceStatus":"VERIFIED","verifiedAt":"2026-08-30","sourceUrl":"https://labs.scale.com/leaderboard/rli","methodologyAnchor":"rli","note":"A project counts only when independent reviewers judge the deliverable client-acceptable."},{"id":"ale","name":"Agents’ Last Exam","shortName":"ALE","publicLabel":"Broad agent tasks","publicFinding":"AI completes fewer than one in three broad agent tasks in full on the frozen test set.","displayResult":"30.6% full pass rate","metric":"Overall Full Pass Rate","rawValue":30.6,"normalizedScore":30.6,"pillarWeight":0.45,"headlineWeight":0.135,"system":"Codex + GPT-5.6 Sol · XHigh","systemBoundary":"agent_system","surface":"ALE-V1-FULL-2026-08-30 · 152 tasks","evaluator":"Agents’ Last Exam","evidenceStatus":"VERIFIED","verifiedAt":"2026-08-30","sourceUrl":"https://agents-last-exam.org/leaderboard","methodologyAnchor":"ale","note":"Full pass rate only; partial-credit score and best-per-task composites are excluded."}]},{"id":"autonomy","name":"Autonomy & Agency","question":"Can AI keep working on its own without constant human help?","weight":0.4,"summary":"AI can reach into longer expert tasks, but strict end-to-end computer-workflow completion stays low.","metrics":[{"id":"metr","name":"METR Time Horizon 1.1","shortName":"METR","publicLabel":"Longer autonomous work","publicFinding":"AI can now tackle tasks that take human experts hours, but reliability falls as work gets longer.","displayResult":"~11h 59m at 50% success","metric":"50% success horizon, normalized to a 40-hour target","rawValue":719,"rawUnit":"human-expert task minutes","normalizedScore":84.51,"pillarWeight":0.4,"headlineWeight":0.16,"system":"Claude Opus 4.6","systemBoundary":"agent_system","surface":"TH1.1 · approximately 11h 59m","evaluator":"METR","evidenceStatus":"VERIFIED","verifiedAt":"2026-08-30","sourceUrl":"https://metr.org/time-horizons/","methodologyAnchor":"metr","note":"This is task difficulty, not literal unattended runtime; uncertainty is substantial near the suite’s edge."},{"id":"osworld2","name":"OSWorld 2.0","shortName":"OSWorld 2.0","publicLabel":"Computer use","publicFinding":"AI often makes partial progress across computer workflows but rarely finishes them cleanly.","displayResult":"20.6% binary accuracy","metric":"Binary Accuracy","rawValue":20.6,"normalizedScore":20.6,"pillarWeight":0.6,"headlineWeight":0.24,"system":"Claude Opus 4.8 · Max · Batched tool","systemBoundary":"agent_system","surface":"v2026.06.24 · 108 tasks · 500-step budget","evaluator":"OSWorld 2.0 benchmark team","evidenceStatus":"VERIFIED","verifiedAt":"2026-08-30","sourceUrl":"https://osworld-v2.xlang.ai/","methodologyAnchor":"osworld-2","note":"Strict binary completion is scored; the 54.8% partial score is display-only."}]}],"liveEdge":[{"metricId":"arc2","label":"Near the V1 ceiling","detail":"Static novel reasoning is the only scored measure above 90.","tone":"high"},{"metricId":"metr","label":"Longer tasks, wider uncertainty","detail":"AI can attempt work that would take human experts almost 12 hours, but uncertainty rises as tasks get longer.","tone":"high"},{"metricId":"ale","label":"Broad work remains brittle","detail":"Fewer than one in three broad agent tasks in the frozen test set pass completely.","tone":"mid"}],"gaps":[{"metricId":"arc3","title":"Adapt when the rules are hidden","detail":"Interactive semi-private adaptation scores 7.78 — the lowest reading in the basket."},{"metricId":"rli","title":"Deliver client-acceptable work reliably","detail":"Only 15.8% of the frozen professional project set is automated to the acceptance threshold."},{"metricId":"osworld2","title":"Finish long computer workflows cleanly","detail":"Strict binary completion is 20.6%, despite materially higher partial progress."}],"evidenceQueues":{"underVerification":[{"title":"New benchmark results","detail":"Checked before they can change the Index","tag":"CHECKING"},{"title":"Company performance claims","detail":"Not counted without independent evidence","tag":"EXCLUDED"},{"title":"Updated test versions","detail":"Reviewed for a fair like-for-like comparison","tag":"REVIEW"}],"open":[{"title":"A reliable history of progress","detail":"This first snapshot is where the series begins","tag":"OPEN"},{"title":"World Ahead comparisons","detail":"Possible after a second approved edition","tag":"V1.1"},{"title":"Human impact signals","detail":"Not published until defensible evidence exists","tag":"OPEN"}]},"capabilityLimits":[{"metricId":"arc3","icon":"◇","title":"Adapt when the rules are hidden","detail":"AI performs poorly when it must discover the rules through interaction instead of seeing them upfront.","valueType":"metric","label":"ARC-AGI-3"},{"metricId":"rli","icon":"▦","title":"Deliver client-acceptable work","detail":"Most professional projects still need human work before independent reviewers call the result client-ready.","valueType":"metric","label":"RLI automation"},{"metricId":"osworld2","icon":"⌘","title":"Finish computer workflows cleanly","detail":"AI often makes useful progress on computer tasks but fails to complete every required step.","valueType":"metric","label":"OSWorld binary"},{"metricId":"ale","icon":"✓","title":"Pass broad agent tasks in full","detail":"Broad agent tasks only count when the whole job is finished correctly; partial progress does not pass.","valueType":"metric","label":"ALE full pass"},{"metricId":"osworld2","icon":"↯","title":"Turn progress into completion","detail":"OSWorld partial progress is 54.8%, but strict completion is only 20.6%.","valueType":"osworld_completion_gap","label":"completion gap"},{"metricId":"arc3","icon":"∞","title":"Generalize across novel environments","detail":"The 84.72-point ARC2-to-ARC3 spread exposes a major adaptation gap.","valueType":"reasoning_balance","label":"remaining balance"},{"metricId":"metr","icon":"◷","title":"Measure longer tasks reliably","detail":"The 11h 59m horizon sits near the current suite’s measurement-quality edge.","valueType":"caution","label":"suite-edge caution"},{"metricId":"rli","icon":"◎","title":"Reach the real-work milestone","detail":"RLI remains 34.2 points short of Choosely’s 50% project-automation milestone.","valueType":"rli_milestone","label":"of milestone"}],"capabilityHistory":[{"date":"Mar 2023","title":"GPT-4 launches","detail":"Broad multimodal reasoning advances","state":"reached","sourceUrl":"https://openai.com/index/gpt-4-research/"},{"date":"Jun 2023","title":"Function calling","detail":"Models begin using external tools","state":"reached","sourceUrl":"https://openai.com/index/function-calling-and-other-api-updates/"},{"date":"Mar 2024","title":"Claude 3 Opus","detail":"Long context and robust analysis","state":"reached","sourceUrl":"https://www.anthropic.com/news/claude-3-family"},{"date":"May 2024","title":"GPT-4o","detail":"Real-time multimodal interaction","state":"reached-warm","sourceUrl":"https://openai.com/index/hello-gpt-4o/"},{"date":"2025–26","title":"Agentic systems","detail":"Autonomous step-by-step workflows","state":"reached-warm","sourceUrl":"/ai-progress/methodology#architecture"},{"date":"Aug 2026","title":"AI Progress Index V1.0","detail":"First canonical measurement","state":"reached-warm","sourceUrl":"/ai-progress/methodology#calculation"},{"date":"Next","title":"Reliable autonomy","detail":"The work ahead","state":"open","sourceUrl":"#open-gaps"}],"milestones":[{"date":"2026-03","title":"Long-horizon measurement enters the edge of the suite","detail":"METR’s usable public reference reaches roughly 11h 59m at 50% success, with explicit uncertainty warnings.","sourceUrl":"https://metr.org/notes/2026-03-20-impact-of-modelling-assumptions-on-time-horizon-results/"},{"date":"2026-06","title":"Desktop work gets a stricter binary surface","detail":"OSWorld 2.0 v2026.06.24 tests 108 workflows with a 500-step budget.","sourceUrl":"https://osworld-v2.xlang.ai/"},{"date":"2026-08","title":"A frozen evidence basket becomes reproducible","detail":"Six exact benchmark surfaces, 10,000 basis points and a 24% concentration ceiling produce the first CAPI snapshot.","sourceUrl":"/ai-progress/methodology"},{"date":"2026-09","title":"The World Ahead becomes a consequence layer","detail":"The approved 2026 edition visualises a plausible world shaped by verified capability — it does not forecast it.","sourceUrl":"#world-ahead"}],"supportingMilestones":[{"title":"Professional deliverable benchmark","status":"REACHED","detail":"The human-anchored GDPval-AA v2 Elo level has been exceeded by frontier systems. This is benchmark-specific, not universal professional-work parity.","sourceUrl":"https://artificialanalysis.ai/evaluations/gdpval-aa"}],"worldAhead":{"editionId":"WA-L00-2026","title":"Choosely World Ahead 2026","status":"frozen","approvalStatus":"APPROVED","frozenAt":"2026-09-04","desktopSha256":"0e44a42e981caaee1cbe4c2d03ab5d1d921f373e804e92b54256cc77da7ada77","mobileSha256":"e66f4d69899564f7069fc09d51ff380b22b3659224d1fd194c5cecfc4fbcc423","positioning":"An evolving glimpse of the future made plausible by today’s verified AI capabilities.","bridge":"The Index tells you how far capability has moved. The World Ahead lets you feel why that might matter.","humanSignals":[],"changes":[]},"calculated":{"exactScore":40.2456,"displayScore":40.2,"pillars":[{"id":"reasoning","name":"Reasoning & Adaptation","exactScore":50.14,"displayScore":50.1,"weight":0.3},{"id":"work","name":"Real-World Work","exactScore":22.46,"displayScore":22.5,"weight":0.3},{"id":"autonomy","name":"Autonomy & Agency","exactScore":46.164,"displayScore":46.2,"weight":0.4}]},"distributions":{"json":"https://choosely.ai/ai-progress/data.json","csv":"https://choosely.ai/ai-progress/data.csv","methodology":"https://choosely.ai/ai-progress/methodology"}}