AI

AGI Progress in 2026: Benchmarks, Agents and Safety Thresholds

The strongest AGI story of 2026 is not that somebody crossed a finish line. It is that researchers became better at exposing where frontier AI still breaks. Interactive environments, longer agent tasks and explicit safety thresholds moved evaluation beyond static question-answer benchmarks without producing a universally accepted declaration of AGI.

That makes 2026 important without requiring an “AGI achieved” headline. The field is getting better at asking what generality would look like operationally.

This dated 2026 record sits alongside our 2027 AGI guide. Its job is to preserve what materially changed in 2026 rather than rewrite the year after later breakthroughs occur.

ARC-AGI-3 moved the benchmark into interactive environments

ARC Prize launched ARC-AGI-3 in March 2026. Instead of presenting a static puzzle with a known input and output format, the benchmark places AI agents in game-like environments with no instructions or stated goals.

The agent has to explore, infer the rules and discover what success means. ARC Prize said humans score 100% while frontier AI scored only 0.51% at launch.

The exact number will change as competition entries improve, but the design shift is the important story. AGI research increasingly cares about adaptation and exploration, not only answer accuracy.

GPT-6 Astra transformed the ARC-AGI-3 picture in September

On 3 September, ARC Prize reported that GPT-6 Astra scored 62.7% on ARC-AGI-3 Semi-Private with its standard harness and 99.9% with a provider-adapter harness that preserves OpenAI’s reasoning state and uses context compaction.

That is a step-function change from the 0.51% launch baseline, but it does not collapse the AGI debate into one leaderboard. ARC Prize explicitly says it is not claiming Astra is AGI and notes that ARC-AGI-3 has bounded, deterministic environments. The result instead turns harness design, memory and state management into part of the measurement story.

METR kept extending the idea of long-horizon evaluation

METR’s 2026 time-horizon work measures the difficulty of software tasks that frontier agents can complete at specified success rates. Its task suite includes software engineering, machine learning and cybersecurity work.

The organisation warns that these numbers should not be read as literal autonomous runtime. They are estimates of task difficulty based on how long human experts take.

Even with that limitation, the research gives a useful view of whether AI systems are becoming reliable on more extended pieces of work.

MirrorCode suggested some very long coding tasks are already possible

METR’s 2026 research programme also reported early evidence that AI can complete some coding tasks that would take humans weeks, including reimplementing a large codebase under benchmark conditions.

This does not mean agents can replace a software team for weeks of real organisational work. Benchmarks remove many sources of ambiguity and institutional context. It does show that simple task-duration assumptions are becoming outdated.

DeepMind added a cognitive AGI framework in March

Google DeepMind’s 2026 cognitive framework proposes measuring learning, metacognition, attention, executive functions, social cognition and other faculties against human baselines. The move broadens AGI evaluation beyond model benchmarks toward a more explicit account of the cognitive abilities a general system would need.

Google DeepMind strengthened its Frontier Safety Framework

DeepMind updated its Frontier Safety Framework in 2026 with tracked capability levels designed to identify emerging risks earlier. The framework covers domains including cyber, biological risk, harmful manipulation, machine-learning R&D and misalignment.

The relevance to AGI is indirect but important. A lab may need to respond to dangerous capability before the system satisfies anyone’s preferred definition of AGI.

Anthropic continued revising its Responsible Scaling Policy

Anthropic’s Responsible Scaling Policy reached version 3.4 in 2026. The framework connects capability thresholds to safeguards and public risk reporting.

Its updates around automated AI R&D are especially relevant to AGI discussions because a system that substantially accelerates AI research could change the pace of capability growth itself.

Safety measurement is becoming part of capability measurement

Traditional benchmarks ask whether a system can do something useful. Frontier-safety evaluations also ask whether a system can do something dangerous.

That may seem like a separate topic, but the two converge as systems become more general. Broad competence means the same system can potentially support both beneficial and harmful objectives.

OpenAI Astra crossed a critical cybersecurity threshold

On 3 September 2026, OpenAI released GPT-6 Astra and said the model met the Critical cybersecurity capability threshold in its Preparedness Framework. OpenAI reported that, with the right tools and access, the model could identify previously unknown vulnerabilities and develop exploit chains against hardened systems without a person guiding every step.

This is important AGI evidence precisely because it is not proof of AGI. Astra demonstrates a major jump in one consequential capability domain, while the broader AGI question still depends on generalisation across many domains. TechnologyBlog’s Astra cybersecurity report covers the threshold in detail.

The AGI definition remains unresolved

OpenAI still frames AGI around highly autonomous systems outperforming humans at most economically valuable work. DeepMind’s academic framework separates levels of performance and generality. IBM describes AGI as hypothetical human-level or greater cognition across tasks.

Because those definitions differ, 2026 progress cannot be reduced to a single percentage.

Agentic AI is the bridge concept to watch

The most important practical shift is that frontier models increasingly operate as agents: they plan, use tools, browse, write code and iterate toward a goal.

Agency does not equal AGI, but it exposes limitations that static chat can hide. Long tasks reveal whether a system maintains state, notices failure and recovers.

What did not happen in 2026?

There was no universal scientific agreement that a deployed system had achieved AGI. There was no universally accepted benchmark that could settle the question. There was no reason for consumers or businesses to treat “AGI” as a normal product category.

That negative result matters because marketing language can run ahead of consensus.

South Africa’s policy process showed why evidence quality matters

South Africa withdrew its draft national AI policy after problems including fictitious and unverifiable references were identified. Cabinet later confirmed the withdrawal for rework and said the policy should establish national standards for ethical AI use.

The episode is a reminder that AI governance needs stronger sourcing, not weaker sourcing. Future AGI policy will be even more sensitive to evidence quality because the stakes will be higher.

Which 2026 developments actually changed the AGI debate?

The developments worth preserving are those that changed how generality or frontier risk is measured: a major benchmark, a material jump in long-horizon work or a revised capability threshold from a leading lab.

Minor model launches may matter greatly to users without changing the AGI evidence. A faster or cheaper product is not automatically a milestone in general intelligence.

The 2026 takeaway

The story of AGI in 2026 is not arrival. It is measurement under rapidly changing conditions. Interactive benchmarks exposed severe weaknesses in March, Astra sharply narrowed one of those gaps in September, long-horizon agent evaluations kept expanding and safety frameworks became more explicit.

That gives readers better tools for judging claims in 2027. Instead of asking whether a model “feels like AGI”, we can ask what unfamiliar tasks it can learn, how long it can work reliably, what safety thresholds it approaches and how much special scaffolding is required.

Why model releases alone are poor AGI milestones

Model names change quickly and vendors optimise different products for different uses. A newer release can be faster, cheaper or better at coding without changing the AGI picture fundamentally.

A model launch is an AGI milestone only when it supplies new evidence about generalisation, long-horizon autonomy, learning or another capability central to the definition.

Independent evaluation is becoming more important

As commercial incentives around AGI grow, independent benchmarks and reproducible methods matter more. Vendor evaluations are useful, but they are strongest when external researchers can inspect the methodology or reproduce the result.

ARC Prize and METR are valuable partly because they create evaluation programmes outside the product-marketing cycle. No benchmark is perfect, but plural independent measurement reduces dependence on a single company’s definition.