Skip to content
← Back to Blog
September 16, 2026

Has AGI Been Achieved? How to Evaluate the Claims and What to Build Now

Ihor Arkhypenko

Digital human face dissolving into purple and blue cubes, illustrating artificial general intelligence and AGI capabilities.

Don't wait for the AGI debate to settle. Whether it's "been achieved" depends entirely on which definition you pick, and no benchmark score proves general intelligence anyway. Build with today's models, measure the full workflow, keep the option to swap in stronger ones.

Answers regarding Artificial General Intelligence (AGI) and whether it already exists vary, as definitions measure different things: cognitive breadth, unfamiliar-task learning, economic work or autonomous operation. There is no definition-independent answer. Recent frontier-model evaluations show broad and sometimes superhuman performance in selected domains, but a benchmark score does not establish dependable general intelligence. Product teams should examine performance, generality, learning efficiency, autonomy and reliability separately.

That is not a reason to wait. Companies do not need agreement on AGI before acting. They need products that create value with current models, and architectures capable of absorbing whatever comes next.

On This Page

  • What Is Artificial General Intelligence?
  • Has AGI Been Achieved In 2026?
  • How Close Are We To AGI?
  • How Is AGI Measured?
  • What Should Businesses Build With AI Right Now?
  • How To Build For Model Change
  • AGI-Readiness Checklist

How This Article Evaluates AGI Claims

We separate benchmark scores from capability claims, tag each source by organization and date, and treat product examples as workflow illustrations, not proof of AGI. Use the same filter when you read any AGI announcement.

It then asks four things: What is the evidence? Under what test conditions? What happens when the system is wrong? What sits between the model’s output and a real action?

What Is Artificial General Intelligence?

Artificial general intelligence, or AGI, describes a machine with broad capabilities that generalize across tasks rather than remaining confined to one narrow function. Disagreement starts on three points: what counts as “broad,” which humans you compare against, and whether autonomy is part of the definition at all.

A common introductory definition describes AGI as a hypothetical system able to understand or learn any intellectual task a human can. Google Cloud’s AGI overview uses this framing and emphasizes generalization, reasoning and adaptation, while leaving reliability, embodiment and independence open to interpretation.

In Levels of AGI for Operationalizing Progress on the Path to AGI, a Google DeepMind research team proposed a more operational framework. It separates generality, the breadth of tasks a system can perform, from performance, how well it performs them, and treats autonomy as a separate deployment characteristic.

The ARC-AGI framework takes a different view. Its definition centers on skill-acquisition efficiency: whether a system can learn and generalize to unfamiliar problems without its developers anticipating each one. In this view, memorizing more facts or acing familiar tasks doesn’t make a system smarter. It only looks smarter until it meets a problem it hasn’t seen.

OpenAI’s Charter defines AGI in economic terms: highly autonomous systems that outperform humans at most economically valuable work. Whether a system can replace paid work depends on cost, tool access, integration, regulation and acceptable error—not cognition alone.

Why The Answer Depends On The Definition

The definitions are not interchangeable:

Definition LensCentral QuestionWhat It CapturesWhat It May Miss
Human-task breadth
Can it perform most intellectual tasks people can?Breadth and adaptabilityReliability, cost and deployment constraints
Performance × generality
How capable is it across how many domains?Uneven capability profiles and progress levelsLearning efficiency on genuinely novel tasks
Skill-acquisition efficiency
How efficiently can it learn an unfamiliar task?Fluid intelligence and generalizationEconomic usefulness across full workflows
Economic work
Can it outperform people at most valuable work?Commercial consequenceSocial, physical and non-market abilities
Operational autonomy
How long can it complete work without intervention?Real-world execution and reliabilityThe full range of cognitive ability

The label “AGI” compresses these five questions into one word. Separating them makes the disagreement easier to evaluate.

Has AGI Been Achieved In 2026?

No broadly accepted scientific test has established that AGI has been achieved. Frontier systems meet parts of several definitions. That shows rapid progress, not agreement that general intelligence has arrived.

What The Evidence Supports

Broad competence: Yes, current model evaluations show performance across writing, coding, mathematics, research, visual interpretation and computer use.

Superhuman performance on selected benchmarks or narrow task classes: Yes. A system can exceed human baselines on particular evaluations while remaining brittle elsewhere.

Human-like learning across unfamiliar environments: Contested. Generalization tests continue to evolve as models saturate earlier benchmarks; the ARC-AGI framework focuses specifically on performance on unfamiliar tasks.

Reliable autonomous work across open-ended situations: Not established as a general capability. Success changes with task duration, environment, tools and error tolerance.

Replacement of humans across most economically valuable work: Not demonstrated across the economy.

As of September 2026, published ARC Prize results and analysis report GPT-6 Astra at 62.7% under the Standard harness and 99.9% under a Provider Adapter harness. OpenAI’s announcement reports 99.9% under its configured evaluation setup. The difference shows why benchmark conditions must be stated explicitly. These published results are not proof that AGI has been achieved.

Benchmark Caveat: Where benchmark results are reported by model developers or benchmark organizations, this article treats them as published claims under stated test conditions rather than independent certification of general intelligence.

ARC-AGI itself represents a particular theory of intelligence. A near-ceiling score can show that a system has become highly effective under that benchmark’s specific task design and evaluation conditions. A high score doesn’t prove the system holds up in environments it never saw, in tasks that run for hours, or in processes where a mistake costs money.

The evidence supports substantial progress across several dimensions, not a settled verdict that AGI has arrived.

How Close Are We To AGI?

There is no reliable single distance-to-AGI measure. Progress depends on whether the milestone is task breadth, unfamiliar-task learning, economic usefulness, autonomy or reliability. Track what models can and can’t do over time. Don’t force five different definitions into a single “months to AGI” number.

Choose the capability your product needs, define the evidence that demonstrates it, and specify what happens when the system fails. You can make a sound build decision without predicting when AGI arrives.

How Is AGI Measured?

AGI cannot be measured credibly with one leaderboard. A credible assessment measures five things: cognitive breadth, novel-task learning, real-world task completion, autonomy and reliability. It also records the exact test conditions behind every result.

MeasureWhat It TestsWhat It MissesProduct Relevance
Cognitive breadth
Performance across mental abilitiesWorkflow reliabilityReveals uneven capability profiles
Novel-task learning
Generalization from limited examplesEconomic usefulnessTests transfer beyond memorized patterns
Task-completion horizon
Sustained work at a stated success rateSocial and physical workShows where human intervention remains necessary
Autonomy and reliability
Independent action and repeatable outcomesFull cognitive breadthSets permission and escalation thresholds

Three measurement approaches are especially useful for product decisions.

Cognitive Capability Profiles

In Measuring Progress Toward AGI: A Cognitive Framework, Google DeepMind proposed a ten-ability taxonomy and an evaluation protocol that compares models with human performance distributions rather than treating one aggregate score as intelligence.

An average score hides uneven skill. A model may reason well but catch its own errors poorly, or write a fluent plan and then fail to update it when conditions change.

Generalization On Unfamiliar Tasks

ARC-AGI tests abstraction from limited examples rather than memorized knowledge. It asks one thing: how much does a model have to be told in advance to solve a problem it hasn’t seen? It remains informative, but no abstract benchmark captures the full complexity of social reasoning and sustained interaction; IEEE Spectrum documents those limits.

Task-Completion Time Horizons

METR’s time-horizon summary and research on longer software-task completion suggest that agent task horizons have been increasing quickly, while also highlighting limits to external validity. The study measures the duration of software-related tasks an agent can complete at a given success rate, expressed in the time a human expert would need. Its results should not be extrapolated directly to all work.

This produces a practical insight for product teams: a model’s ability to execute an individual step does not tell you whether it can own the workflow. A five-minute demonstration and a five-hour process have different failure surfaces.

Intelligence, Autonomy And Reliability Are Different Product Variables

A capable model is not automatically a dependable agent. Intelligence is what a system can figure out. Autonomy is how far it can act without you. Reliability is how often it produces an acceptable result under real conditions.

Consider an agent that can research suppliers, compare terms and prepare a purchase order. Giving it permission to issue payment increases autonomy but does not show it will respect approval thresholds, recover from a changed webpage or detect fraudulent instructions.

The product decision should therefore use three separate tests:

VariableProduct QuestionExample Evidence
Capability
Can the model perform the required reasoning or generation?Task-specific evaluation set
Autonomy
Which actions can it take without approval?Permission map and escalation rules
Reliability
How often does the whole workflow finish within tolerance?End-to-end success rate, exception rate and recovery tests

A fourth variable—consequence—sets how high the bar goes. A wrong internal summary and an unauthorized wire transfer should never share the same acceptance threshold.

This leads to a simple rule: increase autonomy only when measured reliability and the reversibility of failure justify it.

What Should Product Teams Do If AGI Isn’t Here Yet?

Waiting for AGI is usually the wrong strategic response. If a concept has no advantage with current models beyond “the next model will be smarter,” future capability may strengthen competitors just as quickly. The better question: what customer problem can you solve today, and which parts of your product get stronger as models improve?

Better models move the bottleneck; they don’t remove the engineering work. As raw capability rises, your edge shifts to the things a model can’t copy: proprietary context, workflow design, evaluation sets, permissions, distribution, and what your team learns from production.

Any feature that only wraps access to one model is easy to copy. The durable part of an AI product sits elsewhere:

  • a problem important enough that customers change behavior or pay;
  • data or context competitors cannot reproduce easily;
  • integration with the systems where the work already happens;
  • evaluation cases derived from genuine user failures;
  • permission boundaries suited to the consequence of each action;
  • feedback loops that improve the workflow rather than merely store conversations;
  • switching options when a model changes price, latency, policy or capability.

What Should Businesses Build With AI Right Now?

For product teams, the strongest initial opportunities are usually bounded workflows with accessible context, observable outcomes and a clear fallback. Current models create value when the product gives them the right information, tools and limits.

What Bounded Workflows Have In Common

Good candidates often share four conditions:

1. The work is frequent or expensive enough to justify changing the workflow.

2. Success can be evaluated from evidence rather than preference alone.

3. Necessary data and system access can be obtained lawfully.

4. Errors can be detected, reversed or escalated before serious harm occurs.

Examples include research triage, document review, structured content operations, decision support, customer-service assistance, software maintenance and constrained back-office automation. Target a full unit of work—a complete review or a complete triage pass—not a single isolated step.

Portfolio Examples, Not AGI Evidence

The Top Netics venture portfolio illustrates this product-level distinction, not proof that AGI has been achieved. Top Netics is a venture builder and AI product development partner with experience in technology commercialization.

Examples cover bounded workflows: Sapar applies machine learning to Jiu-Jitsu training, Ordder to restaurant operations, Quuannt and Foresee Markets to quantitative analysis, and Senstone Scripter to voice transcription. Each product targets a specific user, context and workflow.

Before engaging an external product engineering partner or committing an internal team, founders should be able to identify:

  • the user decision or action that changes;
  • the minimum context the system needs;
  • the evaluation that predicts useful performance;
  • the human intervention points;
  • the cost of inference and review per completed outcome;
  • the reason customers will choose this workflow over an existing one.

If those answers remain unclear, assess the workflow before funding a full build. Teams comparing external system-design support with internal delivery should use the same evidence.

Teams evaluating an AI product development company should ask for evidence of workflow-evaluation design, model-risk controls, and architecture that can survive provider change.

How To Build For Model Change

A model-independent architecture keeps your business logic, context, evaluation and permissions outside the model. Providers still aren’t drop-in equal. But swapping one becomes a measured test.

1. Define contracts around tasks, not providers

Specify the input, allowed tools, expected output, latency ceiling, cost ceiling and acceptance test for each model-assisted task. Provider-specific calls should sit behind that contract.

2. Build an evaluation set before optimizing prompts

Collect representative cases, difficult edge cases and unacceptable failures. A model upgrade should pass the same regression suite before production traffic moves.

3. Keep business rules outside the model

Eligibility conditions, payment limits, approval requirements and regulatory restrictions belong in deterministic services where possible. Do not rely on a model to remember a policy hidden in a long instruction.

4. Separate retrieval from generation

Record which source supported an output, when it was retrieved and which version was used. This improves auditability and makes stale-context failures easier to diagnose.

5. Give agents the minimum permissions required

Read, draft, recommend and execute are different permission levels. Grant them separately. Require approval for irreversible or high-consequence actions.

6. Design recovery before increasing autonomy

Define what happens when a tool fails, data conflicts, the model exceeds its budget or confidence falls below the operating threshold. An agent without a recovery path is a demo with credentials.

7. Test replacement deliberately

Run at least one alternative model against the evaluation suite. The goal isn’t automatic model-switching. It’s to find where your system quietly depends on one provider.

An experienced product engineering partner should be able to show these boundaries in the system design. “We can swap models later” is not an architectural property until it has been tested.

The AGI-Readiness Checklist

AGI readiness isn’t a forecast. It’s whether you can gain from better models while keeping control of your product, your margins and your safety limits.

Use these questions before approving an AI roadmap:

Customer And Economics

  • Is the customer problem valuable with models available today?
  • Is success defined as a completed outcome rather than model activity?
  • Have willingness to pay and adoption friction been tested?
  • Does the unit cost include inference, tools, review and exception handling?

Evaluation

  • Is there a versioned set of representative production cases?
  • Are capability, reliability and autonomy measured separately?
  • Are failure thresholds stricter where consequences are higher?
  • Can a new model be compared with the current one before migration?

Architecture

  • Are model calls isolated from business rules?
  • Can context sources be traced and refreshed?
  • Are provider-specific dependencies documented?
  • Has an alternative model been tested on the same task contract?

Control

  • Does every tool have the minimum necessary permission?
  • Are irreversible actions approved or otherwise constrained?
  • Can operators inspect, interrupt and recover a workflow?
  • Are data retention, access and jurisdiction requirements enforced outside prompts?

Strategic Advantage

  • Does product value compound through workflow data, integration or distribution?
  • Would better foundation models strengthen the product—or erase its differentiation?
  • Does the team own what it learns from production failures?
  • Is there a clear decision gate between prototype, pilot and scaled deployment?

If your roadmap can’t answer these questions, a stronger model won’t fix the gaps. It will scale them.

Use the checklist to test whether your roadmap is ready for build. Top Netics can map unresolved workflow, risk, evaluation and architecture decisions before engineering begins. Start with a workflow review.

What Should Businesses Build Before The AGI Debate Is Settled?

Build products that solve a verified problem with present capabilities, measure the complete workflow and preserve the option to adopt stronger models. Do not make an AGI forecast carry the business case.

History can decide whether 2026 started the AGI era. Founders face a nearer question: can the product work, earn trust and pay for itself with the models available today?

No general benchmark answers it. Customer evidence, task-level evaluations, clear system boundaries and real production behavior do.

Top Netics works with founders and innovation teams to test those decisions before they harden into expensive architecture. An AI product feasibility assessment maps the user problem, evaluation method, model requirements, data constraints, control points and build path.

Planning an AI product? Book an AI feasibility assessment with Top Netics.

Tagged in:

Frequently asked questions

No universally accepted evaluation has established that AGI has been achieved. Current models demonstrate broad and sometimes superhuman capabilities, but AGI definitions measure different properties, including human-level task breadth, learning efficiency, economic performance and operational autonomy. A system can satisfy parts of one definition without satisfying all the others.

The existence of AGI cannot be answered independently of its definition and test. If AGI means broad competence across many intellectual tasks, current systems show partial evidence. If it also requires reliable learning, sustained autonomy and economic performance across open-ended work, that claim remains unestablished.

ChatGPT is a general-purpose AI assistant, not proof that AGI has been achieved. It can perform many language, reasoning and tool-mediated tasks, but its reliability, autonomy and performance vary by task and operating conditions. Whether it qualifies as AGI depends on the definition and evidence standard being applied.

There is no reliable single distance-to-AGI measure. Progress is rapid across reasoning, software, computer use and long-horizon tasks, but forecasts depend on the milestone selected. The defensible approach is to report capability trends and remaining limitations rather than convert incompatible definitions into one countdown.

Frequently discussed approaches include the ARC-AGI framework for novel-task generalization, Google DeepMind’s cognitive framework for breadth across mental abilities, and METR’s task-completion research for sustained agent performance. Each measures part of the problem. No individual benchmark proves general intelligence across real-world environments.

No. AGI concerns broad intelligence, while agency concerns the ability to pursue goals and take actions through tools. A system can be highly capable but operate under close supervision, or have broad permissions despite limited judgment. Product teams must evaluate capability, autonomy and reliability independently.

No. A viable product should solve a customer problem with models available today. Teams should also isolate model dependencies, create task-specific evaluations and preserve a route to stronger models. Waiting for an uncertain milestone postpones customer learning without guaranteeing a future advantage.

A model-independent architecture places business rules, context management, evaluation and permissions outside the foundation model. Model providers remain technically different, so replacement is never frictionless. The objective is to compare providers against stable task contracts and migrate without rebuilding the product’s core logic.

Ihor Arkhypenko
Written by
Ihor Arkhypenko
Founder & CEO, Top Netics · CIO, Dubai Blockchain CenterTechnology leader known for driving company growth through cross-functional leadership and strategic product development. Passionate about building ecosystems that connect technology, business, and community.
Share