What Is AGI? The Complete Guide to Human-Level AI in 2026
AGI is the most searched, least agreed-upon term in AI. Everyone from Sam Altman to Yann LeCun uses it, but they rarely mean the same system — and that ambiguity is not accidental.
This guide gives you a durable definition you can reuse, the five levels that make comparisons honest, how we actually measure progress today, and where we stand in late 2026 — so you can judge the next “AGI is near” headline yourself.
What AGI Is — and Why the Definition Matters
In 1956 John McCarthy coined “artificial intelligence” to mean, roughly, machines that can do tasks requiring intelligence when done by humans. That was AGI: the original AI was general.
The qualifier “general” reappeared because the field succeeded at narrow tasks: Deep Blue beat Kasparov at chess but could not book a hotel. The “AI effect” — once a system works, we stop calling the underlying skill “intelligence” — made narrow wins feel less like AI each decade.
Modern definitions converge on three invariants, stated most clearly by Bowen Xu (2024) and DeepMind’s levels paper (2023):
- Adaptation to open environments with limited resources. The system must acquire problem-relevant knowledge itself, not have it hand-coded for each new domain.
- No human rewiring per task. An algorithm is general if a developer can reuse it across problems; a system is general if, after deployment, no developer needs to intervene in source code to handle a new task. AGI is the latter.
- Breadth + depth. Not just specialized performance but the versatility and proficiency of a well-educated adult across reasoning, memory and perception — the standard borrowed from human psychometrics.
In one sentence from a 2025-2026 framework adopted by several labs:
AGI is an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult across the ten broad abilities that define human intelligence.
If the system still needs a new architecture per domain, it is powerful narrow AI, not AGI.
Narrow vs General vs Super — The Levels That Prevent Hype
DeepMind’s 2023 survey (Morris et al., refined 2024-2025) maps AGI on two axes — how broad the capability is and how well it performs:
| Performance → Capability ↓ | Emerging | Competent | Expert | Virtuoso | Superhuman |
|---|---|---|---|---|---|
| Narrow (one domain) | Emerging Narrow | Competent Narrow ← most frontier LLMs today on many benchmarks | Expert Narrow | Virtuoso Narrow | Superhuman Narrow (e.g., AlphaGo on Go) |
| General (across domains) | Emerging General (weak AGI) | Competent General = AGI | Expert General | Virtuoso General | Superhuman General = ASI |
How to read this: Today’s frontier models are Competent or even Superhuman on narrow tasks — coding on HumanEval, image classification, Go — but they remain Emerging or below on general breadth. They transfer poorly without scaffolding, forget long contexts, and lean heavily on retrieval. Calling them “AGI” collapses the matrix.
Practitioners often simplify to five bands for communication:
No AI → Narrow AI (today) → Emerging AGI → Competent AGI (human-level) → ASI (superhuman general)
↑ we are here on many tasks
The simplification is honest only if you keep breadth and performance separate. A model can be expert-level at mathematics and emerging at embodied common-sense — that jaggedness is the signature of current systems.
How We Test for AGI — The CHC Framework and the Scores That Exist
Humans are the only proven example of general intelligence, so the most durable AGI tests borrow from human psychometrics — not trivia quizzes.
The 2025 framework by Hendrycks et al. (updated Oct-Dec 2025, co-authored by Bengio, Tegmark, Brynjolfsson and others) grounds AGI evaluation in Cattell-Horn-Carroll (CHC) theory, the most empirically validated model of human cognition. It decomposes intelligence into ten broad abilities and dozens of narrow abilities — fluid reasoning, crystallized knowledge, working memory, long-term storage and retrieval, visual processing, auditory processing, processing speed and others — then adapts established batteries for AI.
What the numbers say in 2026:
- GPT-4: ~27% AGI on this battery — proficient in knowledge-intensive domains, weak in long-term memory storage and retrieval.
- GPT-5: ~57-58% AGI — rapid progress (+30pp generation-over-generation) but still a jagged profile: strong reasoning with tools, fragile on tasks requiring persistent memory over weeks, embodied planning, and causal world models.
- FrontierCoral / HCI (ArXiv 2609.11873): Uses Headroom-Closed Index to show where RSI would matter most — domains with largest remaining gaps (long-context memory, scientific discovery). On hard coding/science holdouts, closed-loop gains are smallest where verifiers are weakest.
Why naturalistic benchmarks beat LMSYS-style leaderboards for AGI: Leaderboards sample narrow tasks with internet-leaked data. CHC batteries require transfer, few-shot adaptation and memory over time — exactly the skills that define general intelligence. The strongest signal, per the 1,250-paper RSI survey (ArXiv 2607.07663), is whether gains hold on write-protected holdouts the model never saw or edited. If not, capability moved via grader loosening, not intelligence.
Unique angle — The Disagreement Matrix: When you see an AGI prediction, check which column the predictor is optimizing. Our table distills public statements 2024-2026 into testable commitments.
| Who | Public stance (paraphrased, 2024-2026) | Operational definition they use | Implicit test they would accept | Year band they imply |
|---|---|---|---|---|
| DeepMind / Hassabis | Levels framework; SRO for testing powerful systems | Competent General across human-solvable tasks | CHC Competent General + safety evals | Early-mid 2030s if verified |
| Anthropic / Amodei | Pacing the frontier; RSI brake pedal needed | Closed-loop RSI + general task transfer | Independent holdout gains + verifier integrity | 2027-2030 headroom focus, IPO pressure |
| OpenAI / Altman | US-led standards; “not pursuing autonomous RSI unsafely” | Same CHC-style breadth + RSI safety | US standards + CAISI verification | 2028-2033, conditional on safety |
| Meta / LeCun | AGI via world models; LLM sceptic | Grounded, embodied intelligence | Robotics + causal physics benchmarks | 2035+ (LLM path insufficient) |
| Gartner (enterprise lens) | Three stances: imminent / impossible / unpredictable | Organization-ready capability | ROI on composite/neuro-symbolic systems | ”Prepare for all, predict none” |
If two predictions differ but use different definitions, they are not disagreeing — they are measuring different systems.
Timeline: From Dartmouth to 2026 in Ten Beats
- 1956 Dartmouth: McCarthy coins “artificial intelligence” seeking human-level machines.
- 1958-1976: Symbolic era: Lisp, Advice Taker, commonsense logic. Intelligence = formalizable reasoning.
- 1980s: Expert systems: Narrow success (XCON, MYCIN) → Lighthill report → AI Winter. The AI effect begins.
- 1997: Deep Blue beats Kasparov. World calls it “not really AI” within years.
- 2012-2017: Deep learning returns: AlexNet → ResNet → Transformer (2017). Narrow superhuman spreads.
- 2017-2022: Scaling hypothesis: Pre-training compute predicts capability; AlphaGo/AlphaFold narrow superintelligence.
- 2022-2023: ChatGPT moment: General-purpose interfaces convince public that breadth is near — but evaluation shifts to leakage-aware tests.
- 2024: DeepMind Levels of AGI. Field agrees to stop arguing words and start measuring breadth × performance.
- 2025: CHC batteries for AI. GPT-4 at 27%, GPT-5 near 57% — quantifiable gap, not vibes.
- 2026: RSI enters policy language. UN Security Council briefing Sep 23, OpenAI/Anthropic standards talk, California SB 813/AB 1405 — AGI governance becomes deployment governance.
What AGI Is Not — Three Trapdoors
Trapdoor 1: “AGI = chatbot that never fails a demo.” Demos are selected for narrow excellence. General intelligence requires unselected competence — handling the next unseen task without prompt engineering. Ask: was the evaluator fixed and independent?
Trapdoor 2: “AGI = consciousness.” Operational AGI tests capability, not phenomenology. A system can score Competent General on CHC tasks without any claim about felt experience. Keep consciousness debates in philosophy, AGI in measurement.
Trapdoor 3: “AGI is the finish line.” DeepMind’s matrix shows AGI (Competent General) is a midpoint, not an end. ASI (Superhuman General) is a separate level with distinct governance because it exceeds the best humans everywhere — which is where control and economic displacement arguments actually activate.
Living With Narrow AI Until AGI Arrives — What to Build Today
- Design for generality, deploy narrowly. Build systems that could generalize (retrieval, tools, verifiers) but ship them gated to one domain with audited evaluators. That’s how you satisfy both capability and safety.
- Budget for memory. The 27% → 57% gap is not trivia recall — it is long-term storage and retrieval, planning over weeks, and world-model coherence. If your agent forgets last month’s decision, no parameter scaling will save it; architectural memory will.
- Pick your AGI definition and publish it. When customers ask “are you AGI-ready?” answer with levels and with a test you’d accept — not a date. “We define AGI as Competent General on CHC; we track progress via independent holdouts audited by [verifier].” That sentence alone puts you ahead of most marketing.
- Track the right dashboard:
- Breadth × performance matrix position, not single benchmark score
- Evaluator lineage (who wrote the test, who guarded the holdout)
- Verifier strength (formal tests > judge models > self-assessment)
Why This Guide Ages Well (the “Forever” Design)
This page is built to remain useful whether AGI arrives in 2028 or 2045:
- Definitions anchor in CHC and DeepMind levels, not model names. When GPT-6 and Claude 5 ship, re-score them in the same matrix.
- Links point to durable topics, not news: What Is Artificial Intelligence?, How Large Language Models Actually Work, Transformer Architecture Explained, RAG vs Fine-Tuning, Recursive Self-Improvement.
- UpdatedDate bumps only on meaningful measurement changes (new CHC release, verified benchmark on write-protected holdout), not on press releases.
Bookmark it. The next viral “AGI leaked” thread will be easier to parse after you know which level the poster is actually claiming.
Sources
- DeepMind — Levels of AGI — Morris et al. 2023 (updated 2024-2025) — Narrow/General × performance matrix
- ArXiv 2510.18212 — A Definition of AGI — Hendrycks et al. Oct 2025 (Hendrycks, Bengio, Tegmark, Schmidt et al.) — CHC framework, GPT-4 27% / GPT-5 57-58%
- ArXiv 2404.10731 — What Is Meant by AGI? — Bowen Xu 2024 — adaptation to open environments, principles P_G
- Google Cloud & IBM Think — What Is AGI? (updated July 2026) — enterprise definitions, generalization vs narrow
- Gartner — Artificial General Intelligence: 5 Perspectives — imminent/impossible/unpredictable frames
- Merit & arXiv 2607.07663, 2609.11873 — verifier hierarchy, HCI headroom, RSI context for AGI acceleration
- Reuters — AI leaders to brief UN — Sep 23, 2026 — why definitions are now governance terms, not academic trivia
Further Reading on Father of AI
- What Is Artificial Intelligence? Complete 2026 Guide
- Recursive Self-Improvement (RSI): How AI That Builds Itself Works
- How Large Language Models Actually Work
- Transformer Architecture Explained Simply
- What Is Retrieval-Augmented Generation?
- AI 2027 Predictions: What Comes Next
- Who Is the Father of AI? · Who Is the Godfather of AI? · Father of Modern AI — Yann LeCun