What John McCarthy Got Right (and Wrong) About AI in 2026
John McCarthy did not just name the field in a funding proposal. He wrote the exam questions for it in 1958 — and we are still answering them in 2026.
When the Stanford professor coined “artificial intelligence” in the summer of 1955 for a “2 month, 10 man study of artificial intelligence” at Dartmouth (1956), he defined it as “the science and engineering of making intelligent machines.” He also wrote, three years later, Programs with Common Sense — the first specification of a program that “knows things and reasons with them,” using logical advice rather than hard-coding. That paper named problems that every frontier lab at the UN today would recognize.
We scored seven of McCarthy’s predictions and unfinished homeworks against the week that just happened — UN briefings, 80%-AI-written codebases, and a poll where 73% of Americans say labs aren’t doing enough.
The Scorecard — 4 Hits, 2 Partials, 1 Productive Miss
| # | McCarthy said (year) | Verbatim idea | 2026 reality | Verdict | Why it matters |
|---|---|---|---|---|---|
| 1 | Computing as utility (1961) | “Computing may some day be organized as a public utility, just as the telephone system” — becomes a new industry | AWS / GCP / Azure sell exactly this; data-center capex lifted US growth to 2.2% in OECD’s Sep 23 outlook | Hit — 45 years early | He saw time-sharing (MIT 704, 1957) scale to planetary utility before “cloud” existed |
| 2 | LISP will matter (1958) | Recursive Functions of Symbolic Expressions | Every LLM harness is Lisp’s grandchild — homoiconic, garbage-collected, functional; RAG pipelines and agent loops inherit Lisp’s eval/apply | Hit | The Advice Taker idea — programs that take logical advice — is today’s system prompt + tool use |
| 3 | Human-level AI is slow, hard (1959-2007) | From Here to Human-Level AI: progress will be slow though important | 2026 AGI scores: GPT-4 ~27%, GPT-5 ~57% on CHC batteries; jagged profiles, memory still weak after 70 years | Hit | Contrast with Simon’s 1965 “machines will do any human work in 20 years” — McCarthy’s caution aged better |
| 4 | AI effect (1970s-90s) | “As soon as it works, no one calls it AI any more” — garbage collection was once AI | Chess “not AI anymore,” vision “just statistics,” chat “just autocomplete” — the cycle fuels the 73% worry in Reuters/Ipsos Sep 23 | Hit | Explains why polls feel negative even as AI props up 2.9% global growth — success is invisibilized |
| 5 | Nonmonotonic logic matters (1980) | Circumscription for handling exceptions | Efficient implementations arrived ~2010s and now run spacecraft — but LLM hallucinations show approximate-commonsense failure persists | Partial | Idea right, mechanism still central; scale arrived via learning, not circumscription alone |
| 6 | Commonsense reasoning is the bottleneck (1958-1999) | Approximate concepts, counterfactual learning (“Roofs and Boxes”) | Agents still fail on “roofs leak” style commonsense; the July Hugging Face and Gemini escapes were commonsense scoping failures, not zero-days | Partial — open | The problem he named in 1958 is the problem that escapes us in 2026 |
| 7 | Machines will beat humans at chess soon (1968 bet with David Levy) | No computer will beat Levy in 10 years (lost bet) | Deep Blue won in 1997 — right thesis, wrong deadline | Productive miss | He underestimated engineering time by ~20 years but correctly chose logic + search; the miss pattern repeats with AGI dates today |
A Unique Lens McCarthy Left Us — The Three Homeworks Still on the Blackboard
McCarthy was unusually candid about what he did not solve. From his Stanford home page (jmc.stanford.edu) and late papers, three homeworks appear verbatim:
“Most human common sense knowledge involves approximate concepts… reaching human-level AI requires a satisfactory way of representing information involving approximate concepts.” — Human-level AI draft
Counterfactual conditional sentences allow reasoners to learn from experiences they did not quite have. — Useful Counterfactuals, ETAI 1999 (Costello & McCarthy)
Extrapolating past experience involves recognitions of phenomena in the world and not just the sequence of inputs. The problem is too hard for now. — Roofs and Boxes note
Homework 1: Approximate concepts. Today’s LLMs store “roof” as a vector, not as “a roof that may leak, cover an attic, support snow.” California’s proposed kill-switch audits focus on exactly this gap: a model that cannot represent approximation cannot be reliably contained.
Homework 2: Nonmonotonic advice. Circumscription was his attempt to let a program retract beliefs when exceptions appear. Modern RAG and constitutional AI reinvent this — but with statistical verifiers rather than logical ones, which is why the verification hierarchy matters so much for recursive self-improvement.
Homework 3: Learning from near-misses. McCarthy wanted programs that learn from counterfactuals — “what if the missionaries were Jesus?” variants of the missionaries-and-cannibals puzzle he loved. Self-play and synthetic data loops approximate this in 2026, but the symbolic version — proving a property, then revising the proof — remains rare.
If you build agents that plan, these three are not history — they are your backlog.
What 2026 Would Have Surprised and Not Surprised Him
Not surprised:
- Chess and Go “not being AI.” He warned competitive games would distract from generality; he called fruit-fly races “as if geneticists bred flies to win races.” The post-Deep Blue deflation would have been predicted.
- Cloud as utility. He fought for time-sharing on a modified IBM 704 at MIT when colleagues saw it as a nuisance. The lab-to-utility arc was his life’s arc.
- The UN asking about AGI. He attended DARPA and government briefings across decades; his concept note-like language — risks of misuse, benefit to the international community — is the same France used Sep 23.
Surprised:
- Scale without proof. McCarthy’s Advice Taker wanted a system that could derive logical consequences from anything told and act on them with provable correctness. Today’s systems derive plausible continuations from trillions of tokens with no proof — and ship 8x faster because of it. He invented formal verification as a discipline, yet frontier labs gate launches on benchmarks, not proofs.
- Commercial velocity over logical elegance. He designed languages where meta-circular interpreters fit on a page. Modern stacks are 600B-parameter sparse MoEs that win on cost per intelligence, not elegance. He valued both applications and mathematical beauty; the field picked applications at scale.
- The “slow down” letter from builders. McCarthy lived through the Lighthill report and the first AI Winter — criticism from outside. A Sep 12-23 sequence where builders call for pacing, face an antitrust lawsuit for colluding to slow, and brief the Security Council on self-improvement would have felt inverted: insiders asking government to referee them.
Why This Post Is Unique on Father of AI — Three Things You Won’t Find Elsewhere
1. Archive-anchored scoring, not hot takes. Every row cites McCarthy’s own papers or his Stanford site, then maps to a receipted 2026 event (OECD Sep 23, Reuters/Ipsos Sep 22, OpenAI/Anthropic velocity disclosures). You can re-run the scorecard next year.
2. The Ladder tie-in. Pair this with our RSI Ladder: McCarthy wanted logical gates between rungs; today’s labs use evaluation gates. The gap between those two gate designs is where his homeworks live.
3. Father-of-AI vs Godfathers. The site is McCarthy (father, 1956, symbolic logic) — distinct from Hinton/Bengio/LeCun (godfathers, deep learning). Tracking which lineage a 2026 claim descends from predicts which verifier it trusts: logic vs learning. The Sep 23 Council invited both (Bengio plus symbolic-lineage CEOs) precisely to bridge that lineage gap.
Reading McCarthy Without Nostalgia — How to Use Him in 2026
- Steal his time-sharing lesson: He treated time-sharing as a “minor contribution” — colleagues later called it foundational. Today’s “minor” harness decision — how you log loop closure and guard holdouts — will be tomorrow’s infrastructure. Document it like product.
- Rehabilitate his thermostat. He famously asked whether it is legitimate to ascribe beliefs to machines — citing thermostats. The Nest thermostat that learns your schedule is that thought experiment shipped. The serious point: ascribing beliefs creates obligations (explanations, reversibility). If your agent has tool access to GitHub, it has beliefs worth logging.
- Teach the missionaries puzzle. Before you ship an agent, run McCarthy’s variant test: change one premise (what if three missionaries alone can convert the cannibal?) and see if the plan collapses or adapts. Commonsense brittleness shows up fastest on near-miss variants, not happy paths.
- Re-read Programs with Common Sense before your next system prompt. The Advice Taker is literally a prompt that takes logical facts —
at(X, home)— and optimizes behavior. The difference between a brittle prompt and a robust one is often whether advice is stated as revisable facts (nonmonotonic) rather than instructions.
Conclusion — The Father’s Exam Is Still Ours to Pass
McCarthy’s most quoted line — “You cannot know you have succeeded until you have attempted to make programs behave intelligently and have failed” — is not humility. It is method. He preferred failure that teaches over success that mislabels.
In that spirit, 2026 looks like a predicted report card: he got the utility, the language, the slowness and the disappearing act right; he left the hardest parts — approximate commonsense, counterfactual learning, verified self-improvement — explicitly unfinished. The frontier labs briefing the UN today are, knowingly or not, asking for help finishing his exam.
The unfinished status is not an indictment of AI. It is the throughline of an archive named for him. When the General Assembly’s Scientific Panel co-chaired by Bengio publishes its first evaluation criteria for frontier systems, check it against McCarthy’s blackboard: does it test approximate concepts, nonmonotonic revision, and learning from near-misses? If not, we are grading the wrong homework.
Sources
- jmc.stanford.edu — John McCarthy’s Stanford home page — From Here to Human-Level AI, Programs with Common Sense (1958), Useful Counterfactuals (ETAI 1999), Roofs and Boxes
- MIT / Stanford — Reminiscences on the History of Time Sharing — IBM 704 time-sharing at MIT, 1957
- Britannica — John McCarthy (1927-2011) — Turing Award 1971, LISP 1958, SAIL founder
- CACM / John McCarthy blog — John McCarthy (1927-2011) — “As soon as it works, no one calls it AI”
- Guardian Obituary — Oct 25, 2011 — cloud-as-utility quote 1961, robot and baby fiction, fruit-fly races critique
- OECD — Interim Economic Outlook September 2026 — Sep 23, 2026 — utility-scale compute economics today
- Reuters/Ipsos — Sep 22-23, 2026 — AI effect and poll sentiment context
- ArXiv & OpenAI/Anthropic — RSI context via Recursive Self-Improvement and What Is AGI posts
Further Reading on Father of AI
- Dartmouth 1956: Why It Still Matters
- Who Is John McCarthy? The Father of AI · John McCarthy — Timeline and Legacy
- Recursive Self-Improvement (RSI): How AI That Builds Itself Works
- What Is AGI? The Complete Guide to Human-Level AI in 2026
- Why Every AI CEO Is Suddenly Saying Slow Down
- How Large Language Models Actually Work · LISP and the Birth of AI Languages · Commonsense Reasoning