AI & Tech 🔥 Trending

An Open Harness Beats the Human Expert Baseline on ARC-AGI 3

Prime Agent scored 95.5% on ARC-AGI 3, just past the 95.4% human expert baseline, and it is fully open source. Two ideas carry it. A Recursive Language Model treats context as a variable and subagent delegation as function calls inside a REPL, letting the model manage its own history. A Continual Harness lets the agent rewrite its own prompts, skills, memory and subagents mid-run. It also built Sega Genesis and Game Boy Color emulators from scratch. The sentence to dwell on is Prime Intellect’s own: no model has yet been trained around this harness — the result comes from wrapping stock frontier models.

Read the original — via Prime Intellect ↗

← All shorts