AI & Tech

Lilian Weng Argues the Harness Matters as Much as the Model

The post makes the case that the system around a model — execution orchestration, context management, tool integration, workflow coordination — is as load-bearing for self-improvement as the weights. It catalogues the patterns: goal-oriented plan-execute-observe-improve loops, the filesystem as persistent memory to survive long-horizon tasks, sub-agents with inspectable logged outputs, structured context playbooks instead of monotonically growing prompts, and treating harness code itself as an optimisation target. It is equally direct about what is unsolved: weak evaluators for fuzzy tasks, memory lifecycle management, diversity collapse in evolutionary loops, and reward hacking. 312 points on Hacker News.

Read the original — via Lil'Log ↗

← All shorts