AI & Tech

DeepSeek Re-Post-Trains V4-Flash for Agent Work

DeepSeek shipped V4-Flash-0731 on 31 July with no architecture or size change — the gains come entirely from re-post-training. The vendor-published scores are agentic rather than academic: 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, 76.7 on Cybergym, 54.4 on DeepSWE and 70.3 on Toolathlon verified. The model natively speaks the Responses API format and is specifically adapted for Codex, which is an unusually direct statement about where the company expects it to be used. None of these figures have been independently reproduced. Existing endpoints are unchanged; set the model parameter to deepseek-v4-flash.

Read the original — via DeepSeek ↗

← All shorts