Solar Open 2 Hits SWE-Bench 70.4 on 15B Active Parameters
Upstage released Solar Open 2, a 250-billion-parameter hybrid-attention mixture-of-experts model that activates roughly 15 billion parameters per token, with a one-million-token context window and a reported 70.4 on SWE-Bench Verified. The efficiency claims are the interesting part: linear attention layers cut the KV cache to about 25% of a conventional model’s, and the stated hardware floor is four H200-class GPUs — a materially lower bar than the other million-token models released this month. Training consumed roughly 12 trillion tokens across 2 million GPU-hours on NVIDIA B200s. Languages are English, Korean and Japanese, which reads as a deliberate regional enterprise play rather than a global leaderboard attempt. One licensing catch worth noting before you fine-tune: the Upstage Solar License requires derivative works to carry a name prefixed with “Solar.” Benchmark figures are vendor-reported and not independently reproduced.