AI & Tech

Kimi K3 Is 2.8 Trillion Parameters and Trained Natively in 4-Bit

Moonshot’s Kimi K3 is a 2.8T-parameter mixture of experts activating roughly 104B per token, across 93 layers mixing Kimi Delta Attention with Gated MLA, 896 experts with 16 selected per token, 1M context and a native vision path. The architectural story is not the size but the numerics: MXFP4 weights and MXFP8 activations via quantization-aware training, meaning the model was trained to be served at 4-bit rather than squeezed afterwards. That is how a frontier-scale MoE becomes deployable at all. Reported scores include GPQA Diamond 93.5, Terminal-Bench 2.1 88.3 and DeepSWE 67.5. The licence is bespoke, not OSI-approved.

Read the original — via Hugging Face ↗

← All shorts