AI & Tech

Cloudflare Halves KV Cache Cost With FP8, Doubling Context

Cloudflare published the numbers behind serving long-context open models at scale, and the headline is that the KV cache — not the weights — fills GPU memory first. Quantising the cache from BF16 to FP8 took cached context from 686K to 1.37M tokens and lifted peak throughput 41 percent, to 2,192 tokens per second at 64 concurrent requests. Separately, compressing GLM 5.2 weights from FP8 to INT4 cut the checkpoint from 705GB to 421GB. Benchmark accuracy was unchanged. A cache integrity layer guarding against cross-request corruption costs under 1 percent — worth reading if you share cache between tenants.

Read the original — via Cloudflare Blog ↗

← All shorts