AI & Tech
Mistral's Shieldstral Reads Its Moderation Policy at Inference Time
Shieldstral 1.0 is a 3B multimodal safety classifier released under Apache 2.0 with open weights. Unlike guard models that bake policy into weights, it takes policy as input: an instruction setting context, a yes-or-no query, and the text or image being judged. One forward pass returns a calibrated probability. Mistral reports it matching or beating open guard models up to seven times its size — vendor-reported, not yet replicated. It runs on a single 16GB GPU, developed with NVIDIA under the Open Secure AI Alliance. The trade nobody has priced: policy in natural language shares a context window with content that may be adversarial.