AI & Tech

Distilling From a Censored Model Did Not Transfer the Censorship

CTGT distilled a 120B open-weight American model using outputs from DeepSeek V4 Flash to improve quantitative finance reasoning, then tested whether the teacher’s political censorship came along. It did not. Across 152 matched prompt pairs, the teacher showed a 45.45-point censorship gap on China-sensitive topics while the student showed 0.26 — statistically indistinguishable from the untouched base model. Four judge models from four different labs scored the responses, correlating 0.948 with human raters. The finding is narrow but useful: capability transfer through task-specific distillation does not automatically carry the teacher’s refusal behaviour. The eval set, rubric and code are released.

Read the original — via CTGT ↗

← All shorts