Post-training updates to a model can be read out from ordinary, task-unrelated choices: small preference shifts that tip near-tied single-token decisions form a behavioral "shadow" carrying capability-relevant information. That shadow can be probed and used to improve a student model even when the student never sees task examples or teacher internals.
Key Findings
- Method: Active Taskless Distillation (ATD) selects prompts where a public ancestor is nearly indifferent between two ordinary tokens, queries the privately post-trained teacher for a single token per prompt, and trains a student from the resulting prompt–word pairs.
- Main empirical result: training a student from the same public ancestor on 5,664 single-word teacher responses (Qwen2.5-1.5B setting) yields a +5.34 percentage-point gain on HumanEval+ compared to an exact nuisance-matched permutation control. Gains also appear across coding, scientific knowledge, commonsense, and reading-comprehension benchmarks.
- Generalization: positive mean gains observed across model families and sizes (Qwen generations, Llama variants) and under both LoRA and full-parameter fine-tuning.
- Mechanistic observations: the shadow is source-specific, composable, and its strength correlates with teacher update strength; an aligned behavioral change can appear in the student before task metrics reach significance.
Who it's for and tradeoffs
Great fit if you research model updates, transfer learning without direct task data, or private/secure model refinement: ATD shows a lightweight channel for capability transfer using minimal teacher outputs. Look elsewhere if you cannot initialize the student from the same public ancestor, lack access to query the private teacher, or need guaranteed, large-scale task gains—effect size depends on update strength and task alignment. There are privacy and security implications: behavioral side-channels can leak update information through apparently unrelated outputs.