Back to feed
arXiv cs.CL
arXiv cs.CL
8/3/2026
The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

Short summary

This paper shows that knowledge distillation in small instruction-tuned LLMs has asymmetric bias effects: it improves context-following on unambiguous tasks but degrades refusal calibration on ambiguous ones. The authors trace the calibration loss to insufficient refusal-shaped training data and show that aggregate bias metrics conceal per-item harm. They propose PCCD, a three-step protocol that catches both asymmetric bias and trivial-refuser failures missed by standard evaluations.

  • Distillation improves unambiguous-task accuracy but causes 15% of correctly-abstained ambiguous items to receive stereotype answers
  • Silence-loss and filled-silence effects are uncorrelated, arising from distinct mechanisms
  • Proposed PCCD protocol detects calibration failures that aggregate metrics like CrowS-Pairs and BBQ miss

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more