arXiv cs.CL
7/24/2026

Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought
Short summary
The paper reveals that MoE routing behaves like Huffman coding, with models allocating sparse experts for common tokens and diverse expert committees for rare, complex chain-of-thought tasks. The authors identify a redundancy trap in Qwen3.5-35B-A3B where load-balancing masks Huffman efficiency, and propose Subset Difference Pruning to eliminate functional duplicates. They argue next-gen MoEs should move toward Minimum Description Length optimality rather than forced load-balancing.
- •MoE routing follows a Frequency-Diversity Law analogous to Huffman coding
- •Load-balancing can mask underlying Huffman efficiency in MoE models
- •Subset Difference Pruning eliminates redundant experts without degrading reasoning
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

