Dev.to
7/23/2026

Take your benchmark to the people who can kill it
Short summary
The author applied BOLT/PGO-style weight reordering to MoE model binaries, clustering co-activated experts to reduce SSD read scatter. Independent validation on MLX yielded +32.3% decode throughput and −26.3% TTFT on a 235B model. Two of three original optimization pitches were refuted by community measurements, but the file-layout reordering claim survived and migrated to cold-prefill and large-expert-table decode scenarios.
- •MoE expert weight reordering by co-activation clustering reduces SSD read scatter during inference
- •Independent validation on MLX showed +32.3% decode throughput and −26.3% TTFT on a 235B model
- •Community refutation process killed two of three pitches but sharpened the surviving layout-reordering claim
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



