Back to feed
Dev.to
Dev.to
7/23/2026
Take your benchmark to the people who can kill it

Take your benchmark to the people who can kill it

Short summary

The author applied BOLT/PGO-style weight reordering to MoE model binaries, clustering co-activated experts to reduce SSD read scatter. Independent validation on MLX yielded +32.3% decode throughput and −26.3% TTFT on a 235B model. Two of three original optimization pitches were refuted by community measurements, but the file-layout reordering claim survived and migrated to cold-prefill and large-expert-table decode scenarios.

  • MoE expert weight reordering by co-activation clustering reduces SSD read scatter during inference
  • Independent validation on MLX showed +32.3% decode throughput and −26.3% TTFT on a 235B model
  • Community refutation process killed two of three pitches but sharpened the surviving layout-reordering claim

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more