Dev.to
7/21/2026

VIDRAFT Releases Aether-7B-5Attn: A Fully Open-Source MoE LLM with Five Heterogeneous Attention Mechanisms
Short summary
VIDRAFT released Aether-7B-5Attn, a 6.59B-parameter Mixture-of-Experts LLM under Apache-2.0 with five heterogeneous attention mechanisms arranged via a 7×7 Latin Square layout across 49 layers. The model is explicitly bilingual (Korean/English) with ~144B training tokens and ships full training data, code, logs, and checkpoints for reproducibility. Base and instruct variants are available on Hugging Face alongside a live demo.
- •6.59B-parameter MoE LLM with five heterogeneous attention types (full, differential, sliding window, NSA sparse, hybrid) under Apache-2.0
- •7×7 Latin Square layout distributes attention mechanisms across 49 layers to avoid clustering
- •Fully open release includes training data recipes, code, hyperparameters, logs, intermediate checkpoints, and evaluation code
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



