Dev.to
7/22/2026

Aether-7B-5Attn: VIDRAFT's Fully Open-Source MoE Foundation Model with Five Heterogeneous Attention Mechanisms
Short summary
Korean AI startup VIDRAFT released Aether-7B-5Attn, a fully open-source MoE foundation model with 6.59B total parameters (2.98B active per token) that integrates five distinct attention mechanisms in a Latin-square layer arrangement. The release includes weights, training code, data recipes, checkpoints, and evaluation code under Apache 2.0. Trained on 144.2B tokens with a multilingual mix emphasizing math (37.8%) and Korean (21.6%).
- •6.59B parameter MoE model with five heterogeneous attention types in a 7×7 Latin-square arrangement
- •Full open-source release: weights, training code, data recipes, checkpoints, eval code under Apache 2.0
- •Trained on 144.2B tokens with Korean and math emphasis; part of VIDRAFT's broader Darwin model family ranked #1 on K-AI Leaderboard
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



