Back to feed
Dev.to
Dev.to
7/22/2026
Aether-7B-5Attn: VIDRAFT's Fully Open-Source MoE Foundation Model with Five Heterogeneous Attention Mechanisms

Aether-7B-5Attn: VIDRAFT's Fully Open-Source MoE Foundation Model with Five Heterogeneous Attention Mechanisms

Short summary

Korean AI startup VIDRAFT released Aether-7B-5Attn, a fully open-source MoE foundation model with 6.59B total parameters (2.98B active per token) that integrates five distinct attention mechanisms in a Latin-square layer arrangement. The release includes weights, training code, data recipes, checkpoints, and evaluation code under Apache 2.0. Trained on 144.2B tokens with a multilingual mix emphasizing math (37.8%) and Korean (21.6%).

  • 6.59B parameter MoE model with five heterogeneous attention types in a 7×7 Latin-square arrangement
  • Full open-source release: weights, training code, data recipes, checkpoints, eval code under Apache 2.0
  • Trained on 144.2B tokens with Korean and math emphasis; part of VIDRAFT's broader Darwin model family ranked #1 on K-AI Leaderboard

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more