Dev.to
8/3/2026

The original headline is: "Qwen3.8-Max Just Went GA: A Developer's Guide to Alibaba's 2.4T Model"
Original: Qwen3.8-Max Just Went GA: A Developer's Guide to Alibaba's 2.4T Model
Short summary
Alibaba released Qwen3.8-Max GA on August 3, 2026: a 2.4T-parameter sparse MoE model with ~95B active parameters, 1M-token context, multimodal input, and OpenAI-compatible API at $2/M input and $6/M output. The article explains why active-parameter count drives cost, how explicit caching reduces input cost 8x, and why prefix stability matters more than prompt length for production architectures.
- •2.4T total / ~95B active MoE model with 1M context window at $2/M input, $6/M output
- •Explicit cache reads cost $0.17/M vs $2.00/M fresh — prefix stability cuts agent-loop costs ~9x
- •Open weights promised next week; pricing undercuts Kimi K3 ($3/$15) significantly
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



