Back to feed
Dev.to
Dev.to
8/3/2026
The original headline is: "Qwen3.8-Max Just Went GA: A Developer's Guide to Alibaba's 2.4T Model"

The original headline is: "Qwen3.8-Max Just Went GA: A Developer's Guide to Alibaba's 2.4T Model"

Original: Qwen3.8-Max Just Went GA: A Developer's Guide to Alibaba's 2.4T Model

Short summary

Alibaba released Qwen3.8-Max GA on August 3, 2026: a 2.4T-parameter sparse MoE model with ~95B active parameters, 1M-token context, multimodal input, and OpenAI-compatible API at $2/M input and $6/M output. The article explains why active-parameter count drives cost, how explicit caching reduces input cost 8x, and why prefix stability matters more than prompt length for production architectures.

  • 2.4T total / ~95B active MoE model with 1M context window at $2/M input, $6/M output
  • Explicit cache reads cost $0.17/M vs $2.00/M fresh — prefix stability cuts agent-loop costs ~9x
  • Open weights promised next week; pricing undercuts Kimi K3 ($3/$15) significantly

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more