Alibaba Qwen open-sourced Qwen3.8-Flash, a 125B-parameter multimodal MoE activating only 6B parameters per token. Delivering 62.5 on SWE-bench Pro and 58.7 on DeepSWE 1.1 at 1/9 the training cost of Qwen3.7-Plus, it previews the Qwen4 architecture with 262K native context (1M with YaRN) and low API pricing ($0.16 / $0.47 per 1M tokens).
Key Takeaways
- โ125B/6B active MoE footprint achieves 62.5 on SWE-bench Pro and 58.7 on DeepSWE 1.1.
- โQSA hybrid attention kernel provides 7.6x prefill and 4.9x decode speedups at 1M context.
- โOpen weights live on Hugging Face with instant day-0 hosting across major API routers.