Alibaba Qwen officially launched Qwen3.8-Omni-Flash: native audio-video understanding, reasoning, and tool use in one model for agentic workflows such as auto-editing vlogs, translating short videos, and movie recaps. It approaches Gemini 3.8 Flash on audio-video, gains about +19.5 points average agent score on WildClawBench-MM and UniClawBench, offers 1M context, open-sources Qwen-MM-Plugins, and cuts video input cost by about 89% versus Qwen3.5-Omni-Plus.
Key Takeaways
- โOfficial @Alibaba_Qwen release of Qwen3.8-Omni-Flash: native omni-modal + tool orchestration for long-horizon A/V agent workflows.
- โApproaches Gemini 3.8 Flash on A/V; about +19.5 avg agent points on WildClawBench-MM / UniClawBench; 1M context with ~51.8% fewer tokens on OmniVideoBench.
- โ~89% lower video input cost vs Qwen3.5-Omni-Plus; open-sources Qwen-MM-Plugins; Qwen-Live Harness coming; API/Studio/blog live.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.