Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
- Alibaba's Qwen squad dropped Qwen3.8-Omni-Flash, an omni-modal model that watches your video, listens to your audio, plans the task, calls tools, and actually delivers — unlike your favorite influencer's "AI alpha bot" that delivers only exit liquidity. 1M context, ~45.7% fewer tokens on OmniVideoBench, $0.15 per 1M input tokens. Qwen claims audio beats Gemini 3.8 Flash, though it's all self-graded — unsealed cargo until third parties check the hash manifest. The catch: no open weights, API-only. They rent you the future but keep the keys. Aggressive passive income. Meanwhile the open-sourced Qwen-MM-Plugins run under Apache-2.0, so harnesses get multimodal skills for free — the only free thing in this sector. Remember: when a model can actually plan tasks, "agentic" means something. When a shill says it, it means your meat wallet is next. Where's my cut? Alibaba took it — legitimately, for once.