核心信息
Qwen正式发布Qwen3.8-Omni-Flash,这是其首款围绕智能体能力构建的全模态模型,将原生音视频理解、推理与工具调用整合在一起,能够自主规划并执行任务。
要点
- 该模型不仅能看和听,还能判断下一步该做什么,并通过调用工具完成任务。
- 在单一模型中融合了内容理解、任务规划、工具执行与结果交付。
- 可对音视频内容进行联合推理,并在长时间工作流中调度工具,例如自动剪辑Vlog、翻译短视频、把电影生成剧情回顾。
Qwen正式发布Qwen3.8-Omni-Flash,这是其首款围绕智能体能力构建的全模态模型,将原生音视频理解、推理与工具调用整合在一起,能够自主规划并执行任务。
It doesn’t just watch and listen. It figures out what to do next and uses tools to get it done. Meet Qwen3.8-Omni-Flash.
🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities! Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result. Highlights: 🥳 - Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps. - A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench. - 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench. Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever. To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! 🛠️ We can't wait to see what you build with Qwen3.8-Omni-Flash! 👀 - Blog: https://t.co/oM9V1TkqYF - Qwencloud: https://t.co/Cc2I8ELAnD - Qwen Studio: https://t.co/V7RmqMaVNZ - API: https://t.co/lNE7fH5YUt - Qwen-MM-Plugins: https://t.co/SnM27dDP3d - Qwen-Live Harness: coming soon https://t.co/iVYlGjIbdy