Key Info

Qwen has unveiled Qwen3.8-Omni-Flash, its first omni-modal model designed around agentic capabilities, combining native audio-video understanding, reasoning, and tool use to autonomously plan and execute tasks.

Highlights

  • The model goes beyond watching and listening: it figures out what to do next and uses tools to get it done.
  • It integrates understanding, task planning, tool execution, and result delivery in a single model.
  • It jointly reasons over audio and video and can orchestrate tools across long workflows, such as auto-editing vlogs, translating short videos, and turning movies into recaps.