Qwen’s New Omni-Flash Model Prices Multimodal AI at a Fraction of Gemini Flash
egamers.io Sep 24, 2026

Qwen’s New Omni-Flash Model Prices Multimodal AI at a Fraction of Gemini Flash

AI-summarised brief · reviewed before publication

Qwen has launched its Qwen 3.8‑Omni‑Flash multimodal model, priced at $0.15 per million input tokens and $0.47 per million output tokens, compared with Google’s Gemini 3.8 Flash at $0.75 and $3.75 respectively. The model targets AI agents, processing audio and video simultaneously, and claims performance close to Gemini on audio‑video tasks. Qwen estimates audio input costs under $0.01 per hour and a 720p video with audio at roughly $0.20, excluding response generation. The model is available via Qwen Studio, Qwen Cloud, and API, with open‑source plugins for video editing and speaker recognition.

💡 Why It Matters

  • · By offering comparable multimodal capabilities at a fraction of Gemini’s cost, Qwen lowers the barrier for deploying AI agents in media‑heavy applications, potentially accelerating adoption in video editing, translation, and summarization.