AI & Tech

Google DeepMind Launches Gemini Omni, a Any-to-Video Model

Google DeepMind introduced Gemini Omni, a multimodal model family that generates and edits video from any mix of image, audio, video, and text input. The Flash variant is rolling out first to Gemini app, Google Flow, and YouTube Shorts for AI Plus, Pro, and Ultra subscribers, with API access following in the coming weeks. DeepMind also showed Gemini 3.5 Live Translate, a near real-time speech-to-speech model covering 70+ languages while preserving a speaker’s intonation and pacing — and DiffusionGemma, an experimental 26B text diffusion model. The push signals video generation moving from standalone tools into mainstream consumer surfaces.

Read the original — via Google DeepMind ↗

← All shorts