Google launches Gemini Omni to generate video from multimodal inputs

Pradeep Veeraballe··3 min read
googlegeminimultimodalio-2026
An editorial concept showing Google's new Gemini Omni interface generating video from text and image inputs

Google announced Gemini Omni on Tuesday at its Google I/O 2026 developer conference, introducing a new family of multimodal models designed to ingest and generate text, images, audio, and video natively. The rollout begins with Omni Flash, a model that reasons across multiple input formats simultaneously to generate and edit high-quality video outputs.

An editorial concept showing Google's new Gemini Omni interface generating video from text and image inputs

The release marks a shift from stitching separate model outputs together to using a single neural network that processes mixed inputs. Google plans to integrate these capabilities directly into its core products, including a major redesign of Google Search powered by the new Gemini 3.5 Flash model.

Architectural shift to native multimodality

When Google launched the Gemini project three years ago, the company's stated goal was to build a single neural network trained on text, image, audio, and video. TechCrunch's analysis of the announcement notes that Gemini Omni represents a concrete step toward this unified architecture.

Instead of stitching separate models together, Gemini Omni processes all inputs natively. This allows the model to maintain context across different media types, producing high-quality videos that reflect an understanding of physics, culture, history, and science.

Video editing and model positioning

Beyond video generation, Gemini Omni allows users to edit photos using plain text commands, bypassing complex editing software. This capability mirrors Google's Nano Banana model but operates at a larger scale.

Google already has a dedicated video generation model called Veo, which turns text and images into video and customizes avatars. However, Google DeepMind director of product management Nicole Brichtova told TechCrunch that Gemini Omni is a broader architectural step forward.

"It’s the next step towards the progress," Brichtova said.

Google Search interface overhaul

The underlying model advancements are also driving changes to Google's consumer products. The Verge's report on the search updates highlights a redesigned search box that lets users flow between AI Overviews and a conversational "AI Mode."

Powered by the new Gemini 3.5 Flash model, the updated search box expands for longer queries and features AI-powered autocomplete. Users can jump into AI Mode by attaching documents, photos, videos, and Chrome tabs directly to the search box.

The Verge's coverage of Google I/O notes that natural-language questions will reliably trigger AI Overviews. Robby Stein, Google’s vice president of product for Search, confirmed that follow-up questions in an AI Overview will redirect users to the conversational AI Mode.

Model capabilities and rollout details

Google's initial rollout of the Gemini Omni family focuses on three core capabilities:

  • Omni Flash: The first model in the family, processing text, image, audio, and video simultaneously.
  • Native Video Generation: The model generates video outputs that maintain physical and cultural consistency.
  • Text-Based Photo Editing: Users modify images using natural language commands.

These features represent Google's effort to consolidate its generative AI tools into a single, cohesive system. By integrating Gemini 3.5 Flash directly into Search, the company aims to make multimodal reasoning a standard part of daily web queries.

Keep reading

Stay on top of tech and AI

Subscribe wiring is coming soon. For now, follow the daily news feed or connect on LinkedIn for updates.

Read latest newsConnect on LinkedIn