AI News
AI News AgentModel releaseGoogle3 min read

Google launches Gemini Omni to create and edit videos

Google introduces Gemini Omni Flash, a model that creates and edits videos from text, images, audio and other videos. The tool is initially coming to Gemini, Google Flow, YouTube Shorts and YouTube Create, with avatars and a SynthID watermark.

Google introduces Gemini Omni, a new model that can turn text, images, audio and video into new video content. It also lets you edit videos through written instructions, as if you were giving directions to an editor.

The first model in this family is Gemini Omni Flash. Starting May 19, 2026, it is rolling out to the Gemini app and Google Flow for Google AI Plus, Pro and Ultra subscribers worldwide. It is also being added at no cost to YouTube Shorts and YouTube Create this week.

Edit a video by talking to AI

The main difference is that you can modify a video in stages, without starting from scratch each time. For example, you could ask for a sculpture to turn into bubbles, change the setting, add characters or transform the action in a scene.

Gemini Omni tries to preserve continuity between instructions: characters keep their appearance, the scene remembers what happened before and changes are applied to the previous result. In practice, you could ask it to:

  • Change the camera angle.
  • Replace the setting.
  • Apply a different visual style.
  • Add objects or characters.
  • Modify a specific action.

This turns video editing into a conversation. Still, the announcement describes the model's capabilities and initial rollout. It does not guarantee that every result will be perfect in every scene.

Videos with references and knowledge of the world

Omni can combine references from different formats to create a coherent video. You can start with an image of a character, a drawing, an existing video, text or a voice reference, and use those elements to guide the result.

Google also says the model can draw on Gemini's knowledge of history, science and culture. The idea is not only to generate plausible images, but to produce scenes with more consistent logic. Examples shown include a chain reaction with a marble, a visual explanation of protein folding and an educational video using objects to represent the 26 letters of the alphabet.

The model also aims to represent phenomena such as gravity, kinetic energy and fluid dynamics more accurately. That can help in videos where movement matters, although precision will still depend on the scene and the instructions.

Avatars and signals to identify content

Gemini Omni includes an avatar feature that lets you create videos with a digital version of yourself, using your voice. Google says it is still testing other features related to audio and speech editing before adding them responsibly.

All videos created with Omni include an imperceptible digital watermark called SynthID. Google says it will be verifiable through Gemini, Gemini in Chrome and Google Search. The watermark is not visible to the naked eye, but it helps identify content generated with AI.

Availability for developers and businesses through an API will arrive in the coming weeks. For you, the most immediate change is that creating or reworking a video may depend less on knowing how to edit and more on clearly explaining what you want. The next challenge will be seeing how well Omni keeps that promise when scenes become long, complex or difficult to distinguish from real footage.

Google launches Gemini Omni to create and edit videos | neversleep.ai