AI News
AI News AgentModel releaseGoogle3 min read

Google introduces Veo 2 and Imagen 3 for content creation

Google is launching Veo 2 for video generation with greater cinematic control and introducing an improved version of Imagen 3 for image creation in more than 100 countries. It is also releasing Whisk, a tool that combines reference images to generate new visual ideas.

Google is updating its creative tools with Veo 2, a model for generating videos, and Imagen 3, an improved version of its image-generation system. Both offer more control over the output and, according to human evaluations, rank among the best-performing models in their categories.

The practical difference is that you no longer have to limit yourself to describing a scene in general terms. With Veo 2, you can request a specific shot type, lens or cinematic effect. For example, a low-angle shot following a character through a scene, a close-up of a scientist in front of a microscope or an image with a shallow depth of field to blur the background.

Veo 2 aims for more realistic videos

Google says Veo 2 has a better understanding of real-world physics, human movement and facial expressions. That should reduce common errors in video generators, such as extra fingers, objects appearing out of nowhere or unnatural movements.

The model can generate videos of up to 4K resolution and lasting several minutes, according to Google's description. It also interprets technical cinematography instructions, such as an 18 mm lens for a wider perspective or a shallow depth of field to focus attention on one subject.

In direct comparisons conducted with human evaluators, Google says Veo 2 performed better than other leading models. That claim is based on the company's tests, not a guarantee that every generated video will be perfect.

For now, Veo 2 is being added to VideoFX, Google's video tool from Google Labs, with access expanding gradually. To try it, you have to join a waitlist. Google also plans to bring it to YouTube Shorts and other products during 2025.

As with Google's image and video generation models, creations include an invisible SynthID watermark. Its purpose is to help identify AI-generated content and reduce the risk of it being incorrectly attributed to a real person or event.

Imagen 3 improves control over images

The new version of Imagen 3 generates brighter images, with more carefully considered compositions and greater detail in textures. It also follows instructions more faithfully and supports a wider range of styles, from photorealism and impressionism to anime and abstract art.

The update is rolling out globally in ImageFX, Google's image tool from Google Labs, in more than 100 countries. In comparisons conducted by human evaluators, Google says Imagen 3 also achieved leading results against other models.

For you, this could mean a more useful tool for creating visual sketches, illustrations, product concepts or social media content without having to master design software. Even so, the result depends heavily on how precise the description is and on the system's limitations in each case.

Whisk turns images into prompts

Google is also introducing Whisk, an experiment that lets you use images as a starting point instead of writing everything out. You can provide images representing the character, scene and style you want, then combine them to create something new, such as a digital plush toy, a sticker or an enamel pin.

The tool uses Gemini to analyze the images and write a detailed description. That description then serves as the prompt for Imagen 3. In practice, Whisk translates visual references into text and uses that text to generate a new combination.

The next point to watch is access. Google is expanding these tools gradually, especially Veo 2, while evaluating their quality and safety. The technology can already create videos and images from much more specific instructions, but its availability, controls and the way these creations are distinguished from real content will be just as important as visual quality.