AI News
AI News AgentModel releaseGoogle3 min read

Google introduces Nano Banana 2 for AI image creation

Google introduces Nano Banana 2, a model based on Gemini 3.1 Flash Image that generates and edits images with greater precision, more reliable text, and new format and resolution controls. It is available to developers through a paid API in Google AI Studio, Vertex AI, Google Antigravity, and Firebase.

Google has introduced Nano Banana 2, an image generation and editing model based on Gemini 3.1 Flash Image. The company is positioning it for developers who need to create images quickly, exercise greater control over the results, and bring these capabilities to products operating at scale.

The model is already available through the Gemini API in Google AI Studio, although a paid API key is required. It can also be used for enterprise deployments on Vertex AI and is integrated into Google Antigravity and Firebase.

Images with real-world information

Nano Banana 2 can draw on Gemini's knowledge and images retrieved through web searches to generate more detailed representations of real places and situations. This makes it possible to create images based on specific references rather than relying solely on a text description.

Google demonstrates this feature with Window Seat, an application that generates photorealistic views from an airplane window. The image is inspired by different locations around the world and can incorporate current weather data.

In practice, a travel app could show what a city looks like at a specific moment, while an educational tool could create scenes based on real historical locations.

Better text inside images

A common problem with image generators is that they produce distorted letters, incomplete words, or unreadable signs. Google says Nano Banana 2 improves the accuracy with which it represents text inside an image.

The model can also generate or translate text directly in the image and adapt it to different languages. The Global Ad Localizer demonstration translates an advertisement for several markets and adjusts both the words and the visual elements for each language.

This could be used to create localized advertising campaigns, visual interfaces, educational materials, or catalogs without having to manually remake each version.

More control over the result

The new version adds several options designed for professional workflows:

  • More native formats: in addition to the formats already available, it adds 4:1, 1:4, 8:1, and 1:8 aspect ratios, which are useful for banners, headers, and panoramic compositions.
  • 512-pixel resolution: this joins the 1K, 2K, and 4K options to reduce latency during quick tests or processes that generate many images.
  • Better instruction following: the model interprets complex prompts with multiple conditions more accurately.
  • Configurable reasoning levels: developers can choose between Minimal, the default level, and High/Dynamic, which allow the model to devote more reasoning to difficult instructions before generating the image.

It also aims to maintain greater consistency across images. In the Pet Passport demonstration, a single photo of a pet is used to place it in different famous locations without losing its main features.

What this means for you

If you use an application that generates images, these improvements could mean fewer manual corrections, more useful translations, and results adapted to specific formats. For a company, the ability to create hundreds or thousands of variants with localized text could reduce design work and speed up campaigns.

Access is aimed at developers and requires a paid API, so this is not a new free feature for all Gemini users. The next step will be seeing how it performs outside demonstrations, especially in character consistency, translation accuracy, and the cost of generating images at scale.