Nano Banana 2.1: Specs, Prompts, API & Pricing
Nano Banana 2.1 specs, prompts, API model ID, pricing, comparisons with Nano Banana 2 and Pro, and a practical image-to-video workflow.
Tags
Try Nano Banana 2.1 while you read.
Developer: Google DeepMind · API model ID: gemini-nano-banana-2.1 · Output: 1K, 2K, and 4K · Current VioEvo route: Nano Banana 2.1 · Official sources: Google Gemini API model page · Google DeepMind Nano Banana
Version note: Nano Banana 2.1 is the current high-efficiency update to Nano Banana 2. For the family naming map and the short differences between Nano Banana, Pro, 2, 2.1, and Lite, see the Nano Banana family guide. The API facts below follow Google's current documentation; the recommendations and test rubric are editorial guidance.
Nano Banana 2.1 is built for image generation and conversational editing when you want Flash-level efficiency without giving up higher-resolution output or reference control. The useful distinction from a generic “AI image generator” is the workflow: you can combine text, images, video, or PDF inputs, ask for a focused edit, and choose how much thinking the model should use.
Nano Banana 2.1 Timeline
| Version | Public release or official API-page date | What changed |
|---|---|---|
| Nano Banana | Gemini 2.5 Flash Image API page last updated October 2025; the model page does not state the original launch date | Established the high-volume, low-latency conversational image workflow. |
| Nano Banana Pro | Gemini 3 Pro Image API page last updated November 2025; the model page does not state the original launch date | Added a higher-capability route for demanding generation and editing. |
| Nano Banana 2 | Gemini 3.1 Flash Image API page last updated February 2026; TechCrunch reported its launch on February 26, 2026 | Made the Flash-level image route the mainstream speed and throughput choice, with image generation, editing, and search grounding. |
| Nano Banana 2.1 | Stable model; Gemini API page latest update October 2026. Google does not state a separate launch date on the page. | Improves visual quality, realism, prompt adherence, text rendering, multi-turn consistency, wide-aspect output, reference fusion, grounding, and thinking controls. |
Official model pages: Nano Banana, Nano Banana Pro, Nano Banana 2, and Nano Banana 2.1. “Latest update” is a documentation date, not necessarily a release date; where Google does not publish a launch date on the model page, the timeline says so.
What Is New in Nano Banana 2.1?
Google describes 2.1 as an update to Nano Banana 2 (gemini-3.1-flash-image) with improvements in visual quality and realism, prompt adherence, multi-turn character consistency, and text rendering. The API page also lists four concrete changes that matter in production:
- 1K, 2K, and 4K output: 1K is the default; 2K and 4K are available when the asset needs more pixels.
- Wide and panoramic fixes: Google says tiling artifacts were fixed for 1:4, 4:1, 1:8, and 8:1 aspect ratios at 2K and 4K.
- Multi-image fusion: Up to 14 reference images, described as supporting character consistency for up to four characters and object fidelity for up to ten objects.
- Grounded generation and thinking: Google Web and Image Search grounding is available, and thinking can be set to minimal, medium (default), or high.
These are documented capabilities, not a promise that every prompt will pass a production review. For a quality decision, run the test plan below with the same references and acceptance criteria.

Nano Banana 2.1 sample: one subject appears in a full-body portrait and three framed views. This is a single generated composition, not a multi-turn consistency benchmark.
Nano Banana 2.1 Specs
| Specification | Current documented behavior |
|---|---|
| API model ID | gemini-nano-banana-2.1 |
| Inputs | Text, image, video, and PDF |
| Outputs | Image and text |
| Output resolutions | 1K (default), 2K, and 4K |
| Input token limit | 131,072 |
| Output token limit | 32,768 |
| Reference images | Up to 14 in multi-image fusion |
| Search grounding | Google Web Search and Image Search |
| Thinking | Minimal, medium (default), and high |
| Audio generation | Not supported |
| Function calling | Not supported |
| Batch API | Supported |
Source: Gemini Nano Banana 2.1 model documentation. Google can change limits and availability; re-check the model page before implementation.
Nano Banana 2.1 API
For API work, use the exact model ID gemini-nano-banana-2.1 with the Gemini API. The official image generation guide documents the request structure, image configuration, multi-turn editing pattern, and safety behavior. The model page lists text, image, video, and PDF inputs, image and text outputs, search grounding, thinking, and Batch API support.
Google's API documentation is the authority for parameters. VioEvo's selector, nano-banana-2.1, is a product identifier and should not be copied into a direct Gemini API request.
Nano Banana 2.1 Pricing and Free Access
Google's current Gemini API pricing page lists no free-tier model price for Nano Banana 2.1 and a standard paid rate of $30 per 1 million image-output tokens. Google's own image-equivalent estimates are $0.0336 per 1K image, $0.0504 per 2K image, and $0.113 per 4K image. Text and thinking output are listed at $7.50 per 1 million tokens, and text/image/video input at $1.50 per 1 million tokens. Batch rates are lower: $15 per 1 million image-output tokens, with image equivalents of $0.0168, $0.0252, and $0.0567 for 1K, 2K, and 4K.
The pricing page also lists 5,000 free Google Web and Image Search grounding requests per month shared across Gemini 3.x models, then $14 per 1,000 requests for text and image-based grounding. That is a grounding allowance, not a free allowance for generated images.
| Route / tier | Free access | Published image-output reference |
|---|---|---|
| Nano Banana 2.1, Google API standard | No free-tier image price listed | $0.0336 (1K), $0.0504 (2K), or $0.113 (4K) per image equivalent |
| Nano Banana 2, Google API standard | No free-tier image price listed | $0.067 per 1K image equivalent; current pricing page lists 0.5K, 1K, 2K, and 4K output |
| Nano Banana Pro, Google API standard | No free-tier image price listed | $0.134 per 1K/2K image equivalent; $0.24 per 4K image |
| GPT Image 2.5, OpenAI API | Account access and billing depend on OpenAI API eligibility | $30 per 1 million image-output tokens; per-image cost varies with image token usage |
| Qwen Image 2.1 | Not confirmed | No official 2.1 API rate found; Qwen's official repository currently documents Qwen-Image-2.0 |
| VioEvo route | Current credit cost |
|---|---|
| Nano Banana 2.1, 1K | 20 credits per image |
| Nano Banana 2.1, 2K | 30 credits per image |
| Nano Banana 2.1, 4K | 50 credits per image |
| Eligible new account | 150 welcome credits when the current welcome-gift policy grants them; credits expire after 30 days |
The Google and OpenAI rows are provider API prices; they use different billing units and are not direct per-image comparisons. Qwen's row is intentionally left without a price because no official Qwen Image 2.1 API rate was confirmed. The VioEvo rows are the current product contract, not provider API prices. New-account credits are subject to eligibility, abuse checks, and the active welcome-gift policy. Check the VioEvo pricing page and live tool controls for current account-specific access, credits, and plan limits. Google rates can change; see the official Gemini API pricing table and OpenAI image generation guide before budgeting an API integration.
Nano Banana 2.1 vs Nano Banana 2
Nano Banana 2.1 is the documented update to Nano Banana 2, so the comparison should focus on the delta rather than repeating the family overview.
| Test dimension | Nano Banana 2 | Nano Banana 2.1 | How to measure |
|---|---|---|---|
| Realism | Gemini 3.1 Flash Image baseline | Google says visual quality and realism improved | Blind-score skin, materials, lighting, geometry, and object interactions on the same 20 prompts |
| Text rendering | Search-grounded image generation and editing | Google specifically calls out more accurate text rendering | Count exact character, word, placement, and line-break passes in the same poster and label set |
| Instruction following | Flash image generation and editing baseline | Google says prompt adherence improved | Check every required attribute and forbidden attribute against a prewritten checklist |
| Consistency | Image editing and multi-turn workflows | Google adds multi-turn character consistency and up to 14-image fusion | Measure identity and product-feature retention across three edits and a multi-reference set |
| Price | Google's current table lists a $0.067 1K image equivalent | Google lists $0.0336 per 1K, plus 2K and 4K equivalents | Compare cost per accepted image, not cost per attempt; include retries and review failures |
Sources: Google's Nano Banana 2 API model page, Nano Banana 2.1 API model page, and Gemini API pricing. The test protocol is editorial guidance; quality deltas are Google's claims.
The quality deltas above are Google's release claims. No universal public benchmark establishes a winner for every prompt. Run the same prompts, references, resolutions, thinking level, and review rubric before migrating.
Nano Banana 2.1 vs Nano Banana Pro
Pro and 2.1 occupy different positions in Google's lineup: Pro is the higher-capability choice, while 2.1 is the high-efficiency choice with a detailed Flash-family API contract. A useful comparison asks which model produces an accepted asset within the workflow's latency and budget constraints.
| Test dimension | Nano Banana Pro | Nano Banana 2.1 | How to measure |
|---|---|---|---|
| Realism | Pro-level image generation and editing positioning | Google reports improved realism over Nano Banana 2 | Use the same portrait, product, and material prompts; score artifact rate and detail preservation blind |
| Text rendering | Test the current Pro endpoint's text behavior | Google specifically highlights accurate text rendering | Use identical multilingual labels, small type, and layout constraints; record exact-copy pass rate |
| Instruction following | Higher-capability tier, with current limits documented separately | Configurable thinking and improved prompt adherence | Require every constraint to pass; record omissions, substitutions, and extra objects |
| Consistency | Evaluate Pro's reference and multi-turn behavior from current docs | Up to 14 references and up to four characters/ten objects documented | Run three-turn edits and reference-fusion prompts; score identity, product geometry, and style drift |
| Price | Google's current table lists $0.134 per 1K/2K image and $0.24 per 4K image | $0.0336 / $0.0504 / $0.113 image equivalents for 1K / 2K / 4K | Compare accepted-image cost, including retries and human review time |
Sources: Google's Nano Banana Pro API model page, Nano Banana 2.1 API model page, and Gemini API pricing. The test protocol is editorial guidance.
Do not describe Pro as universally better from the product name alone. The official sources do not publish a single apples-to-apples benchmark covering every dimension above. Use the rubric and keep the model, prompt, resolution, and reference inputs controlled.
Nano Banana 2.1 Prompt Examples
The following prompts are designed to be copied into a text-to-image or conversational-editing request. Put the subject first, then the composition and constraints. For iterative edits, state what must change and what must remain fixed.
The onsite samples illustrate related tasks; they are not outputs from these exact prompts.
Text rendering prompt
Create a clean 4:5 editorial poster for a fictional coffee brand. Render the exact headline "NORTHLINE COFFEE" at the top, the exact subheading "Small batch. Bright roast." below it, and the exact price "$18" in a small badge. Use only those words, spell every character correctly, keep all text legible at poster scale, and leave generous margin around the typography. Warm daylight, cream paper texture, dark green ink, premium modern layout.

Nano Banana 2.1 text-to-image sample: a breakfast menu board combines large headlines, smaller copy, and food images. Check every small line as well as the headline when reviewing generated text.
Character consistency prompt
Use the supplied reference images to create a three-panel storyboard of the same four characters in a city bakery. Preserve each character's face, hair, age, clothing colors, and distinguishing accessories across all three panels. Change only the camera angle and action: ordering, carrying a tray, then sitting at a window. Keep the bakery layout, morning light, and illustration style consistent. Do not add or remove characters.
4K prompt
Set the requested output size to 4K in the image configuration; prompt text alone does not set the API output resolution.
Generate a 4K landscape product photograph of a brushed-aluminum camping lantern on a wet granite ledge after rain. Show fine water droplets, realistic brushed metal, a soft warm internal glow, and a distant mountain valley in the background. Preserve clean edges around the lantern, physically plausible reflections, and enough negative space on the left for a headline. No text, logos, or extra products.
Image-editing prompt
Edit the supplied product photo. Replace only the background with a sunlit pale-blue studio wall and add a soft contact shadow beneath the product. Preserve the product silhouette, logo, label text, camera angle, scale, highlights, and all existing edges. Do not change the product color or add props. If any background detail is ambiguous, prefer a clean studio surface.

Nano Banana 2.1 image-to-image sample combining a street portrait with a coffee-pouring mural. Review the hand, cup, and overlap between the person and the painted figure.
Nano Banana 2.1 vs GPT Image 2.5 and Qwen Image 2.1
The name comparison needs a source-quality distinction. OpenAI's current VioEvo guide documents GPT Image 2.5 as two API choices, Flare and Sunburst, with token-based output pricing. Qwen's official image repository documents Qwen-Image-2.0; it does not establish a separate official Qwen Image 2.1 API or price in the sources checked for this guide. I therefore do not invent a Qwen 2.1 specification.
| Model name | Confirmed identity | Useful comparison question |
|---|---|---|
| Nano Banana 2.1 | Google's gemini-nano-banana-2.1; 1K, 2K, and 4K; up to 14 reference images | Does the current Flash-family route meet the quality and consistency bar at the required cost? |
| GPT Image 2.5 | OpenAI's Flare and Sunburst API models; see ChatGPT Images 2.5 | Does Flare's speed or Sunburst's quality produce more accepted images on the same edit and text tests? |
| Grok Imagine Image 2.0 | xAI's current VioEvo image route; see Grok Imagine Image 2.0 | Does its image-editing workflow preserve the reference subject and required copy on the same test set? |
| Qwen Image 2.1 | No separate official Qwen Image 2.1 model page or price was confirmed; Qwen's official repository documents Qwen-Image-2.0 | Is the search result referring to Qwen-Image-2.0, a provider alias, or an unreleased/third-party label? Verify before comparing. |
Sources: Google's Nano Banana 2.1 model page, OpenAI's image generation guide, and the Qwen-Image official repository. The comparison questions are editorial evaluation criteria.
For a fair test, use identical prompts and references, then score realism, text rendering, instruction following, consistency, latency, and accepted-image cost. Do not treat different model names as evidence of a quality ranking.
From a Nano Banana 2.1 Image to Video
The practical workflow is to generate a clean still first, review the subject and composition, then send that approved image to VioEvo's image-to-video tool. Describe only the motion you want: camera movement, subject action, timing, and any audio or atmosphere the selected video model supports.
- Generate a 2.1 still with the subject, framing, lighting, and negative constraints in the prompt.
- Use 2K or 4K when the still must survive a crop or a detailed first frame; review text and identity before animating.
- Open Image to Video, upload the approved still, and choose a video model suited to the motion and duration.
- Prompt the movement separately: “slow push-in, subject keeps the same face and wardrobe, subtle fabric movement, stable background, no new objects.”
- Review the first and last seconds for identity drift, unwanted camera motion, and text deformation before export.
FAQ
Which is better, nano banana or nano banana 2?
Nano Banana is the family label, while Nano Banana 2 is the specific Gemini 3.1 Flash Image generation. For a current API or production workflow, use a versioned model and compare it against your own realism, text, instruction, consistency, latency, and cost rubric. Nano Banana 2.1 is the newer documented update to that Flash route.
What's the best version of nano banana?
Nano Banana 2.1 is the practical default when you need 1K, 2K, or 4K output, reference-image fusion, improved text rendering, and Flash-level efficiency. Pro may be the better fit when your controlled tests show that its capability advantage justifies its cost and latency; Lite is the better fit for fast 1K volume.
Is nano banana 2.1 better than pro?
Not universally. Google positions Pro as the higher-capability tier and 2.1 as the high-efficiency route. Test the same prompts, references, resolution, thinking setting, and acceptance rubric, then compare cost per accepted image rather than relying on the names.
Does Nano Banana 2.1 have a free API tier?
Google's current pricing table lists no free-tier image price for the model. It does list 5,000 free Web and Image Search grounding requests per month shared across Gemini 3.x models. VioEvo may grant eligible new accounts 150 welcome credits under its own policy.
What is the Nano Banana 2.1 API model ID?
Use gemini-nano-banana-2.1 in the Gemini API. The VioEvo selector is nano-banana-2.1, which is a separate product identifier.
Sources: Google Gemini API model page, Google image-generation guide, Google pricing, and Google DeepMind Nano Banana.