← All Guides

Google Image Model Guide

Nano Banana 2.1: reference images, editing and API pricing

A practical guide to Google's new image model, with separate integration paths for Hermes Agent and OpenClaw.

Checked October 7, 2026: Google released Nano Banana 2.1 on October 6. Its stable Gemini API ID is gemini-nano-banana-2.1. This is an image generation and editing model; keep your agent's main chat model separate.

What changed, and which model should you choose?

The release notes confirm availability. Google's model card describes a Gemini 3.6 Flash base. The model page emphasises better instruction following, text rendering and consistency through successive edits. These are Google's claims, not an independent quality comparison performed for this guide.

ModelGemini API IDRole
Nano Banana 2.1gemini-nano-banana-2.1Updated efficient image model
Nano Banana 2gemini-3.1-flash-imagePrevious generation; Google recommends 2.1 for new projects
Nano Banana Progemini-3-pro-imageSeparate premium image model

2.1 defaults to 1K output and supports 2K and 4K. It accepts up to 14 reference images, with documented consistency support for up to four characters and ten objects. More references do not automatically produce a better result: use only those needed to explain the task.

Try one product edit in Google AI Studio

  1. Open Google AI Studio and check that Nano Banana 2.1 is available to your account. Confirm billing and the selected model before generating.
  2. Upload a clear source photograph. Describe the requested change and the details to preserve.
  3. Start with a 1K draft, compare it with the source, and save the approved version. Use that version as the reference for the next edit.
  4. Request a higher-resolution final only after approving the composition. Inspect labels and fine edges again at the final size.
Use the attached mug photo as the product reference.
Replace only the background with a pale blue studio backdrop.
Keep the handle shape, printed logo, camera angle and crop.
Use soft light from the left. Add no props or extra text.

This prompt is an original starting point. For multiple references, assign roles explicitly: “image 1 is the product; image 2 provides the room style.” Change one variable per revision so you can identify which instruction caused an unwanted alteration.

Local editing: check the result, not just the prompt

Google's editing documentation supports text instructions alongside an image. Describe a region in words; this example does not supply a separate mask parameter. “Preserve everything else” is an instruction, not a guarantee that every other pixel stays unchanged.

For a shop image, compare the silhouette, logo spelling, number of parts, shadows and reflections with the approved source. Test a barcode with a scanner rather than trusting its appearance. For a poster, prepare the exact copy first and check each line after generation. Google's model card still reports pose and spatial-reasoning limitations. Generated images include SynthID.

Direct Google API: a reference-image script

Use your own billed Gemini API project and key. Install or update google-genai in a dedicated Python environment, save a real PNG as product.png, and provide GEMINI_API_KEY through your environment. Do not put the key in the script or prompt.

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade google-genai
python edit_product.py

Save the following as edit_product.py. It follows Google's model-specific Interactions editing example, requests one edit and refuses to overwrite the saved result. Remove the image input and use a scene description for text-to-image.

import base64
import os
from pathlib import Path
from google import genai

source = Path("product.png")  # A real PNG reference image
output = Path("product-edited.png")
if output.exists():
    raise FileExistsError(output)
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
result = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=[
        {"type": "text", "text": "Replace only the background with pale blue. "
         "Preserve the product shape, label, camera angle and crop."},
        {"type": "image", "mime_type": "image/png",
         "data": base64.b64encode(source.read_bytes()).decode("utf-8")},
    ],
)
image = result.output_image
if image is None or not image.data:
    raise RuntimeError("No image returned; inspect the response before retrying")
with output.open("xb") as target:
    target.write(base64.b64decode(image.data))
print(output.resolve())

The response helper returns the last generated image. More complex interleaved outputs require inspecting the response steps. This example has been checked against documentation; no paid generation was performed. If your SDK lacks interactions, update it in the dedicated environment rather than changing the model ID to an older one.

Google API image-output prices

ResolutionStandardBatch
1K$0.0336$0.0168
2K$0.0504$0.0252
4K$0.113$0.0567

These are the published per-image output charges, checked October 7. The API has no free tier for 2.1. Standard input costs $1.50 per million tokens; text and thinking output costs $7.50 per million. Optional search grounding may add charges. Batch is an asynchronous route, not a switch used by the script above.

Twenty 1K standard outputs calculate to $0.672 in image-output charges; twenty 4K outputs calculate to $2.26. Add input and other billable work. A rejected draft or another edit can mean another request, so set a project budget and review drafts before making variants. Other providers set their own prices.

Hermes Agent setup, then OpenClaw setup

Self-hosted Hermes: the checked v2026.9.24 image plugin directory has no Google backend. Use the direct script above from a custom skill, with your own key in the process environment. Have the skill accept a reference path and edit instruction, invoke the script, and return the saved file path. Limit the number of calls per task and verify the script manually before automating it. Do not place this image ID in the main chat-model configuration.

Self-hosted OpenClaw: the checked v2026.5.7 Google image provider passes the requested model name to generateContent. Its compatibility with 2.1 has not been tested here, while Google's current model-specific example uses Interactions. A skill invoking the direct script is the documented workflow used in this guide; native adapter support needs a separate bounded test with your own key.

Managed hosting: the Image picker is shared, but runtime support differs. The hosted Hermes Google wrapper currently only accepts older model IDs; entering google/gemini-nano-banana-2.1 can fall back to the old default. OpenClaw accepts custom image IDs, but this alone does not verify generation, editing or billing. Use AI Studio or the direct API for 2.1 today. See the general image guide for existing hosted workflows and the Hermes / OpenClaw fal guides for a different provider route.

When a request fails

  • 403: check project billing, API permissions and regional eligibility.
  • 404: check the exact model ID and account access; an old preview name is a different model.
  • 429: check project quota and use bounded backoff. Avoid unattended retry loops.
  • No image: inspect the response and any refusal before retrying. Do not save text as an image.
  • The wrong model runs: check the actual provider request and wrapper allowlist, not just the picker label.

Sources were checked October 7, 2026. Framework findings describe the repository versions linked above and the hosted wrapper source, not a live audit of every customer instance.

Nano Banana 2.1 FAQ

Does Nano Banana 2.1 work on Hermes Agent?

Self-hosted Hermes can call your own Google API script through a skill. The checked upstream release has no Google image plugin. The hosted Google wrapper accepts older model IDs and falls back when given 2.1, so a custom ID does not enable it today.

Can I select Nano Banana 2.1 on OpenClaw?

The Google image adapter accepts model names, but its generateContent route has not been tested with 2.1 here. The model-specific example in this guide uses Google Interactions directly. An accepted custom image ID is not proof of successful generation.

Is Nano Banana 2.1 free?

The Gemini API has no free tier for this model. The listed 1K standard image-output charge is $0.0336; input, text/thinking output and optional grounding can add charges. Gemini app allowances and hosting subscriptions are separate.

Will every background edit preserve the subject pixel for pixel?

No such guarantee is established by the official documentation. Reference consistency helps guide an edit, but small details, poses, text, reflections or spatial relationships can still change. Compare every deliverable with the original.

Related guides

Explore managed Hermes Agent

Compare hosted agent tools with self-hosting. Google API calls are billed separately; the current hosted Hermes image plugin does not enable Nano Banana 2.1.

Explore Hermes hosting