← Home

Guide

Hermes Agent + Gemini: Use Google's Gemini Models with Hermes

Gemini — Google DeepMind's flagship model family — is a strong choice for Hermes Agent. Gemini 3.1 Pro brings million-token context for deep research, Gemini 3.8 Flash delivers fast, low-cost responses, and the whole lineup covers multimodal vision, audio, video, and code workloads.

What Is Gemini?

Gemini is Google DeepMind's family of multimodal large language models. Built from the ground up to handle text, images, audio, and video in a single model, Gemini is known for its long-context performance and tight integration with Google's search and tool ecosystem. For agent workloads, the 1M-token window across the lineup stands out — useful for multi-document analysis and long sessions.

Hermes Agent reaches Gemini through two paths: the Google AI Studio API (via GOOGLE_API_KEY or GEMINI_API_KEY) or via OpenRouter, which routes to Gemini alongside 200+ other models on a single key.

Gemini Model Lineup for Hermes

ModelBest ForContextCost (Input)
Gemini 3.1 ProDeep research, multi-document analysis, complex tool use1M tokens~$2/M tokens
Gemini 3.8 FlashBalanced everyday agent work, coding, tool use — recommended default1M tokens~$0.75/M tokens
Gemini 3 FlashFast chat, high-volume messaging, low-cost triage1M tokens~$0.50/M tokens
Gemini 3.1 Flash LiteHighest-volume bots, classification, routing1M tokens~$0.25/M tokens
Gemini 2.5 Flash / Flash-LitePrevious generation — still available, cheapest options1M tokens~$0.30 / ~$0.10/M tokens

For most Hermes deployments, Gemini 3.8 Flash is the right starting point. It handles tool calls, produces coherent multi-step responses, and costs well under a dollar per million input tokens — a fraction of frontier-model pricing. Upgrade to Gemini 3.1 Pro when you need its reasoning depth; drop to Flash Lite when volume eclipses everything else.

Option 1: Hermes Agent on OpenClaw Launch (Easiest)

The fastest path to a Gemini-powered Hermes Agent. No API key required, no Docker setup, no config file editing.

  1. Go to openclawlaunch.com/hermes-hosting and start a Hermes deploy.
  2. Select Gemini 3.8 Flash (or 3.1 Pro / 3 Flash) from the model dropdown.
  3. Connect your messaging channel — Telegram, Discord, WhatsApp, or others.
  4. Click Deploy. Your Gemini-powered Hermes Agent is live in roughly 30 seconds.
Tip: OpenClaw Launch routes Gemini requests through OpenRouter by default. AI credits are included in every paid Hermes plan — no separate Google billing required unless you bring your own key. The 30-minute free trial runs on the free trial models; Gemini needs a paid plan.

Option 2: Google AI Studio API Direct (Self-Hosted)

If you're running Hermes on your own server with a direct Google AI Studio key, set the environment variable and tell Hermes to use the google provider:

# Hermes reads GOOGLE_API_KEY or GEMINI_API_KEY
export GOOGLE_API_KEY=AIza...

# Pick the provider and default model (interactive wizard)
hermes model

# Or configure /opt/data/config.yaml directly:
# model:
#   provider: google
#   default: gemini-2.5-flash

Generate an API key at aistudio.google.com/apikey. The free tier covers most prototyping; paid tier billing is usage-based with no monthly minimum.

Option 3: Gemini via OpenRouter (Self-Hosted)

OpenRouter lets you reach every Gemini model with a single key, alongside Claude, GPT, DeepSeek, Grok, and 200+ others.

export OPENROUTER_API_KEY=sk-or-...

# Pick the provider and default model (interactive wizard)
hermes model

When to Choose Each Gemini Model

Choose Gemini 3.8 Flash as your default. It handles everyday chat, coding, tool use, and short research tasks at conversational speed for roughly $0.75 per million input tokens. For high-volume bots, this is one of the cheapest viable options in the frontier-quality tier.

Choose Gemini 3.1 Pro when you need its reasoning depth for hard tasks — long research sessions, multi-PDF analysis, complex agentic plans. The pricing is competitive with Claude Sonnet for similar quality on reasoning benchmarks.

Choose Gemini 3.1 Flash Lite (or the older 2.5 Flash-Lite) for the highest-volume scenarios: classification, routing, intent detection, or simple Q&A where responses need to be both fast and very cheap.

Switching Models at Runtime

/model google/gemini-2.5-flash
/model google/gemini-2.5-pro
/model google/gemini-3-pro

BYOK on OpenClaw Launch

On managed OpenClaw Launch deploys, you can use your own OpenRouter key instead of bundled AI credits, and Gemini requests then bill to your OpenRouter account. On a hosted Hermes bot, Gemini models picked in the dashboard are served through OpenRouter, not through a saved Google AI Studio key. To bill Gemini to your own Google account, use the self-hosted Google AI Studio setup above; see Hermes Agent + Google AI Studio for the full walkthrough.

What's Next?

Deploy Hermes with Gemini

Get a Gemini-powered Hermes Agent running in 30 seconds on OpenClaw Launch.

Deploy Hermes