Guide
Hermes Agent + Gemini: Use Google's Gemini Models with Hermes
Gemini — Google DeepMind's flagship model family — is a strong choice for Hermes Agent. Gemini 3.1 Pro brings million-token context for deep research, Gemini 3.8 Flash delivers fast, low-cost responses, and the whole lineup covers multimodal vision, audio, video, and code workloads.
What Is Gemini?
Gemini is Google DeepMind's family of multimodal large language models. Built from the ground up to handle text, images, audio, and video in a single model, Gemini is known for its long-context performance and tight integration with Google's search and tool ecosystem. For agent workloads, the 1M-token window across the lineup stands out — useful for multi-document analysis and long sessions.
Hermes Agent reaches Gemini through two paths: the Google AI Studio API (via GOOGLE_API_KEY or GEMINI_API_KEY) or via OpenRouter, which routes to Gemini alongside 200+ other models on a single key.
Gemini Model Lineup for Hermes
| Model | Best For | Context | Cost (Input) |
|---|---|---|---|
| Gemini 3.1 Pro | Deep research, multi-document analysis, complex tool use | 1M tokens | ~$2/M tokens |
| Gemini 3.8 Flash | Balanced everyday agent work, coding, tool use — recommended default | 1M tokens | ~$0.75/M tokens |
| Gemini 3 Flash | Fast chat, high-volume messaging, low-cost triage | 1M tokens | ~$0.50/M tokens |
| Gemini 3.1 Flash Lite | Highest-volume bots, classification, routing | 1M tokens | ~$0.25/M tokens |
| Gemini 2.5 Flash / Flash-Lite | Previous generation — still available, cheapest options | 1M tokens | ~$0.30 / ~$0.10/M tokens |
For most Hermes deployments, Gemini 3.8 Flash is the right starting point. It handles tool calls, produces coherent multi-step responses, and costs well under a dollar per million input tokens — a fraction of frontier-model pricing. Upgrade to Gemini 3.1 Pro when you need its reasoning depth; drop to Flash Lite when volume eclipses everything else.
Option 1: Hermes Agent on OpenClaw Launch (Easiest)
The fastest path to a Gemini-powered Hermes Agent. No API key required, no Docker setup, no config file editing.
- Go to openclawlaunch.com/hermes-hosting and start a Hermes deploy.
- Select Gemini 3.8 Flash (or 3.1 Pro / 3 Flash) from the model dropdown.
- Connect your messaging channel — Telegram, Discord, WhatsApp, or others.
- Click Deploy. Your Gemini-powered Hermes Agent is live in roughly 30 seconds.
Option 2: Google AI Studio API Direct (Self-Hosted)
If you're running Hermes on your own server with a direct Google AI Studio key, set the environment variable and tell Hermes to use the google provider:
# Hermes reads GOOGLE_API_KEY or GEMINI_API_KEY
export GOOGLE_API_KEY=AIza...
# Pick the provider and default model (interactive wizard)
hermes model
# Or configure /opt/data/config.yaml directly:
# model:
# provider: google
# default: gemini-2.5-flashGenerate an API key at aistudio.google.com/apikey. The free tier covers most prototyping; paid tier billing is usage-based with no monthly minimum.
Option 3: Gemini via OpenRouter (Self-Hosted)
OpenRouter lets you reach every Gemini model with a single key, alongside Claude, GPT, DeepSeek, Grok, and 200+ others.
export OPENROUTER_API_KEY=sk-or-...
# Pick the provider and default model (interactive wizard)
hermes modelWhen to Choose Each Gemini Model
Choose Gemini 3.8 Flash as your default. It handles everyday chat, coding, tool use, and short research tasks at conversational speed for roughly $0.75 per million input tokens. For high-volume bots, this is one of the cheapest viable options in the frontier-quality tier.
Choose Gemini 3.1 Pro when you need its reasoning depth for hard tasks — long research sessions, multi-PDF analysis, complex agentic plans. The pricing is competitive with Claude Sonnet for similar quality on reasoning benchmarks.
Choose Gemini 3.1 Flash Lite (or the older 2.5 Flash-Lite) for the highest-volume scenarios: classification, routing, intent detection, or simple Q&A where responses need to be both fast and very cheap.
Switching Models at Runtime
/model google/gemini-2.5-flash
/model google/gemini-2.5-pro
/model google/gemini-3-proBYOK on OpenClaw Launch
On managed OpenClaw Launch deploys, you can use your own OpenRouter key instead of bundled AI credits, and Gemini requests then bill to your OpenRouter account. On a hosted Hermes bot, Gemini models picked in the dashboard are served through OpenRouter, not through a saved Google AI Studio key. To bill Gemini to your own Google account, use the self-hosted Google AI Studio setup above; see Hermes Agent + Google AI Studio for the full walkthrough.
What's Next?
- Hermes Agent + Claude — Use Anthropic's Claude family with Hermes
- Hermes Agent + OpenAI — Run Hermes on GPT-5.6 and the OpenAI lineup
- Hermes Agent + OpenRouter — One key for 200+ models, Gemini included
- Hermes Agent + Telegram — Connect your Gemini-powered Hermes bot to Telegram