tool

llama.cpp

Runs quantized language models efficiently across CPU, Metal, CUDA, Vulkan, and other hardware backends.

What it provides

LLM inferenceGGUFQuantization

Connect an HTTPS inference API. Model weights stay on CPU/GPU infrastructure and are never installed inside the 4 GB bot.

Connection requirements

A llama-server deployment with HTTPS

Protocol
OpenAI-compatible REST
Endpoint
https://llama-cpp.example.com
Authentication
Choose the mode required by your service
Open official setup guide ↗

Where to add it

Open Dashboard Tools, select a running OpenClaw or Hermes instance, then use the card’s Install or Connect action. External services may require an HTTPS endpoint and credentials; the dashboard shows those fields before anything is saved.

Related Models & routing