LL
toolllama.cpp
Runs quantized language models efficiently across CPU, Metal, CUDA, Vulkan, and other hardware backends.
What it provides
Connect an HTTPS inference API. Model weights stay on CPU/GPU infrastructure and are never installed inside the 4 GB bot.
Connection requirements
A llama-server deployment with HTTPS
- Protocol
- OpenAI-compatible REST
- Endpoint
https://llama-cpp.example.com- Authentication
- Choose the mode required by your service
Where to add it
Open Dashboard Tools, select a running OpenClaw or Hermes instance, then use the card’s Install or Connect action. External services may require an HTTPS endpoint and credentials; the dashboard shows those fields before anything is saved.