Image Model Guide · Checked October 10, 2026
Qwen-Image 2.1 Turbo: ComfyUI, LoRA and agent setup
Build an eight-step image workflow, then connect it to self-hosted Hermes Agent or OpenClaw with the appropriate tools.
What is Qwen-Image 2.1 Turbo?
Qwen's official model card describes an accelerated 7B image-generation checkpoint supporting text-to-image and image editing. The recommended path uses eight denoising steps, CFG=1 and prefix KV caching. Keep your agent's reasoning/chat model separate: Hermes or OpenClaw prepares the request, and an image backend renders it.
The October 10 article by JimmyMo prompted this guide. It distinguishes the official Turbo release from earlier third-party acceleration LoRAs. The setup below follows the linked primary repositories rather than assuming those adapters are interchangeable.
Choose the full Turbo model or the extracted LoRA
| Route | Load | Use when |
|---|---|---|
| Diffusers | Qwen/Qwen-Image-2.1-Turbo | You want a Python pipeline with the saved schedule. |
| ComfyUI full Turbo | qwen_image_2.1_turbo_bf16.safetensors or the provided INT8 variant | You want the full Turbo diffusion model. |
| ComfyUI base + LoRA | Qwen-Image 2.1 base + qwen_image_2.1_turbo_lora_avg_rank_178_bf16.safetensors | You already have the base files and want the extracted adapter. |
The Comfy-Org repository documents both ComfyUI routes. Its LoRA listing contains a roughly 913 MB adapter uploaded by Kijai. That download still requires the base diffusion model, text encoder and VAE. Do not apply the extracted Turbo LoRA on top of the full Turbo model as the default setup.
ComfyUI setup: eight steps with the correct sigmas
- Update ComfyUI and any custom nodes used by your graph. Start from the official text-to-image template or image-editing template. These are base-model templates; adapt their sampler for Turbo.
- Download matching files from Comfy-Org. Put diffusion weights in
ComfyUI/models/diffusion_models/, theqwen3vl_8bencoder inmodels/text_encoders/, andqwen_image_2.1_vae_bf16.safetensorsinmodels/vae/. - For the LoRA route, keep the base diffusion model and put the adapter in
models/loras/. Insert a model-only LoRA loader before sampling. For the full-model route, select the Turbo diffusion weights instead. - Use a custom-sigma sampler path with Euler and CFG=1. The article uses KJNodes' custom sigmas; update that extension if the node is missing. Replace the base template's ordinary KSampler path, retaining its model, conditioning and latent connections.
- Supply the following schedule, including the terminal zero. It defines eight denoising intervals. Queue one fixed-seed draft and inspect the result before adding editing references or batch jobs.
1.0, 0.978453, 0.954180, 0.926626, 0.895080, 0.845148, 0.704534, 0.414568, 0.0The eight nonzero values are recorded in the official pipeline configuration; the ComfyUI custom-sigma list ends at zero. Reducing an ordinary sampler's step count alone does not reproduce this schedule. For editing, use the edit template's reference-image path and preserve its preprocessing.
Python / Diffusers quick start
Use a CUDA-compatible PyTorch environment. The release needs a recent Diffusers source build with pipeline-configured sigma support; see the supporting Diffusers change.
pip install git+https://github.com/huggingface/diffusers.git
pip install "transformers>=5.17.0" accelerate pillow
# Save as qwen_turbo.py and run: python qwen_turbo.py
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1-Turbo", dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="A studio photograph of a blue ceramic cup on a white table.",
width=2048,
height=2048,
use_kv_cache=True,
generator=torch.Generator("cpu").manual_seed(42),
).images[0]
image.save("qwen-turbo.png")The checkpoint loads its saved schedule automatically. The official pipeline does not override that schedule merely because you set num_inference_steps. Image editing uses the same pipeline with an image input; follow the model card's example for preprocessing and reference handling.
Self-hosted Hermes Agent: call your ComfyUI workflow
The checked Hermes image tool dispatches through provider plugins or its fal route. A Hugging Face checkpoint name does not create a ComfyUI provider. Use the terminal tool with your own workflow API client, or install a custom integration you have tested.
For a ComfyUI server on the same machine, the upstream Hermes configuration example supports this terminal backend. Merge it into your existing configuration and enable terminal access with hermes tools:
# ~/.hermes/config.yaml — self-hosted example
terminal:
backend: local
cwd: /absolute/path/to/your/qwen-workflowsExport your tested graph with Export (API) as qwen-turbo-api.json. Its API node IDs can differ from the visual template. Make request.json containing {"prompt": YOUR_EXPORTED_GRAPH}, after changing the actual prompt input in that graph. Then have your API client submit it once:
curl --fail-with-body http://127.0.0.1:8188/prompt \
-H 'Content-Type: application/json' \
--data-binary @request.json
# Replace PROMPT_ID with the ID returned above; poll this request.
curl --fail-with-body http://127.0.0.1:8188/history/PROMPT_IDComfyUI's official API client example shows waiting for completion and downloading output via /view. Give Hermes the tested client and ask it to return the saved image. Inspect job errors before resubmitting. Inside a container, 127.0.0.1 refers to that container; use an authenticated, reachable GPU endpoint instead.
Self-hosted OpenClaw: use the ComfyUI provider
Follow the upstream Comfy provider documentation. Newer versions install it with openclaw plugins install @openclaw/comfy-provider; older releases may bundle it. Check your installed version, then merge this local example into ~/.openclaw/openclaw.json:
{
"plugins": {
"entries": {
"comfy": {
"enabled": true,
"config": {
"mode": "local",
"baseUrl": "http://127.0.0.1:8188",
"image": {
"workflowPath": "/absolute/path/qwen-turbo-api.json",
"promptNodeId": "REPLACE_WITH_YOUR_PROMPT_NODE_ID",
"promptInputName": "prompt",
"outputNodeId": "REPLACE_WITH_YOUR_SAVE_IMAGE_NODE_ID"
}
}
}
}
},
"agents": {
"defaults": {
"mediaModels": {
"image": { "primary": "comfy/workflow" }
}
}
}
}Replace both node placeholders with IDs from your exported API graph. This example targets a literal prompt input, such as an unconnected prompt on TextEncodeQwenImage21. If the template routes its prompt through an enhancer or string node, target that upstream input and set promptInputName to its actual name. Dimensions and sampler settings belong to the graph. For editing, also configure inputImageNodeId and inputImageInputName for the reference loader; the documented provider upload route accepts one reference image. Keep existing chat-model and channel configuration.
What is available on OpenClaw Launch?
The API Keys page has a ComfyUI credential path for OpenClaw. The platform syncs the credential, while the user still supplies a compatible cloud workflow and prompt node. Saving a key does not publish your local Turbo graph to the cloud or confirm that the cloud has its weights and custom nodes.
Hermes differs here: the current managed image parser explicitly rejects comfy/workflow. There is no verified one-click Qwen-Image 2.1 Turbo choice for hosted Hermes in this guide. Use a separate tested GPU service for a custom self-hosted integration, or choose one of the image models actually listed in your managed dashboard. Do not install diffusion weights on the managed agent server.
Hardware, costs and licensing
A small agent server can orchestrate requests, but image inference needs separate GPU resources. VRAM use depends on precision, resolution, encoder residency and offloading; this guide does not establish a minimum VRAM or a fivefold speed guarantee. Start with one draft and measure peak memory and elapsed time on your hardware.
No universal per-image API price is established by the checkpoint release. Local inference costs GPU time; a cloud provider sets its own charges. The Qwen Research License grants non-commercial research/evaluation use and requires a separate license for commercial use. Check those terms before planning production product images or a paid generation service.
Troubleshooting
- Missing node or model: update the relevant ComfyUI components and match every filename to the graph. The LoRA alone is insufficient.
- Poor eight-step output: check the full sigma list, CFG, base-versus-Turbo choice and LoRA wiring before changing the prompt.
- Out of memory: reduce resolution or batch size and inspect the model's documented quantized/offload options.
- Workflow rejected by the API: export the executable API graph, not the UI canvas JSON. Check the server's returned node errors.
- Agent cannot reach ComfyUI: verify the endpoint from the agent's runtime and check authentication. A chat-model selection cannot repair a missing image backend.
Qwen-Image 2.1 Turbo FAQ
Does Qwen-Image 2.1 Turbo work with Hermes Agent?
A self-hosted Hermes agent can submit an exported workflow to your own ComfyUI server through its terminal tool or a custom integration. The checked OpenClaw Launch Hermes image parser rejects comfy/workflow; this guide does not claim a built-in hosted Hermes Turbo option.
Can OpenClaw use this model?
OpenClaw has a ComfyUI workflow provider. Configure a working Turbo API-format graph and its prompt/output node IDs, then use comfy/workflow as the image route. The checkpoint name is not your primary chat-model ID.
Is the Turbo LoRA a complete model?
No. The extracted LoRA still needs the Qwen-Image 2.1 base diffusion model, text encoder and VAE. Alternatively, use the full Turbo diffusion model without also applying the Turbo LoRA.
Is Qwen-Image 2.1 Turbo free for commercial use?
The published weights use the Qwen Research License, which permits non-commercial research and evaluation. Commercial use requires a separate license. GPU hosting and any third-party API billing are separate costs.