llmman in two minutes
llmman is a command-line tool written in Rust with the slogan "Run any agent on any model".
llmman is a project initiated by Eric Curtin 👏
Its 3 main functions are as follows:
- It manages models like OCI images. A model is retrieved from Docker Hub, Hugging Face, ... or your private registry, and is stored locally in a standard OCI Image Layout. So, no proprietary format, we stay standard.
- The part that interests me the most: It serves these models behind a single API (
llmman serve, port17434) that knows how to speak OpenAI, Ollama and Anthropic. - It can launch code agents (Claude Code, Codex, OpenCode, Gemini, ...) pointed at a local model or a remote endpoint, in a single command.
And llmman knows how to work with llamacpp as an inference engine (otherwise known as: as a runtime).
This is convenient, because my preferred engines are llamacpp and Docker Model Runner (which itself uses
llamacpp).
So, today let's see how to use llmman with llamacpp to serve my favorite model.
llmman comes with many features, I encourage you to consult the excellent documentation.
Installation
I tested this installation on both macOS and Linux.
I used the following options:
# Official script
curl -fsSL https://llmmanorg.github.io/install.sh | sh
# or via Homebrew
brew install llmmanorg/tap/llmman
For Linux, I used the curl option, and for macOS, I used the curl option on one Mac and the brew option on another Mac.
Choosing the llamacpp runtime
There are several solutions here as well. I chose to let llmman handle it, which will download and cache the official llamacpp binary (once), and then use it at every startup. (It is also possible to specify the path to your existing llamacpp installation).
Starting up
It's simple, just run this command:
llmman serve
To verify that everything is working correctly, you can use the following commands:
curl -s http://127.0.0.1:17434/api/version
curl -s http://127.0.0.1:17434/v1/models # {"data":[],"object":"list"} if you haven't installed any model
We need a model
To retrieve models, it's also simple, and afterwards llmman will allow you to easily switch from one to another. For example, if you want to install Qwen3.5-0.8B-GGUF, the version provided by Unsloth on the Hugging Face platform, type the following command:
llmman pull hf.co/unsloth/Qwen3.5-0.8B-GGUF
Wait a little while for the download, and once it's finished, you can verify that everything is fine and that your model is indeed there with these commands:
llmman list
llmman show hf.co/unsloth/Qwen3.5-0.8B-GGUF
There, you are now ready to query your model.
Testing the three API versions (Ollama, Anthropic, OpenAI)
You can query the various APIs using curl:
# OpenAI
curl -s http://127.0.0.1:17434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"hf.co/unsloth/Qwen3.5-0.8B-GGUF","messages":[{"role":"user","content":"Explain why the hawaiian pizza is the best pizza in the world"}],"max_tokens":300,"chat_template_kwargs":{"enable_thinking":false}}'
# Anthropic
curl -s http://127.0.0.1:17434/v1/messages \
-H "Content-Type: application/json" \
-d '{"model":"hf.co/unsloth/Qwen3.5-0.8B-GGUF","max_tokens":200,"messages":[{"role":"user","content":"Explain why the hawaiian pizza is the best pizza in the world"}],"thinking":{"type":"disabled"}}'
# Ollama
curl -s http://127.0.0.1:17434/api/chat -d '{\n "model": "hf.co/unsloth/Qwen3.5-0.8B-GGUF",\n "messages": [{"role": "user", "content": "Explain why the hawaiian pizza is the best pizza in the world"}],\n "stream": false\n}'
As you can see, it's easy to implement llmman and then use it in generative AI applications or with your favorite agents.
Go Example with the OpenAI Go SDK
You can therefore use the API exposed by llmman with the usual frameworks, as here with Go and the OpenAI API:
package main
import (
"context"
"fmt"
openai "github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
ctx := context.Background()
model := "hf.co/unsloth/Qwen3.5-0.8B-GGUF"
baseURL := "http://127.0.0.1:17434/v1"
client := openai.NewClient(
option.WithBaseURL(baseURL),
option.WithAPIKey("not-needed"),
)
completion, err := client.Chat.Completions.New(
ctx,
openai.ChatCompletionNewParams{
Model: model,
Messages: []openai.ChatCompletionMessageParamUnion{
openai.SystemMessage("You are a pizza expert."),
openai.UserMessage("Tell me more about Hawaiian pizza."),
},
Temperature: openai.Float(0.0),
TopP: openai.Float(0.9),
MaxTokens: openai.Int(2048),
},
option.WithJSONSet(
"chat_template_kwargs",
map[string]any{
"enable_thinking": false,
},
),
)
if err != nil {
fmt.Println("[error:", err, "]")
return
}
if len(completion.Choices) == 0 {
fmt.Println("[error: empty completion]")
return
}
fmt.Println(completion.Choices[0].Message.Content)
}
One last thing before you go
llmman provides a particularly well-made web UI. To access it, it's simple: open this URL in your browser: http://127.0.0.1:17434/ and you will be able to search for and download models, and "chat with your models":


That's all for today. I'll leave you to play with llmman. Feel free to ask questions 🙂
Written by
Keep reading

Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙
Serve a GGUF model with llama.cpp, point Docker Agent at it with a short agent.yaml, and get a fully local code agent running in minutes.
Sep 15, 2026
Zed Editor + Docker Agent + ACP: coding with a local agent plugged into llmman
Plug Zed's agent panel into Docker Agent over ACP, then point it at llmman to chat with a fully local, OpenAI-compatible coding agent.
Oct 3, 2026
Docker Model Runner, standalone edition (`dmr`)
Install and use dmr, the standalone Docker Model Runner: one binary with the inference daemon and the full model CLI, no Docker Desktop needed.
Jul 31, 2026
No comments yet. Be the first to comment!