Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
28 commits
Select commit Hold shift + click to select a range
3e6b145
docs: align configuration reference with the config schema and curren…
thushan Jun 9, 2026
82431c1
docs: correct API reference endpoints, headers and response shapes ag…
thushan Jun 9, 2026
5c78416
docs: verify concepts against engine, balancer, sticky, health and tr…
thushan Jun 9, 2026
bc2be7c
docs: validate integration configs against the schema and tool sources
thushan Jun 9, 2026
ddb2adb
docs: fix development docs to match real structs, make targets and ci…
thushan Jun 9, 2026
36e6fe4
docs: correct comparison pages and fix invalid config snippets
thushan Jun 9, 2026
e99ec08
docs: correct getting-started and top-level pages against the code
thushan Jun 9, 2026
85eb7f2
docs: fix notes and troubleshooting (proxy paths, logging env vars, i…
thushan Jun 9, 2026
ed2077f
docs: rework readme with supported-backends rundown and accurate claims
thushan Jun 9, 2026
f088269
docs: align CLAUDE.md backend list, headers and config facts with the…
thushan Jun 9, 2026
d1f8df5
docs: trim CLAUDE.md to the essentials
thushan Jun 9, 2026
f02f43b
Revert "docs: trim CLAUDE.md to the essentials"
thushan Jun 9, 2026
6a322be
docs: list /olla/openai in the readme endpoints table
thushan Jun 9, 2026
0efdd25
docs: clarify which circuit breaker treats backend 5xx as a failure
thushan Jun 9, 2026
601d3bc
docs: address coderabbit feedback on crush and opencode integration docs
thushan Jun 9, 2026
90f87f3
docs: correct circuit-breaker 5xx behaviour and drop fabricated consu…
thushan Jun 9, 2026
faa3120
docs: fix crush models requirement, opencode guidance and api-transla…
thushan Jun 9, 2026
0221674
docs: add missing routes (/version, sglang version) and tidy api-refe…
thushan Jun 9, 2026
04a08c2
docs: correct struct source paths and make-target descriptions in dev…
thushan Jun 9, 2026
52a34a4
docs: configuration house-style cleanup (no em-dashes)
thushan Jun 9, 2026
b1c2507
docs: compare, getting-started and troubleshooting house-style cleanup
thushan Jun 9, 2026
8c491b3
docs: remove em-dashes from the readme
thushan Jun 9, 2026
46bae43
examples: fix crush config to the real catwalk schema (models list, o…
thushan Jun 9, 2026
fd1bcad
examples: fix opencode config (models map, correct path, no --provide…
thushan Jun 9, 2026
5101483
examples: replace fictional claude code config with real settings.jso…
thushan Jun 9, 2026
9ff8af2
examples: correct open webui env vars and drop deprecated olla keys
thushan Jun 9, 2026
10a3657
ignore examples too
thushan Jun 10, 2026
e4d562a
ignore examples too
thushan Jun 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ on:
paths-ignore:
- '**.md'
- 'docs/**'
- 'examples/**'
- 'LICENSE'
- '.gitignore'
- '.dockerignore'
Expand All @@ -17,6 +18,7 @@ on:
paths-ignore:
- '**.md'
- 'docs/**'
- 'examples/**'
- 'LICENSE'
- '.gitignore'
- '.dockerignore'
Expand Down
25 changes: 21 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# CLAUDE.md

## Overview
Olla is a high-performance proxy and load balancer for LLM infrastructure, written in Go. It intelligently routes requests across local and remote inference nodes (Ollama, LM Studio, LiteLLM, vLLM, SGLang, Llamacpp, Lemonade, Anthropic, and OpenAI-compatible endpoints).
Olla is a high-performance proxy and load balancer for LLM infrastructure, written in Go. It intelligently routes requests across local inference nodes (Ollama, LM Studio, LiteLLM, vLLM, vLLM-MLX, SGLang, llama.cpp, Lemonade, LMDeploy, Docker Model Runner, and OpenAI-compatible endpoints). The Anthropic Messages API is also supported via passthrough or translation.

The project provides two proxy engines: Sherpa (simple, maintainable) and Olla (high-performance with advanced features).

Expand Down Expand Up @@ -40,8 +40,11 @@ olla/
│ │ ├── lmstudio.yaml # LM Studio configuration
│ │ ├── lemonade.yaml # Lemonade SDK configuration
│ │ ├── litellm.yaml # LiteLLM gateway configuration
│ │ ├── lmdeploy.yaml # LMDeploy configuration
│ │ ├── vllm.yaml # vLLM configuration
│ │ ├── vllm-mlx.yaml # vLLM-MLX (Apple Silicon) configuration
│ │ ├── sglang.yaml # SGLang configuration
│ │ ├── dmr.yaml # Docker Model Runner configuration
│ │ └── openai-compatible.yaml # OpenAI-compatible generic profile (type: "openai" is an alias)
│ ├── models.yaml # Model configurations
│ └── config.local.yaml # Local configuration overrides (user, not committed to git)
Expand Down Expand Up @@ -147,6 +150,20 @@ olla/
- `/olla/proxy/` - Olla API proxy endpoint (POST)
- `/olla/proxy/v1/models` - OpenAI-compatible models listing (GET)

### Provider Proxy Prefixes
Profile-driven per-backend namespaces (see `internal/app/handlers/server_routes.go`):
- `/olla/ollama/`
- `/olla/lmstudio/` (also `/olla/lm-studio/`, `/olla/lm_studio/`)
- `/olla/vllm/`
- `/olla/vllm-mlx/`
- `/olla/sglang/`
- `/olla/lmdeploy/`
- `/olla/llamacpp/`
- `/olla/lemonade/`
- `/olla/litellm/`
- `/olla/dmr/`
- `/olla/openai/` and `/olla/openai-compatible/` (both served by `openai-compatible.yaml`)

### Translator Endpoints
Dynamically registered based on configured translators (e.g., Anthropic Messages API)

Expand All @@ -157,7 +174,7 @@ Dynamically registered based on configured translators (e.g., Anthropic Messages
## Response Headers
- `X-Olla-Endpoint`: Backend name
- `X-Olla-Model`: Model used
- `X-Olla-Backend-Type`: ollama/openai/openai-compatible/lm-studio/vllm/sglang/llamacpp/lemonade
- `X-Olla-Backend-Type`: ollama, lm-studio, litellm, vllm, vllm-mlx, sglang, llamacpp, lmdeploy, lemonade, openai, openai-compatible, docker-model-runner, omlx
- `X-Olla-Request-ID`: Request ID
- `X-Olla-Response-Time`: Total processing time
- `X-Olla-Mode`: Translator mode used (`passthrough` or absent for translation) - set on Anthropic translator requests
Expand Down Expand Up @@ -210,7 +227,7 @@ Always run `make ready` before committing changes.
- **Translator Layer**: Enables API format translation (e.g., OpenAI ↔ Anthropic) with passthrough optimisation for backends with native support
- **Passthrough Mode**: When a backend natively supports the Anthropic Messages API (vLLM, llama.cpp, LM Studio, Ollama), requests bypass translation entirely
- **Translator Metrics**: Thread-safe per-translator statistics tracking passthrough/translation rates, fallback reasons, latency, and streaming breakdown (`internal/adapter/stats/translator_collector.go`)
- **Sticky Sessions**: Optional decorator on the endpoint selector that pins multi-turn LLM conversations to the backend that handled the first turn, maximising KV-cache reuse. FNV-64a hashed keys, TTL + LRU bounded, purged on routable→non-routable health transitions (`internal/adapter/balancer/sticky.go`)
- **Sticky Sessions**: Optional decorator on the endpoint selector that pins multi-turn LLM conversations to the backend that handled the first turn, maximising KV-cache reuse. 64-bit FNV-1a hashed keys, TTL + LRU bounded, purged on routable→non-routable health transitions (`internal/adapter/balancer/sticky.go`)
- **Proxy Engines**: Choose Sherpa (simple) or Olla (high-performance)
- **Load Balancing**: Priority-based recommended for production
- **Version Management**: Build-time version injection via `internal/version`
Expand Down Expand Up @@ -259,4 +276,4 @@ CRITICAL: Always delegate tasks to the appropriate subagent. Do NOT perform work
- Research/exploration → Use the explore subagent
- Testing → Use the test subagent

Only use the main context for orchestration and task decomposition.
Only use the main context for orchestration and task decomposition.
17 changes: 10 additions & 7 deletions docs/content/api-reference/anthropic.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ The Anthropic translator accepts requests in Anthropic Messages API format at `/
**Key Features**:

- ✅ Full Anthropic Messages API compatibility
- ✅ **Passthrough mode** for backends with native Anthropic support (vLLM, llama.cpp, LM Studio, Ollama)
- ✅ **Passthrough mode** for backends with native Anthropic support (vLLM, vLLM-MLX, llama.cpp, LM Studio, Ollama, oMLX, Docker Model Runner)
- ✅ **Translation mode** for OpenAI-compatible backends without native support
- ✅ Automatic fallback from passthrough to translation when needed
- ✅ Streaming via Server-Sent Events (SSE)
Expand Down Expand Up @@ -58,7 +58,7 @@ sequenceDiagram
Olla->>Client: Response (Anthropic format - unchanged)
```

**Compatible backends**: vLLM (v0.11.1+), llama.cpp (b4847+), LM Studio (v0.4.1+), Ollama (v0.14.0+)
**Compatible backends**: vLLM (v0.11.1+), vLLM-MLX (recent), llama.cpp (b4847+), LM Studio (v0.4.1+), Ollama (v0.14.0+), oMLX (v0.4.2+), Docker Model Runner (recent)

**Observability**: Responses include `X-Olla-Mode: passthrough` header.

Expand Down Expand Up @@ -669,11 +669,11 @@ Standard Olla rate limits apply:
Configure in `config.yaml`:

```yaml
security:
rate_limit:
enabled: true
requests_per_minute: 100
burst: 50
server:
rate_limits:
global_requests_per_minute: 1000
per_ip_requests_per_minute: 100
burst_size: 50
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.


Expand Down Expand Up @@ -792,9 +792,12 @@ When `passthrough_enabled` is `true` (the default), Olla forwards requests direc
| Backend | Profile | Min Version | Notes |
|---------|---------|-------------|-------|
| vLLM | `config/profiles/vllm.yaml` | v0.11.1+ | No token counting |
| vLLM-MLX | `config/profiles/vllm-mlx.yaml` | recent | Supports token counting |
| llama.cpp | `config/profiles/llamacpp.yaml` | b4847+ | Supports token counting |
| LM Studio | `config/profiles/lmstudio.yaml` | v0.4.1+ | No token counting |
| Ollama | `config/profiles/ollama.yaml` | v0.14.0+ | No token counting |
| oMLX | `config/profiles/omlx.yaml` | v0.4.2+ | No token counting (native endpoint exists but Olla uses local estimator) |
| Docker Model Runner | `config/profiles/dmr.yaml` | recent | Supports token counting |

To disable passthrough for a specific backend, set `anthropic_support.enabled: false` in the profile:

Expand Down
22 changes: 12 additions & 10 deletions docs/content/api-reference/docker-model-runner.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ The `/engines/v1/...` paths use automatic engine selection. The explicit `/engin

List models available on the Docker Model Runner instance.

Returns an empty `data` array when no models have been loaded yet — this is normal behaviour due to lazy model loading and does not indicate an unhealthy endpoint.
Returns an empty `data` array when no models have been loaded yet. This is normal behaviour due to lazy model loading and does not indicate an unhealthy endpoint.

### Request

Expand Down Expand Up @@ -279,15 +279,17 @@ All responses include:
## Configuration Example

```yaml
endpoints:
- url: "http://localhost:12434"
name: "local-dmr"
type: "docker-model-runner"
priority: 95
model_url: "/engines/v1/models"
health_check_url: "/engines/v1/models"
check_interval: 10s
check_timeout: 5s
discovery:
static:
endpoints:
- url: "http://localhost:12434"
name: "local-dmr"
type: "docker-model-runner"
priority: 95
model_url: "/engines/v1/models"
health_check_url: "/engines/v1/models"
check_interval: 10s
check_timeout: 5s
```

See the [Docker Model Runner Integration Guide](../integrations/backend/docker-model-runner.md) for full configuration and setup instructions.
9 changes: 1 addition & 8 deletions docs/content/api-reference/lemonade.md
Original file line number Diff line number Diff line change
Expand Up @@ -361,14 +361,7 @@ curl http://localhost:40114/internal/health
**Response:**
```json
{
"status": "healthy",
"endpoints": {
"lemonade-npu": {
"url": "http://npu-server:8000",
"healthy": true,
"last_check": "2025-01-15T10:30:00Z"
}
}
"status": "healthy"
}
```

Expand Down
58 changes: 31 additions & 27 deletions docs/content/api-reference/llamacpp.md
Original file line number Diff line number Diff line change
Expand Up @@ -537,7 +537,7 @@ All responses from Olla include these headers:
- `X-Olla-Backend-Type: llamacpp` - Identifies the backend type
- `X-Olla-Endpoint: <name>` - Backend endpoint name (e.g., "llamacpp-server")
- `X-Olla-Model: <model>` - GGUF model used for the request
- `X-Olla-Request-Id: <id>` - Unique request identifier
- `X-Olla-Request-ID: <id>` - Unique request identifier
- `X-Olla-Response-Time: <ms>` - Total processing time in milliseconds
- `Via: 1.1 olla/<version>` - Olla proxy version

Expand All @@ -546,39 +546,43 @@ All responses from Olla include these headers:
## Configuration Example

```yaml
endpoints:
- url: "http://192.168.0.100:8080"
name: "llamacpp-llama-8b"
type: "llamacpp"
priority: 95
# Profile handles health checks and model discovery
headers:
X-API-Key: "${LLAMACPP_API_KEY}"
discovery:
static:
endpoints:
- url: "http://192.168.0.100:8080"
name: "llamacpp-llama-8b"
type: "llamacpp"
priority: 95
# Profile handles health checks and model discovery
headers:
X-API-Key: "${LLAMACPP_API_KEY}"
```

### Multi-Instance Setup

llama.cpp serves one model per instance. For multiple models, run multiple instances:

```yaml
endpoints:
# Instance 1: Chat model
- url: "http://192.168.0.100:8080"
name: "llamacpp-chat"
type: "llamacpp"
priority: 90

# Instance 2: Code model
- url: "http://192.168.0.101:8080"
name: "llamacpp-code"
type: "llamacpp"
priority: 85

# Instance 3: Embedding model
- url: "http://192.168.0.102:8080"
name: "llamacpp-embed"
type: "llamacpp"
priority: 80
discovery:
static:
endpoints:
# Instance 1: Chat model
- url: "http://192.168.0.100:8080"
name: "llamacpp-chat"
type: "llamacpp"
priority: 90

# Instance 2: Code model
- url: "http://192.168.0.101:8080"
name: "llamacpp-code"
type: "llamacpp"
priority: 85

# Instance 3: Embedding model
- url: "http://192.168.0.102:8080"
name: "llamacpp-embed"
type: "llamacpp"
priority: 80
```

Olla will automatically route requests to the appropriate instance based on the model name.
Expand Down
4 changes: 2 additions & 2 deletions docs/content/api-reference/lmdeploy.md
Original file line number Diff line number Diff line change
Expand Up @@ -232,7 +232,7 @@ curl -X POST http://localhost:40114/olla/lmdeploy/generate \

## POST /olla/lmdeploy/pooling

Reward or score pooling for embedding-style tasks. This is the correct path for pooling operations — `/v1/embeddings` is not supported.
Reward or score pooling for embedding-style tasks. This is the correct path for pooling operations; `/v1/embeddings` is not supported.

### Request

Expand Down Expand Up @@ -269,7 +269,7 @@ curl -X POST http://localhost:40114/olla/lmdeploy/pooling \

## GET /olla/lmdeploy/is_sleeping

Probe whether the LMDeploy engine is in sleep mode. Sleeping instances return HTTP 503 on generation endpoints — Olla's health checker treats this as a transient failure rather than a hard outage.
Probe whether the LMDeploy engine is in sleep mode. Sleeping instances return HTTP 503 on generation endpoints, which Olla's health checker treats as a transient failure rather than a hard outage.

### Request

Expand Down
5 changes: 3 additions & 2 deletions docs/content/api-reference/lmstudio.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,11 +6,12 @@ Proxy endpoints for LM Studio servers. Available through multiple prefixes: `/ol

| Method | URI | Description |
|--------|-----|-------------|
| GET | `/olla/lmstudio/v1/models` | List available models |
| GET | `/olla/lmstudio/v1/models` | List available models (OpenAI format) |
| GET | `/olla/lmstudio/api/v1/models` | List available models (OpenAI format, alternate path) |
| GET | `/olla/lmstudio/api/v0/models` | List available models (LM Studio enhanced format) |
| POST | `/olla/lmstudio/v1/chat/completions` | Chat completion |
| POST | `/olla/lmstudio/v1/completions` | Text completion |
| POST | `/olla/lmstudio/v1/embeddings` | Generate embeddings |
| GET | `/olla/lmstudio/api/v0/models` | Legacy models endpoint |

## Alternative Prefixes

Expand Down
30 changes: 23 additions & 7 deletions docs/content/api-reference/models.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,9 +24,13 @@ Returns all available models across all configured and healthy endpoints.

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `format` | string | `unified` | Response format (unified/openai/ollama/lmstudio/vllm) |
| `provider` | string | all | Filter by provider type |
| `capability` | string | all | Filter by capability (chat/completion/embedding/vision) |
| `format` | string | `unified` | Response format (`unified`, `openai`, `ollama`, `lmstudio`, `vllm`) |
| `endpoint` | string | all | Filter to models served by a specific endpoint name (e.g. `local-ollama`) |
| `family` | string | all | Filter by model family (e.g. `llama`, `qwen`, `mistral`) |
| `type` | string | all | Filter by model type (`llm`, `vlm`, `embeddings`) |
| `capability` | string | all | Legacy alias for `type` (`vision`/`multimodal` → `vlm`, `embedding`/`embeddings`/`vector_search` → `embeddings`, `chat`/`text_generation`/`completion` → `llm`) |
| `available` | bool | unset | When `true`, return only models reported available right now; when `false`, return only currently unavailable models |
| `include_unavailable` | bool | `false` | When `true`, include models from endpoints that aren't currently healthy. The response then carries availability metadata per model |

### Request Examples

Expand All @@ -40,14 +44,26 @@ curl -X GET http://localhost:40114/olla/models
curl -X GET http://localhost:40114/olla/models?format=openai
```

#### Filter by Provider
#### Filter by Endpoint
```bash
curl -X GET http://localhost:40114/olla/models?provider=ollama
curl -X GET 'http://localhost:40114/olla/models?endpoint=local-ollama'
```

#### Filter by Capability
#### Filter by Family or Type
```bash
curl -X GET http://localhost:40114/olla/models?capability=chat
curl -X GET 'http://localhost:40114/olla/models?family=llama'
curl -X GET 'http://localhost:40114/olla/models?type=vlm'
```

#### Filter by Capability (legacy)
```bash
curl -X GET 'http://localhost:40114/olla/models?capability=vision' # → type=vlm
curl -X GET 'http://localhost:40114/olla/models?capability=chat' # → type=llm
```

#### Include Unhealthy Endpoints
```bash
curl -X GET 'http://localhost:40114/olla/models?include_unavailable=true'
```

### Response Formats
Expand Down
Loading
Loading