Skip to content

Commit b608942

Browse files
Andrey Cheptsovclaude
andcommitted
Improve preset creation harness and CLI output
A preset claims to be a verified serving configuration, but three of the properties that decide what "verified" means were optional or defaulted, so two creations with the same configuration could produce very different artifacts. - `max_ttft` and `min_context_length` are now required. Without a latency bound, maximising throughput has a degenerate optimum; without a context floor, one creation served 64K where another served 1M under identical constraints. - `concurrency` is now required rather than defaulting to 8. - `trials` in the configuration, `trials_num` in the constraints the agent reads, since it is the number of trials rather than a ceiling. - `baseline: true` makes the first trial a reference point rather than an optimization attempt. - `input_tokens`/`output_tokens` pin the benchmark workload so trials and the final service are comparable; both default to 1024. - rename the config's `context_length` to `min_context_length`, since it is a requirement, and keep `context_length` for the measured value. Per-user speed is now the steady decode rate, `1/TPOT`, as the serving literature defines it. Dividing aggregate throughput by concurrency folded in time to first token and read about 9% low. Both display paths now use the same definition; they previously disagreed. The trial record gains `learned`, required for every trial, and `failed` for a benchmark that broke a constraint. A failed trial keeps its benchmark, since that is what the next trial learns from, and is excluded from best-trial selection. `findings.md` is removed. It asked the agent to enumerate what it had not tried, which is unbounded and produced an arbitrary subset presented as complete. What a trial taught now lives on the trial record. Trials themselves are measured more honestly: - record the largest context each trial handles, found by sending real requests - run the final benchmark inside the service replica, directly against the engine, so it is comparable with the trial benchmarks - require that a benchmark not reuse the previous one's prompts, which had been inflating later trials through the engine's prefix cache - record final service attempts in `verifications.jsonl` and mirror them out, so the CLI reads the phase instead of inferring it from a spent trial budget Also fixes a real bug: session constraints were read from the agent workspace, which is deleted when a session finishes, so a verified preset could never show what it was created against. `dstack preset` output is reworked: `ps`-style filtering, CONSTRAINTS and BENCHMARK columns, and one sparkline glyph per trial. The constraints are dimmed so the measurement leads, and a trial that broke a constraint is marked with a yellow bar. Maintainer notes written as `<!--!...-->` are stripped from the rendered agent prompt. The `endpoints` to `presets` rename in #4058 deleted two docs pages without adding redirects, so `/docs/concepts/endpoints/` and `/docs/reference/cli/dstack/endpoint/` returned 404. Both now redirect to their `preset` equivalents. `shared_prefix_tokens` is documented on the concepts page. It existed only in the generated schema reference, so the property that decides the benchmark's prefix cache hit rate was invisible to anyone reading the concept. A run where no trial met the constraints showed neither hardware nor a number, because both were read from the best trial and a failed one cannot become best. Such a run now shows its fastest failed benchmark, dimmed and marked `*` so it does not read as a result where styling is absent. That is the answer such a run produced: `ttft=4.3s` against a 675ms bound is why a card is unusable. The listing always shows `prefix`, including `prefix=0%`. It decides how much of each request the engine serves from its prefix cache, so two rows are only comparable when it matches. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 752de28 commit b608942

22 files changed

Lines changed: 1255 additions & 257 deletions

File tree

‎mkdocs.yml‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -103,6 +103,8 @@ plugins:
103103
"docs/tasks.md": "docs/concepts/tasks.md"
104104
"docs/services.md": "docs/concepts/services.md"
105105
"docs/fleets.md": "docs/concepts/fleets.md"
106+
"docs/concepts/endpoints.md": "docs/concepts/presets.md"
107+
"docs/reference/cli/dstack/endpoint.md": "docs/reference/cli/dstack/preset.md"
106108
"docs/examples/llms/llama31.md": "docs/examples/inference/vllm.md"
107109
"docs/examples/llms/llama32.md": "docs/examples/inference/vllm.md"
108110
"docs/examples/llms/qwen36.md": "docs/examples/models/qwen36.md"

‎mkdocs/docs/concepts/presets.md‎

Lines changed: 105 additions & 42 deletions
Original file line numberDiff line numberDiff line change
@@ -5,7 +5,7 @@ description: Creating and reusing optimized model inference configurations
55

66
# Presets
77

8-
A preset configuration lets you use an agent to create a preset: a validated and optimized model inference configuration. Once created, the preset can be reused to deploy model inference on validated hardware without an agent.
8+
A preset configuration lets you use an agent to create a preset: a verified and optimized model inference configuration. Once created, the preset can be reused to deploy model inference on verified hardware without an agent.
99

1010
The value of presets comes from combining two fundamental features: agent-driven model inference optimization and the `dstack` [service](services.md) primitive, which can deploy model inference to any cloud, Kubernetes, or on-prem cluster.
1111

@@ -25,14 +25,26 @@ The filename must end with `.dstack.yml` (e.g. `.dstack.yml` or `preset.dstack.y
2525

2626
```yaml
2727
type: preset
28-
name: qwen25-7b
28+
name: dsv4-flash
2929

3030
# The agent picks a compatible variant of the base model
31-
base: Qwen/Qwen2.5-7B-Instruct
31+
base: deepseek-ai/DeepSeek-V4-Flash
3232

3333
# The number of benchmarked trials
34-
max_trials: 3
34+
trials: 5
3535

36+
# The requirements the preset must meet (time to first token is in milliseconds)
37+
min_context_length: 1048576
38+
max_ttft: 675
39+
40+
# The number of simultaneous requests every benchmark uses
41+
concurrency: 1
42+
43+
# The request shape every benchmark uses (defaults to 1024 and 1024)
44+
input_tokens: 10000
45+
output_tokens: 1500
46+
47+
# The environment variables the agent may pass to runs
3648
env:
3749
- HF_TOKEN
3850
```
@@ -44,19 +56,23 @@ To create the preset, pass the configuration to the `dstack preset create` comma
4456
<div class="termy">
4557

4658
```shell
47-
$ dstack preset create -f preset.dstack.yml
48-
Create the preset qwen25-7b? [y/n]: y
49-
[2026-07-15 11:32:01] Starting preset creation for Qwen/Qwen2.5-7B-Instruct. Allowed fleets: gpu-fleet.
50-
[2026-07-15 11:41:06] Prototype task qwen25-7b-a1b2c3-2 verified vLLM on an L4:24GB.
51-
[2026-07-15 11:52:06] Final service qwen25-7b-a1b2c3-3 verified with context length 32768.
52-
[2026-07-15 11:52:18] Benchmark via guidellm 0.7.1: 32/32 requests succeeded.
59+
$ dstack preset create -f preset.dstack.yml --fleet b200-fleet
60+
Create the preset dsv4-flash? [y/n]: y
61+
[2026-08-04 11:38:34] Starting preset creation for deepseek-ai/DeepSeek-V4-Flash. Allowed fleets: b200-fleet.
62+
[2026-08-04 12:31:19] Trial 3 switched from vLLM to SGLang: 319 tok/s per user, 2.2x the baseline.
63+
[2026-08-04 13:04:52] Final service dsv4-flash-c83375b4-4 verified with context length 1048576.
64+
[2026-08-04 13:12:07] Benchmark via sglang.bench_serving: 32/32 requests succeeded.
5365
```
5466

5567
</div>
5668

57-
The command executes entirely locally and uses the locally installed `claude` CLI along with `dstack`'s bundled skills. The agent uses a `dstack` task to find the best serving configuration for the available fleet offers, then submits it as a `dstack` service for a final benchmark. The validated preset is saved locally under `~/.dstack/presets`.
69+
> It's highly recommended to specify the exact hardware you want the preset to use, so that the
70+
> optimization is done against that hardware. Point `dstack preset create` to a fleet configured
71+
> correspondingly, via `fleets` inside the preset configuration or via `--fleet` in the CLI.
72+
73+
The command executes entirely locally and uses the locally installed `claude` CLI along with `dstack`'s bundled skills. The agent uses a `dstack` task to find the best serving configuration for the available fleet offers, then submits it as a `dstack` service for a final benchmark.
5874

59-
You can stop watching with <kbd>Ctrl+C</kbd> at any time. The agent keeps running, and `dstack preset logs -f` follows it again. Resume an interrupted creation with `dstack preset create --resume`:
75+
You can stop watching with `Ctrl`+`C` at any time. The agent keeps running, and `dstack preset logs -f` follows it again. Resume an interrupted creation with `dstack preset create --resume`:
6076

6177
<div class="termy">
6278

@@ -66,26 +82,40 @@ $ dstack preset create -f preset.dstack.yml --resume a1b2c3d4
6682

6783
</div>
6884

85+
When resuming, the constraints are read from the original session, not from the configuration file. Editing them and resuming has no effect. To change any of them, create a new preset.
86+
6987
To stop a creation and its runs, use `dstack preset stop`.
7088

71-
!!! info "Claude configuration"
89+
??? info "Claude configuration"
7290
By default, preset creation uses the existing `claude` login. To use an Anthropic API key instead, set:
7391

7492
```shell
7593
export DSTACK_AGENT_ANTHROPIC_API_KEY=...
7694
```
7795

78-
By default, the agent uses `claude-opus-4-8` and the default `claude` CLI effort. To override them, set:
96+
By default, the agent uses `claude-opus-4-8`. It doesn't set an effort level, so the `claude` CLI default applies. To override them, set:
7997

8098
```shell
81-
export DSTACK_AGENT_ANTHROPIC_MODEL=claude-fable-5
82-
export DSTACK_AGENT_CLAUDE_EFFORT=high
99+
export DSTACK_AGENT_ANTHROPIC_MODEL=claude-opus-5
100+
export DSTACK_AGENT_CLAUDE_EFFORT=max
83101
```
84102

85103
Supported effort levels are `low`, `medium`, `high`, `xhigh`, and `max`.
86104

105+
??? info "Presets directory"
106+
The verified presets are saved locally under `~/.dstack/presets`, and `dstack preset` reads them from there. Presets aren't stored on the server.
107+
87108
## Configuration options
88109

110+
### Fleets
111+
112+
Set `fleets` to restrict creation and reuse to specific [fleets](fleets.md). It's highly recommended to specify a fleet with exactly the hardware that you'd like the preset to use.
113+
114+
Alternatively, pass `--fleet` to `dstack preset create` or `dstack preset apply`.
115+
116+
> Profile settings such as `spot_policy`, `max_price`, and `backends` are ignored during preset
117+
> creation. Configure them on the fleet instead.
118+
89119
### Model
90120

91121
=== "Base"
@@ -104,21 +134,27 @@ To stop a creation and its runs, use `dstack preset stop`.
104134
repo: Qwen/Qwen2.5-7B-Instruct
105135
```
106136

107-
### Trials
137+
### Shared prefix
108138

109-
`max_trials` is required and sets how many benchmarked trials the agent runs before promoting the best one. Set `concurrency` to control the benchmark concurrency.
139+
By default every request is unique, so the cache hit rate is near zero. Set `shared_prefix_tokens` to control how much of each request the serving framework can serve from its prefix cache.
110140

111-
### Context length
141+
<div editor-title="preset.dstack.yml">
112142

113-
Set `context_length` to require a minimum supported context length.
143+
```yaml
144+
input_tokens: 8192
145+
output_tokens: 1024
114146
115-
### Fleets
147+
# Roughly 90% of prompt tokens can be served from cache
148+
shared_prefix_tokens: 7360
149+
```
116150

117-
Set `fleets` to restrict creation and reuse to specific [fleets](fleets.md). Placement properties such as `backends`, `max_price`, and `spot_policy` constrain both creation and reuse too.
151+
</div>
152+
153+
The `shared_prefix_tokens` value is the part of `input_tokens` that is identical across requests, such as a system prompt or conversation history, and must be less than `input_tokens`.
118154

119155
### Prompt
120156

121-
Set `prompt` to guide the agent with custom objectives, target metrics, or an experimentation approach. It accepts inline text or a file `path`.
157+
The `prompt` property is optional. Set it to guide the agent with custom objectives, target metrics, or an experimentation approach. It accepts inline text or a file `path`.
122158

123159
<div editor-title="preset.dstack.yml">
124160

@@ -129,6 +165,10 @@ prompt: |
129165

130166
</div>
131167

168+
### Baseline
169+
170+
Set `baseline: true` to make the first trial a baseline: the agent serves the model the way the chosen serving framework recommends, without tuning it for performance. Later trials are optimization attempts.
171+
132172
!!! info "Reference"
133173
The `preset` configuration supports many more options. See the [`.dstack.yml` reference](../reference/dstack.yml/preset.md).
134174

@@ -139,20 +179,23 @@ To deploy a preset as a service, pass the preset configuration and the preset ID
139179
<div class="termy">
140180

141181
```shell
142-
$ dstack preset apply -f preset.dstack.yml --id 532f3f4b
182+
$ dstack preset apply -f preset.dstack.yml --id c83375b4
143183
Project main
144184
User admin
145185
Type service
146-
Resources cpu=2.. mem=8GB.. disk=100GB.. gpu=RTXPRO4500:32GB:1..
186+
Resources cpu=8.. mem=64GB.. disk=500GB gpu=B200:180GB:2
147187
Spot policy on-demand
148188
Max price off
149-
Model Qwen/Qwen3.5-27B (base)
150-
Preset 532f3f4b (ctx=8K con=8 387 tok/s TTFT 582ms)
189+
Retry policy off
190+
Idle duration 5m
191+
Max duration off
192+
Model deepseek-ai/DeepSeek-V4-Flash (base)
193+
Preset c83375b4 (io=10000/1500 conc=1 tok/s/user=309 tok/s=296 ttft=213ms ctx=1M)
151194
152-
# BACKEND RESOURCES INSTANCE TYPE PRICE
153-
1 runpod (EU-RO-1) cpu=12 mem=54GB disk=100GB gpu=RTXPRO4500:32GB:1 NVIDIA RTX PRO 4500 Blackwell $0.74
195+
# BACKEND RESOURCES INSTANCE TYPE PRICE
196+
1 runpod (US-CA-2) cpu=48 mem=502GB disk=500GB gpu=B200:180GB:2 NVIDIA B200 $11.78
154197
155-
Submit the run qwen35-27b? [y/n]: y
198+
Submit the run dsv4-flash? [y/n]: y
156199
```
157200

158201
</div>
@@ -167,20 +210,32 @@ Use `dstack preset` to list presets:
167210

168211
```shell
169212
$ dstack preset list
170-
BASE ID GPU BENCHMARK STATUS SUBMITTED NAME
171-
Qwen/Qwen2.5-0.5B
172-
bc592b38 clauding (0/3) 23 sec ago qwen05
173-
Qwen/Qwen3-32B
174-
f91d6b60 RTX5090:32GB:1 con=8 576 tok/s TTFT 368ms verified (10/10) 2 days ago qwen3-32b
175-
Qwen/Qwen3.5-27B
176-
3c4d5e6f verifying (3/3) 2 min ago qwen35-27b-2
177-
532f3f4b RTXPRO4500:32GB:1.. con=8 387 tok/s TTFT 582ms verified (4/4) yesterday qwen35-27b
178-
d1c2e12b RTX5090:32GB:1 con=8 266 tok/s TTFT 2.15s verified (7/10) yesterday
213+
ID BASE GPU CONSTRAINTS BENCHMARK STATUS SUBMITTED
214+
c83375b4 deepseek-ai/DeepSeek-V4-Flash B200:180GB:2 io=10000/1500 conc=1 tok/s/user=309 ttft=213ms ctx=1M ▂▁██▇ trialing (5/5) 2 min ago
179215
```
180216

181217
</div>
182218

183-
Presets are grouped by base model. In-progress creations appear too, with a live status like `clauding` or `verifying`. Pass `-w` to watch in realtime. Pass `-v` to include validation resources and all benchmark metrics, or `--json` for complete preset objects. Filter with `--base` or `--repo`.
219+
By default, `dstack preset` shows creations that are still running, or the most recent one if none are. Pass `-a` to show every preset, or `-n` to show the last N:
220+
221+
<div class="termy">
222+
223+
```shell
224+
$ dstack preset list -a
225+
ID BASE GPU CONSTRAINTS BENCHMARK STATUS SUBMITTED
226+
c83375b4 deepseek-ai/DeepSeek-V4-Flash B200:180GB:2 io=10000/1500 conc=1 tok/s/user=309 ttft=213ms ctx=1M ▂▁██▇ trialing (5/5) 2 min ago
227+
092c792b Qwen/Qwen3.5-397B-A17B RTXPRO6000:4 io=8K/1K conc=64 tok/s/user=19.6 ttft=3.43s ctx=32K ▁▂▅▇█·· verified (7) 3 days ago
228+
9ab0fa65 Qwen/Qwen3.6-27B RTXPRO4500:1 io=1K/1K conc=8 tok/s/user=57.1 ttft=499ms ctx=128K ▁▄██▆·█ verified (7) 4 days ago
229+
f91d6b60 Qwen/Qwen3-32B RTX5090:32GB:1 io=1K/512 conc=8 tok/s/user=85.8 ttft=368ms ctx=32K ▁▁▅▅▄▅▇▄▇█ verified (10) 2 weeks ago
230+
```
231+
232+
</div>
233+
234+
The `CONSTRAINTS` column is what the creation was asked for, and `BENCHMARK` is the best trial so far. `tok/s/user` is the steady decode rate, measured as one second divided by the median time per output token, so it excludes the time to the first token.
235+
236+
The glyphs after the benchmark are one per trial: height is throughput, a yellow bar is a trial whose benchmark broke a constraint, and a red `·` is one that produced no benchmark at all. The shape shows whether a run converged or wandered.
237+
238+
Pass `-w` to watch in realtime, `-v` for more detail, or `--json` for complete preset objects. Filter with `--base` or `--repo`.
184239

185240
### Delete presets
186241

@@ -189,14 +244,22 @@ Delete a preset by ID or name, or all presets for a base model with `--base`:
189244
<div class="termy">
190245

191246
```shell
192-
$ dstack preset delete 8f3a12c4
247+
$ dstack preset delete c83375b4
193248
```
194249

195250
</div>
196251

197252
For command options and agent settings, see the [`dstack preset` CLI reference](../reference/cli/dstack/preset.md).
198253

199-
> Presets are experimental, and we’d love your feedback. Report bugs and request features on [GitHub](https://github.com/dstackai/dstack/issues), and ask questions on [Discord](https://discord.gg/u8SmfwPpMd).
254+
!!! info "Roadmap and feedback"
255+
Here's what is coming soon:
256+
257+
* Allow the agent to change the source code, compile binaries, etc.
258+
* Support for PD disaggregation
259+
* Allow passing multiple `--previous <preset ID>` to `dstack preset create` to reuse the insights from previous sessions
260+
* Allow passing ranges to `concurrency`
261+
262+
Report bugs and request features on [GitHub](https://github.com/dstackai/dstack/issues), and ask questions on [Discord](https://discord.gg/u8SmfwPpMd).
200263

201264
!!! info "What's next?"
202265
1. Learn how dstack [services](services.md) work

‎src/dstack/_internal/cli/commands/preset.py‎

Lines changed: 42 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -104,10 +104,10 @@ def _register(self) -> None:
104104
help="Leave the verified service running",
105105
)
106106
create_parser.add_argument(
107-
"--max-trials",
107+
"--trials",
108108
type=int,
109109
metavar="N",
110-
help="The maximum number of benchmarked trials before the best one is promoted",
110+
help="The number of benchmarked trials before the best one is promoted",
111111
)
112112
create_parser.add_argument(
113113
"--debug",
@@ -235,12 +235,26 @@ def _list(self, args: argparse.Namespace) -> None:
235235
verbose = args.verbose
236236
if not getattr(args, "watch", False):
237237
presets, sessions = self._list_presets_and_sessions(base=base, repo=repo)
238-
print_presets(presets, sessions=sessions, verbose=verbose)
238+
print_presets(
239+
presets,
240+
sessions=sessions,
241+
verbose=verbose,
242+
all_presets=args.all_presets,
243+
limit=args.limit,
244+
)
239245
return
240246
with Live(console=console, refresh_per_second=LIVE_TABLE_REFRESH_RATE_PER_SEC) as live:
241247
while True:
242248
presets, sessions = self._list_presets_and_sessions(base=base, repo=repo)
243-
live.update(get_presets_table(presets, sessions=sessions, verbose=verbose))
249+
live.update(
250+
get_presets_table(
251+
presets,
252+
sessions=sessions,
253+
verbose=verbose,
254+
all_presets=args.all_presets,
255+
limit=args.limit,
256+
)
257+
)
244258
time.sleep(LIVE_TABLE_PROVISION_INTERVAL_SECS)
245259

246260
def _list_presets_and_sessions(
@@ -267,18 +281,21 @@ def _create(self, args: argparse.Namespace) -> None:
267281
resume_session = None
268282
if getattr(args, "resume", None):
269283
resume_session = load_resumable_agent_session(args.resume)
270-
if getattr(args, "max_trials", None) is not None:
284+
if getattr(args, "trials", None) is not None:
271285
console.print(
272-
"[warning]--max-trials is ignored when resuming: "
286+
"[warning]--trials is ignored when resuming: "
273287
"the constraints are fixed at creation[/]"
274288
)
275289
api = Client.from_config(project_name=args.project)
276290
allowed_fleets = None
277291
if resume_session is None:
278-
if configuration.max_trials is None:
292+
if configuration.trials is None:
279293
raise ConfigurationError(
280-
"max_trials is required. Set it in the configuration or pass --max-trials"
294+
"trials is required. Set it in the configuration or pass --trials"
281295
)
296+
for field in ("max_ttft", "min_context_length", "concurrency"):
297+
if getattr(configuration, field) is None:
298+
raise ConfigurationError(f"{field} is required")
282299
allowed_fleets = plan_preset(api=api, configuration=configuration)
283300
if not _confirm_preset_creation(store, configuration.name, assume_yes=args.yes):
284301
console.print("\nExiting...")
@@ -408,6 +425,21 @@ def _add_list_args(parser: argparse.ArgumentParser) -> None:
408425
action="store_true",
409426
help="Output in JSON format",
410427
)
428+
parser.add_argument(
429+
"-a",
430+
"--all",
431+
action="store_true",
432+
dest="all_presets",
433+
help="Show all presets. By default, it only shows unfinished creations or the last one.",
434+
)
435+
parser.add_argument(
436+
"-n",
437+
"--last",
438+
metavar="COUNT",
439+
type=int,
440+
dest="limit",
441+
help="Show only the last N presets. Implies --all",
442+
)
411443
model_filter = parser.add_mutually_exclusive_group()
412444
model_filter.add_argument(
413445
"--base",
@@ -490,8 +522,8 @@ def _get_effective_configuration(
490522
require_name: bool = True,
491523
) -> PresetConfiguration:
492524
_apply_name(configuration, args.name, required=require_name)
493-
if getattr(args, "max_trials", None) is not None:
494-
configuration.max_trials = args.max_trials
525+
if getattr(args, "trials", None) is not None:
526+
configuration.trials = args.trials
495527
profile = load_profile_from_args(args=args, repo_dir=Path.cwd())
496528
for field in ProfileParams.model_fields:
497529
if getattr(configuration, field) is None:

0 commit comments

Comments
 (0)