You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
`dataset` selects what every benchmark in a preset session measures: the
synthetic `random` prompts shaped by `input_tokens` and `output_tokens`, a
dataset the benchmark tool supports, or a Hugging Face dataset ID. A custom
dataset provides the requests, so the request-shape properties can't be set
with it, and the preset records the measured means instead.
The dataset is part of the contract: it is written to the session
constraints, the agent reports it with the benchmark, the preset records it,
and `dstack preset` shows it. A session that doesn't set it renders exactly
as before.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Copy file name to clipboardExpand all lines: mkdocs/docs/concepts/presets.md
+32-23Lines changed: 32 additions & 23 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -136,57 +136,67 @@ Alternatively, pass `--fleet` to `dstack preset create` or `dstack preset apply`
136
136
repo: Qwen/Qwen2.5-7B-Instruct
137
137
```
138
138
139
-
### Shared prefix
139
+
### Previous sessions
140
140
141
-
By default every request is unique, so the cache hit rate is near zero. Set `shared_prefix_tokens` to control how much of each request the serving framework can serve from its prefix cache.
141
+
Set `previous` to a list of preset IDs to give the agent the results of earlier creation sessions. It analyzes what they tried and how it worked, and aims to improve on them instead of rediscovering it.
142
142
143
143
<div editor-title="preset.dstack.yml">
144
144
145
145
```yaml
146
-
input_tokens: 8192
147
-
output_tokens: 1024
148
-
149
-
# Roughly 90% of prompt tokens can be served from cache
150
-
shared_prefix_tokens: 7360
146
+
previous:
147
+
- c83375b4
151
148
```
152
149
153
150
</div>
154
151
155
-
The `shared_prefix_tokens` value is the part of `input_tokens` that is identical across requests, such as a system prompt or conversation history, and must be less than `input_tokens`.
152
+
Alternatively, pass `--previous` (repeatable) to `dstack preset create`.
156
153
157
154
### Prompt
158
155
159
-
The `prompt` property is optional. Set it to guide the agent with custom objectives, target metrics, or an experimentation approach. It accepts inline text or a file `path`.
156
+
Set `prompt` to steer what the agent explores: which frameworks or model variants to try, or how deep to go before settling. It accepts inline text or a file `path`. Constraints such as `concurrency` and `max_ttft` can't be changed this way.
160
157
161
158
<div editor-title="preset.dstack.yml">
162
159
163
160
```yaml
164
161
prompt: |
165
-
Optimize for the lowest TTFT at concurrency 32. Consider FP8 quantization.
162
+
Profile the engine before each trial and report how far it is from the
163
+
memory-bandwidth roofline. While that gap is large, prefer patching the
164
+
serving framework over tuning flags.
166
165
```
167
166
168
167
</div>
169
168
170
-
### Baseline
169
+
### Dataset
171
170
172
-
By default, the first trial is a baseline: the agent serves the model the way the chosen serving framework recommends, without tuning it for performance. Later trials are optimization attempts. Set `baseline: false` to make every trial an optimization attempt.
171
+
The requests every benchmark measures.
173
172
174
-
### Previous sessions
173
+
=== "Random"
175
174
176
-
Set `previous` to a list of preset IDs to give the agent the results of earlier creation sessions. It analyzes what they tried and how it worked, and aims to improve on them instead of rediscovering it.
175
+
By default, benchmarks use synthetic prompts shaped by `input_tokens` and `output_tokens`. Set `shared_prefix_tokens` to make part of every request identical, such as a system prompt or conversation history, so the serving framework can serve it from its prefix cache. It must be less than `input_tokens`.
177
176
178
-
<div editor-title="preset.dstack.yml">
177
+
```yaml
178
+
input_tokens: 8192
179
+
output_tokens: 1024
179
180
180
-
```yaml
181
-
previous:
182
-
- c83375b4
183
-
```
181
+
# Roughly 90% of prompt tokens can be served from cache
182
+
shared_prefix_tokens: 7360
183
+
```
184
184
185
-
</div>
185
+
=== "Custom"
186
186
187
-
Alternatively, pass `--previous` (repeatable) to `dstack preset create`.
187
+
Set `dataset` to benchmark on real text instead: a dataset the benchmark tool supports, or a Hugging Face dataset ID.
188
+
189
+
```yaml
190
+
dataset: sharegpt
191
+
```
192
+
193
+
The dataset provides the requests, so `input_tokens`, `output_tokens`, and `shared_prefix_tokens` can't be set with it, and the preset records the measured means. A gated dataset requires `HF_TOKEN` in `env`.
194
+
195
+
### Baseline
196
+
197
+
By default, the first trial is a baseline: the agent serves the model the way the chosen serving framework recommends, without tuning it for performance. Later trials are optimization attempts. Set `baseline: false` to make every trial an optimization attempt.
188
198
189
-
In this case, the baseline trial reproduces the best comparable previous result to confirm it still holds before optimizing further.
199
+
When the session builds on `previous`, the baseline trial reproduces the best comparable previous result instead, to confirm it still holds before optimizing further.
190
200
191
201
!!! info "Reference"
192
202
The `preset` configuration supports many more options. See the [`.dstack.yml` reference](../reference/dstack.yml/preset.md).
@@ -274,7 +284,6 @@ For command options and agent settings, see the [`dstack preset` CLI reference](
274
284
* Currently, the agent doesn't upload compiled binaries anywhere; patches compile at runtime
275
285
* Doesn't support PD disaggregation (coming soon)
276
286
* Presets are saved locally (a preset registry is coming soon)
277
-
* Doesn't allow a custom dataset; always uses `random`
278
287
* Doesn't support ranges for `concurrency`
279
288
280
289
Report bugs and request features on [GitHub](https://github.com/dstackai/dstack/issues), and ask questions on [Discord](https://discord.gg/u8SmfwPpMd).
Copy file name to clipboardExpand all lines: src/dstack/_internal/cli/services/presets/resources/system_prompt.md
+38-4Lines changed: 38 additions & 4 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -39,10 +39,15 @@ Field semantics:
39
39
this session. It is fixed so that benchmark results are comparable.
40
40
<!--!TODO: support a concurrency sweep, so that a trial is measured at several
41
41
concurrencies instead of one.-->
42
+
<!--?if dataset-->
43
+
-`dataset`: the benchmark dataset for every benchmark in this session; see
44
+
`## Benchmark`.
45
+
<!--?else-->
42
46
-`input_tokens`, `output_tokens`: the request shape for every benchmark in
43
47
this session. They are fixed for the same reason.
44
48
-`shared_prefix_tokens`: how many of `input_tokens` are identical in every
45
49
request. `0` means every request is fully unique.
50
+
<!--?end-->
46
51
-`baseline`: whether the first trial must be a baseline rather than an
47
52
optimization attempt; see `# Trials (Main Section)`.
48
53
-`fleets`: use these existing `dstack` fleets only. Do not create, delete,
@@ -293,9 +298,31 @@ no trials remain. In that case, log the failure to `final_report.json` (see
293
298
## Benchmark
294
299
295
300
During trials, run benchmarks via SSH inside the task, directly against the
296
-
serving engine: use `concurrency`, `input_tokens`, `output_tokens`, and
297
-
`shared_prefix_tokens` from `constraints.json` and measure all trials the same
301
+
serving engine: use <!--?if dataset-->`dataset` and `concurrency`<!--?else-->`concurrency`, `input_tokens`, `output_tokens`, and
302
+
`shared_prefix_tokens`<!--?end--> from `constraints.json` and measure all trials the same
298
303
way so that their results are comparable with each other.
304
+
<!--?if dataset-->
305
+
Before any benchmark, reset the serving engine's prefix cache, or restart the
306
+
engine, so it does not reuse what a previous benchmark cached. Do not vary
307
+
which samples the dataset provides between benchmarks.
308
+
309
+
Use the `dataset` for every benchmark. Choose the benchmark tool's options
310
+
that load exactly that dataset, and confirm from the tool's own
311
+
documentation, for the version you run, how it loads the dataset. If the
312
+
dataset fails to load, fix the loading; never fall back to another dataset or
313
+
to synthetic prompts. Prefer the dataset's own output lengths; when the tool
314
+
forces an output length instead, use the same value in every benchmark. For
315
+
example, the dataset options are:
316
+
317
+
| tool | dataset options |
318
+
| --- | --- |
319
+
|`vllm bench serve`|`--dataset-name <dataset>` when `dataset` is the tool's own dataset name, or `--dataset-name hf --dataset-path <dataset>` when it is a Hugging Face dataset ID |
320
+
|`sglang.benchmark.serving`|`--dataset-name <dataset>` when `dataset` is the tool's own dataset name; the tool has no Hugging Face dataset option |
321
+
322
+
The table is an example and not a full command: the remaining options still
323
+
come from `concurrency`, option names and defaults differ between versions,
324
+
and any other tool needs its own equivalent.
325
+
<!--?else-->
299
326
Before any benchmark, ensure it uses a different seed than the previous
300
327
benchmark. Otherwise the benchmark will depend on what has been cached by the
301
328
previous benchmark.
@@ -315,6 +342,7 @@ lengths. For example, the shared-prefix options are:
315
342
The table is an example and not a full command: the remaining options still come
316
343
from `concurrency` and `output_tokens`, option names and defaults differ between
317
344
versions, and any other tool needs its own equivalent.
345
+
<!--?end-->
318
346
319
347
Before any benchmark — a trial one or the final one — verify that the model
320
348
works as expected: send real requests and check the responses, including
@@ -331,8 +359,9 @@ trial benchmarks in `trials/<n>/trial.json`, the final benchmark as
0 commit comments