Skip to content

Add load tests, a browser swarm and a public benchmark site - #4055

Draft
T4rk1n wants to merge 6 commits into
devfrom
feat/benchmarks
Draft

T4rk1n wants to merge 6 commits into
devfrom
feat/benchmarks

Conversation

@T4rk1n

@T4rk1n T4rk1n commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Load tests for streaming and callbacks, a headless-browser swarm for large runs, and a public site that publishes the numbers from CI.

Changes

  • benchmarks/streaming/: how many browsers each server setup can stream to (Flask/gunicorn, FastAPI and Quart on uvicorn, 1 or 4 workers). Simulated browsers follow the renderer's StreamClient and the count climbs until p95 frame latency or errors give out.
  • benchmarks/callbacks/: the same sweep for plain callbacks, HTTP vs websocket, by backend (Flask HTTP only; FastAPI and Quart both). --kind async and --work-ms add per-call work.
  • benchmarks/loadkit.py: the plumbing both share. It runs the server and the clients on separate CPUs, flags points where the clients ran out of CPU, and makes the server import the checkout under test, not the installed dash.
  • benchmarks/swarm/: real headless Chromium users (Playwright), many per machine and many machines, against a deployed app. Users click an HTTP callback, a websocket callback or a stream, timed in the page from click to DOM update, so machines only share a start time, not clocks.
    • Agents run locally or over SSH in a Docker image. They report mergeable histograms.
    • Steps stop when the server saturates or when the swarm stops delivering its load: failed agents or users, agents over 85% CPU, late starts.
    • Meant for a one-off cloud run; see benchmarks/README.md.
  • benchmarks/publish.py + .github/workflows/benchmarks-publish.yml: a static site on gh-pages.
    • Renderer timings on every push to dev, the streaming and callback sweeps weekly and on dispatch.
    • Every chart has an embeddable page (embed/<chart>.html, ?theme=light|dark), with raw data plus history under data/ and shields badges under badges/.
    • GitHub Pages needs enabling on the repo before it goes public.

Numbers

Callbacks, sync, 1 worker, 8-core laptop:

p95 at 1000 browsers p95 at 4000
Flask HTTP 3.7 ms 4.6 s (saturated from 2000)
FastAPI HTTP / websocket 1.2 / 0.9 ms 165 / 27 ms
Quart HTTP / websocket 1.9 / 1.0 ms 511 / 469 ms

With 4 workers no setup saturated at 4000 browsers; the swarm is how to find those limits. The multi-worker Redis streaming numbers need #4054.

Tests

  • Unit tests for the site builder and for histogram merging.
  • Every load test and every setup ran locally.
  • The swarm ran locally and through a fake ssh/docker fan-out. One agent with 50 users measured about 0.05 core and up to 170 MB per user.
  • Not verified: the swarm's Docker image build. The Docker daemon on the machine I used has no DNS; the playwright/python:v1.63.0-noble base tag does exist.

T4rk1n added 4 commits October 7, 2026 12:10
benchmarks/streaming/ measures how many browsers each server setup can
stream to (Flask/gunicorn, FastAPI and Quart on uvicorn, 1 or 4 workers),
with simulated browsers that follow the renderer's StreamClient.
benchmarks/publish.py builds a static site from the renderer timings and
the streaming sweep (embeddable chart pages, raw data with history,
shields badges), and benchmarks-publish.yml pushes it to gh-pages from dev.
benchmarks/callbacks/ sweeps browser counts for plain callbacks over HTTP
and over the websocket on Flask (HTTP only), FastAPI and Quart, 1 and 4
workers, recording round trip, throughput, errors and CPU. The server and
client plumbing it shares with the streaming test moves to
benchmarks/loadkit.py.
benchmarks/swarm/ drives real Chromium users (Playwright) against a
deployed app: plain HTTP callbacks, websocket callbacks and streaming,
timed in the page from click to DOM update. Agents run locally or on remote
hosts over SSH in a Docker image, start together at a shared wall-clock
time, and report mergeable histograms; steps stop when the server
saturates or when the swarm stops delivering its load (failed agents or
users, agents over 85% CPU, late starts).
The benchmark site gains a callbacks section: browsers served per setup
coloured by transport, p95 round trip for 1 and 4 workers (websocket
dashed), server CPU per 1,000 calls a second, the full table, badges and
data with history. The publish workflow runs the callback sweep weekly and
on demand, like the streaming one.
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Dash performance benchmarks

✅ all within thresholds

scenario metric p90 (ms) median growth baseline p90 note
✅ callback_chain chain_ms 434.3 418.8 0.91x 499.1
✅ callback_chain graph_ms 2.2 2.2 1.0x 2.4
✅ callback_fanout fanout_ms 97.1 75.6 0.86x 92.5
✅ deep_nesting render_ms 55.5 53.0 1.11x 56.8
✅ full_children_replace replace_ms 5471.3 2037.4 18.03x 4697.1
✅ initial_render_large render_ms 633.8 567.3 1.3x 694.4
✅ initial_render_small render_ms 104.1 89.3 0.96x 104.0
✅ patch_append_nested append_ms 136.9 94.0 2.41x 192.3
✅ patch_append_toplevel append_ms 120.6 84.1 2.22x 140.2
✅ patch_scalar_update_large update_ms 159.5 147.1 0.91x 202.9
✅ wildcard_all_resolve wildcard_ms 295.6 284.9 0.97x 313.6
✅ wildcard_all_resolve graph_ms 1.3 1.3 1.0x 1.3

growth = late-third / early-third per-op time; ~1 is flat, a large value means the per-op cost scales with accumulated state.

machine scale vs baseline: 0.93x - divided out of the baseline ratios so they compare like for like (the absolute warn/fail ceilings are left un-scaled); calibrated on initial_render_small.

T4rk1n added 2 commits October 7, 2026 13:51
Share the sweep loop between the streaming and callback runners
(loadkit.sweep), split the callback client's and swarm agent's long loops
into small steps, write the swarm report outside the event loop, dedupe
the site builder's sections and literals, run the agent image as its
unprivileged user, and pin the third-party setup-chrome action to a
commit. The site output is byte-identical for the same inputs.
@sonarqubecloud

sonarqubecloud Bot commented Oct 7, 2026

Copy link
Copy Markdown

@camdecoster
camdecoster marked this pull request as draft October 8, 2026 19:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant