The code in this repo was used to produce How continuous batching enables 23x throughput in LLM inference while reducing p50 latency.
This repository was archived by the owner on Jul 8, 2026. It is now read-only.
| Name | Name | Last commit date | ||
|---|---|---|---|---|
The code in this repo was used to produce How continuous batching enables 23x throughput in LLM inference while reducing p50 latency.