guimauve, French for marshmallow, is an inference server built on top of:
- seamless burn integration to import models
- production-ready inference serving thanks to axum
- concurrency control: semaphore-based admission control to prevent overload
- ease of use:
Internal performance testing has proven guimauve to be a lot less resource-hungry and behave
better under load compared to a FastAPI + PyTorch set up.
- GPU inference
- batch inference
Add guimauve as a dependency:
cargo add guimauveImplement the ModelPlugin trait:
// lib.rs
use burn::backend::Flex;
use burn::tensor::{Int, Tensor};
use guimauve::model_plugin::ModelPlugin;
struct MyPlugin;
impl ModelPlugin for MyPlugin {
type Request = serde_json::Value;
type Response = serde_json::Value;
type ModelInput = Tensor<Flex, 2, Int>;
type ModelOutput = Tensor<Flex, 2, Int>;
type Error = anyhow::Error;
fn pre(&self, req: Self::Request) -> Result<Self::ModelInput, Self::Error> {
// parse and prepare model input
todo!()
}
fn infer(&self, input: Self::ModelInput) -> Result<Self::ModelOutput, Self::Error> {
// run inference
todo!()
}
fn post(&self, output: Self::ModelOutput) -> Result<Self::Response, Self::Error> {
// format response
todo!()
}
}Define your entrypoint:
// main.rs
fn main() -> anyhow::Result<()> {
let plugin = MyPlugin;
guimauve::server::Server::builder(plugin)
.address("0.0.0.0:3000")
.build()?
.serve()
}Two endpoints are available:
curl http://0.0.0.0:3000/v1/health
# inference
curl -X POST http://0.0.0.0:3000/v1/infer \
-H 'Content-Type: application/json' \
-d '{"en_sentence": "why are people from Lisboa eating snails?"}'You can check the tutorial below for a full-fledged example or the examples in [the dedicated folder][./examples].
This is a small tutorial to run an inference server on an English to Portuguese translation model.
Download the datasets from https://web.archive.org/web/20240301220426if_/http://www.phontron.com/data/qi18naacl-dataset.tar.gz
and extract them. We only care about the pt_to_en folder.
From the python-onnx-example folder, run:
uv run \
-m train.train_tokenizer \
-i datasets/pt_to_en/pt.dev \
datasets/pt_to_en/pt.test \
datasets/pt_to_en/pt.train \
datasets/pt_to_en/pt.train.r0.125 \
datasets/pt_to_en/pt.train.r0.25 \
datasets/pt_to_en/pt.train.r0.5 \
-o pt_tokenizer.json \
--max-seq-len 128This will create a pt_tokenizer.json file.
Same thing for the English data which will create a en_tokenizer.json file.
From the python-onnx-example folder, run:
uv run \
-m train.train_transformer \
--en-tokenizer models/pt_to_en/en_tokenizer.json \
--en-train datasets/pt_to_en/en.train \
--en-val datasets/pt_to_en/en.dev \
--en-test datasets/pt_to_en/en.test \
--pt-tokenizer models/pt_to_en/pt_tokenizer.json \
--pt-train datasets/pt_to_en/pt.train \
--pt-val datasets/pt_to_en/pt.dev \
--pt-test datasets/pt_to_en/pt.test \
--output-dir models/pt_to_en \
--debugThis might take a while depending on your computer specs.
This will output a transformer.onnx file.
Modify ./crates/onnx-example/build.rs and point the path to where your transformer.onnx file is
located.
Build the Docker image, ARTIFACTS_DIR should contain the tokenizer files built previously.
docker build \
--build-arg CRATE_NAME=onnx-example \
--build-arg ARTIFACTS_DIR=models/pt_to_en \
-t onnx-example .docker run -p 3000:3000 --name onnx-example onnx-exampleCheck it's working:
curl -X POST localhost:3000/v1/infer \
-H 'Content-Type: application/json' \
-d '{"en_sentence": "why are people from Lisboa eating snails?"}'There is also a Docker compose set up with:
- cadvisor for resource usage monitoring
- prometheus for the time series database set up to scrape cadvisor
- grafana to display the data in prometheus in a dashboard
CRATE_NAME=onnx-example \
ARTIFACTS_DIR=models/pt_to_en \
docker compose up --force-recreate --remove-orphans --detach --build guimauve
You can check the Grafana dashboard at http://localhost:8083/.
If you have a model already saved locally or in Databricks using MLflow, this repository offers a small utility to convert it to ONNX:
uv run dbx-to-onnx \
-m models/pt_to_en \
-i "source:int64:1,128" \
-i "target:int64:1,127" \
-o models/pt_to_enThis will create a model.onnx file in the specified output directory.
There is more information in the dedicated README.
