A simple tutorial on deploying fastapi apps using ray_serve
-
Install Docker from here
- Note: If you do not want to use Docker you can run the notebook/scripts directly on your machine also using jupyter lab
-
Run below script
#!/bin/bash git clone https://github.com/abhishek9sharma/ray_serve_fastapi_tutorial.git cd ray_serve_fastapi_tutorial make up_with_build
This notebook demonstrates how to deploy a Hello World FastAPI app Using Ray Serves Python API
-
If you have installed Docker, provided you have run the command
make up_with_buildmentioned before, you can browsethe below link -
If you are using jupyter lab
- run the command
jupyter lab - browse the url http://localhost:8888/lab or http://localhost:8889/lab
- run the command
-
Follow the instructions in the notebook serving_fastapi_ray.ipynb. It is self contained.
Below steps demonstrate how to deploy a Hello World FastAPI app Using Ray Serves CLI
-
Below comamnds should be run frome the cloned folder i.e. ray_serve_fastapi_tutorial
-
Change Directory to workspace ( if not using docker)
cd ray_serve_tutorial/workspace/
-
Install Environment
chmod +x src/install_env.sh bash src/install_env.sh -
Activate the environment
source ray_env/bin/activate -
Spin Up Ray Cluster
ray start --head --dashboard-host 0.0.0.0- Ray cluster should be visible at http://localhost:8265/
- Status can also be verified using
ray status
-
Serving App
-
Serve Ray App From CodeLocation
- Run below commands
serve start --http-host 0.0.0.0 --http-port 8001serve run src.ray_fastapi:rayappadvanced --non-blocking
- The app should be visible at http://localhost:8001/docs and serve at http://localhost:8001/hello
- Run below commands
-
Serve Ray App From Config file
- Run below commands
serve shutdown -yserve build src.ray_fastapi:rayappadvanced -o serve_config_app.yamlserve start --http-host 0.0.0.0 --http-port 8001serve deploy serve_config_app.yaml
- The app should be visible at http://localhost:8001/docs and serve at http://localhost:8001/hello
- Run below commands
-
Serve Replica Autoscaling
- Run app with autoscalilng config
-
Run command
serve shutdown -y -
Remove the num_replicas in serve_config_app.yaml
-
Add below config (You can refer serve_config_app_autoscale.yaml)
max_ongoing_requests: 5 autoscaling_config: target_ongoing_requests: 2 min_replicas: 2 max_replicas: 5 -
Redeploy app using below commands
serve start --http-host 0.0.0.0 --http-port 8001serve deploy serve_config_app.yaml
-
- Simulate AutoScaling
- Start
locust -f src/locust_test.py --web-port 8004. - It should be visible at http://localhost:8004/
- Load test http://localhost:8001 with requests using locust. You can use below settings
- Number of Users 100
- Ramp Up 10
- Host http://localhost:8001/
- You can see auto_scaling of replicas increasing at http://localhost:8265/#/serve/applications/app1/RayAppAdvanced after some time
- Start
- Run app with autoscalilng config
-