FastAPI performance optimization

Logo
Table of contents

Server Runners Comparison

This benchmark compares five ways to run a FastAPI application in production:

All runners use the same underlying ASGI server (Uvicorn). The difference is in how they manage processes and handle startup.

Server Runners

Gunicorn with UvicornWorker

Gunicorn is a mature WSGI HTTP server that runs ASGI applications via the UvicornWorker class. It acts as a process manager, spawning multiple Uvicorn worker processes.

Uvicorn (single process)

Uvicorn is a lightning-fast ASGI server. Running it without --workers uses a single process.

Uvicorn with --workers

Uvicorn natively supports the --workers flag to spawn multiple worker processes. Simpler than Gunicorn but lacks its process management features.

FastAPI CLI

FastAPI CLI (fastapi run --workers N) wraps Uvicorn with --workers.

Measurements

CI run 29815274014 — Python 3.14, Ubuntu latest.

Gunicorn thread matrix

Workers Threads Sync RPS Async RPS
1 0 1819 2228
1 1 1784 2195
1 2 1813 2217
2 0 2654 3415
2 1 2638 3446
2 2 2586 3383

Thread count has negligible impact. Going from 0→1→2 threads changes RPS by <3%, well within noise. The GIL prevents true parallelism for CPU-bound work, and for I/O-bound FastAPI endpoints the async event loop already handles concurrency within each worker.

Doubling workers from 1→2 gives ~46% more sync RPS and ~54% more async RPS.

All runners comparison

Runner Config Sync RPS Async RPS
Gunicorn 1 worker, 0 threads 1819 2228
Gunicorn 2 workers, 0 threads 2654 3415
Uvicorn single-process 1485 1760
Uvicorn –workers 2 workers 2203 2739
FastAPI CLI 1 worker 1594 2039
FastAPI CLI 2 workers 2457 3338

Uvicorn multiprocess

Workers Sync RPS Async RPS
1 1267 1864
2 1850 2803

Verdict

Individual impact: +50-65% throughput by switching from Gunicorn (1 worker, 0 threads) to FastAPI CLI (2 workers) or Gunicorn (2 workers, 0 threads).

Config Best Runner Sync RPS Async RPS
1 worker, 0 threads Gunicorn 1819 2228
2 workers, 0 threads Gunicorn 2654 3415

Gunicorn wins at both 1 and 2 workers. Its process management overhead is negligible and it consistently outperforms all alternatives.

FastAPI CLI is the closest competitor — within 12% of Gunicorn at 1 worker and 6% at 2 workers.

Uvicorn single-process is the slowest — roughly half the throughput of multi-worker setups.

Uvicorn --workers lags Gunicorn by 17-20% despite being pure ASGI.

Thread count is irrelevant — adding threads to Gunicorn changes RPS by less than 3%. Use 0 threads unless you have a specific reason.

Recommendation: Use gunicorn -c gconf.py -k uvicorn.workers.UvicornWorker for production. FastAPI CLI is a good alternative if you prefer simpler configuration.

Re-run the measurements yourself:

git clone git@github.com:KissPeter/fastapi-performance-optimization.git
pip3 install -r test_files/requirements.txt
pytest -vv -rP test_files/ -m server_runners

Versions

Package Version
Python 3.14
FastAPI 0.139.2
Uvicorn 0.51.0
Gunicorn 26.0.0
orjson 3.11.9
ujson 5.13.0