FastAPI performance optimization

Logo
Table of contents

FastAPI performance tuning

FastAPI is a great, high performance web framework but far from perfect. This document is intended to provide some tips and ideas to get the most out of it

Use these techniques to achieve up to ~2.2x performance from your FastAPI application

All tested on the same sized Docker containers (2 CPU cores). The performance numbers below represent real CI-verified measurements.

Combined optimization impact

Each optimization below has a measurable, compounding effect. When applied together, the total improvement is multiplicative, not additive.

Sync endpoints

Optimization Impact Details
Gunicorn (1 worker) → FastAPI CLI (2 workers) +50-65% throughput Server runners
JSON → ORJSON response class +4-13% throughput JSON response classes
BaseHTTPMiddleware → Starlette ASGI +35-44% throughput (avoiding BaseHTTPMiddleware cost) Middleware
Best vs Worst combo +116% throughput (2.2x) Measured below

Async endpoints

Optimization Impact Details
Sync → Async endpoints +20-33% throughput Sync vs Async
Gunicorn (1 worker) → FastAPI CLI (2 workers) +50-65% throughput Server runners
JSON → ORJSON response class +4-13% throughput JSON response classes
BaseHTTPMiddleware → Starlette ASGI +35-44% throughput (avoiding BaseHTTPMiddleware cost) Middleware
Best vs Worst combo +118% throughput (2.2x) Measured below

Measured best vs worst configurations

CI run 36235373898 — all 4 jobs passed, fully green, no non-2xx responses.

Using the big JSON response endpoint (1MB payload) across all tested combinations:

Configuration key: w{N} = N worker processes, t{N} = N threads per worker.

  • Gunicorn w1t0: Gunicorn with 1 worker process, 0 threads (single-process baseline)
  • FastAPI CLI w2: FastAPI CLI (fastapi run --workers 2) with 2 worker processes
  • ORJSON: ORJSONResponse (fast JSON serialization via orjson)
  • Starlette ASGI: Starlette native ASGI middleware (not BaseHTTPMiddleware)

Sync endpoint (/sync/big_json_response)

Configuration RPS Latency
Best: FastAPI CLI (2 workers) + async + ORJSON + Starlette ASGI 27.90 3583.59 ms
Worst: Gunicorn (1 worker, 0 threads) + sync + JSON + BaseHTTPMiddleware 12.94 7729.63 ms
Improvement +115.69% -4146 ms

Async endpoint (/async/big_json_response)

Configuration RPS Latency
Best: FastAPI CLI (2 workers) + async + ORJSON + Starlette ASGI 27.86 3589.66 ms
Worst: Gunicorn (1 worker, 0 threads) + sync + JSON + BaseHTTPMiddleware 12.79 7820.47 ms
Improvement +117.86% -4231 ms

Small payload endpoints

Endpoint Best RPS Worst RPS Improvement
/sync/items/ 3102.39 1391.80 +122.91%
/async/items/ 4304.53 1766.31 +143.70%

The compounding effect

When you combine ALL optimizations:

Factor Worst config Best config
Server runner Gunicorn (1 worker, 0 threads) FastAPI CLI (2 workers)
Endpoint type Sync Async
Response class JSON ORJSON
Middleware BaseHTTPMiddleware Starlette ASGI (or none)
Combined RPS (big JSON) 12.79 27.86
Combined improvement   ~2.2x throughput

The difference between a default FastAPI setup and a properly optimized one is ~2.2x throughput on big JSON responses and ~2.2-2.4x on small payloads. These are not theoretical numbers — they are measured in CI on identical Docker infrastructure (2 CPU cores per container).

Topics by category

Concurrency & workers

Transport & networking

Response & middleware

Tools & debugging

Robustness & Reliability (Generic, not FastAPI-specific)

These patterns apply to any Python web application. They contribute to a robust and reliable application but are not FastAPI performance optimizations.

Test environment