When a FastAPI sync endpoint is called, Starlette dispatches it to a thread pool via anyio.to_thread.run_sync. This thread pool has a global per-worker capacity limit that directly impacts how many sync requests can execute concurrently.
Request arrives at Uvicorn event loop (per worker process)
│
├── async def endpoint → runs directly on event loop
│
└── def endpoint (sync) → Starlette calls anyio.to_thread.run_sync()
│
▼
anyio CapacityLimiter (default: 40 tokens)
│
├── Token available? → dispatch to thread, execute sync code
└── No token? → request waits in queue until a token is released
import anyio.to_thread
limiter = anyio.to_thread.current_default_thread_limiter()
print(limiter.total_tokens) # 40
The default is 40 tokens per worker process. This is a hard ceiling — even if your machine has 64 cores, each worker can only run 40 sync threads simultaneously.
The 40-token limit is shared across all sync operations within a worker:
| Consumer | Blocks a token? |
|---|---|
Sync endpoint (def endpoint()) |
Yes |
Sync dependency (def dep()) |
Yes |
Sync BackgroundTasks |
Yes |
Sync StreamingResponse iterator |
Yes |
UploadFile file I/O |
Yes |
FileResponse / StaticFiles |
Yes |
If your sync endpoint calls a sync dependency that calls an external API, that’s one token held for the entire duration.
A sync endpoint holding a thread pool token while waiting for an HTTP connection means:
Thread pool tokens: 40 (per worker)
Connection pool: N connections (per worker)
Effective concurrency = min(40, N)
If N < 40, the connection pool is the bottleneck — threads wait for connections.
If N > 40, extra connections sit idle — wasted resources.
Important: The connection pool limits connection concurrency, not handler concurrency. Handlers beyond the pool size queue for connections but remain “active” (holding a thread token). This means:
CI-verified (run 29907570007): pool_size=2 showed max_concurrent=10 handlers per worker. The thread pool (40 tokens) is the actual concurrency ceiling, not the pool size.
Increase the token count only if you have a genuine need for more concurrent sync threads:
# In your FastAPI lifespan or startup
import anyio.to_thread
async def lifespan(app):
limiter = anyio.to_thread.current_default_thread_limiter()
limiter.total_tokens = 100 # increase if needed
yield
Each Gunicorn/Uvicorn worker is a separate OS process with its own:
So with 2 workers: total concurrent sync threads across the app = 40 × 2 = 80.
threads settingWhen using UvicornWorker (our setup), Gunicorn’s --threads option is not used for concurrency control. UvicornWorker runs its own async event loop. The threads environment variable in our docker-compose is a no-op for UvicornWorker.
The actual concurrency is controlled by:
When using gthread worker class (not UvicornWorker), Gunicorn’s --threads setting controls the thread pool size directly.
| Scenario | Suggested tokens |
|---|---|
| Mostly async endpoints | 40 (default) |
| Mix of sync/async, moderate load | 40-80 |
| Mostly sync endpoints, high concurrency | 80-200 |
| Sync endpoints with slow I/O (DB, APIs) | Match to connection pool size |
The key insight: tokens should match your connection pool size for sync endpoints that make external calls. If your pool has 50 connections, you need at least 50 tokens to utilize them all.
CI runs 29928705459 and 30024860601 — all concurrency tests passed.
worker_pid and concurrent_at_start (how many handlers were active when it started)| Config | Pool | Tokens | Requests | Concurrent ≥5 per worker |
|---|---|---|---|---|
| anyio_tokens_40_w2 | 100 | 40 | 60 | ✓ Passed |
| anyio_tokens_80_w2 | 100 | 80 | 100 | ✓ Passed |
| anyio_tokens_100_w2 | 100 | 100 | 120 | ✓ Passed |
All three configurations showed handler concurrency reaching well above the minimum threshold per worker, confirming that increasing anyio tokens from 40 to 80/100 allows more concurrent sync handlers when the connection pool is not the bottleneck.
The anyio token limit is the true concurrency ceiling for sync endpoints, not the connection pool size. With pool=100 and tokens=40, only ~40 sync handlers can run concurrently per worker — the remaining 60 pool connections sit idle.