Deploy a FastAPI App on AWS EC2 with Nginx, Gunicorn and systemd
There is a large gap between a FastAPI app that runs with uvicorn main:app --reload on your laptop and one that survives a production weekend. The development server is single-process, restarts on file changes, speaks plain HTTP, and dies the moment your SSH session closes. Production needs the opposite of all four properties.
This guide walks through the whole path on a bare Ubuntu 24.04 EC2 instance: a dedicated system user, a virtualenv, Gunicorn supervising Uvicorn workers, a systemd unit that starts on boot and restarts on crash, and Nginx in front terminating TLS and serving static files. Every command here is one you actually run, in order.
Architecture in one line: the internet talks to Nginx on 443, Nginx proxies to Gunicorn on a Unix socket, Gunicorn supervises N Uvicorn worker processes, and systemd supervises Gunicorn. Each layer restarts the one below it.
Choosing and preparing the instance
For a typical API serving a few hundred requests per second, a t3.small (2 vCPU, 2 GB) is a reasonable starting point, and t3.micro is enough for staging. The burstable t3 family is fine for request-response API workloads because they are idle between bursts; avoid it for anything that pegs CPU continuously, since you will exhaust your CPU credits and get throttled to a fraction of a core with no obvious error message.
Give the instance a security group that opens only what you need. Port 22 should be restricted to your own IP or, better, closed entirely in favour of AWS Systems Manager Session Manager. Ports 80 and 443 open to 0.0.0.0/0. Nothing else — in particular, do not open 5432 or 8000 to the world, which is the single most common way a hobby deployment gets compromised.
| Port | Source | Purpose |
|---|---|---|
| 22 | Your IP only | SSH administration |
| 80 | 0.0.0.0/0 | HTTP, redirects to HTTPS and serves ACME challenges |
| 443 | 0.0.0.0/0 | HTTPS, the only port that serves real traffic |
| 8000 | Nobody | Gunicorn — reached through a Unix socket, never exposed |
| 5432 | App SG only | PostgreSQL, if it runs on a separate instance |
Also attach an Elastic IP before you point DNS at the box. A stopped-and-started EC2 instance gets a new public IP otherwise, and you will spend an afternoon wondering why your domain resolves to nothing.
A dedicated system user, not root
Running your application as root means any remote code execution bug in a dependency is an instant full-machine compromise. Create an unprivileged user that owns the code and nothing else. It gets no login shell, because it is never meant to be logged into.
sudo adduser --system --group --home /opt/myapi --shell /usr/sbin/nologin myapi
sudo mkdir -p /opt/myapi/app
sudo chown -R myapi:myapi /opt/myapiInstall the system packages you need. Ubuntu 24.04 ships Python 3.12, which every current FastAPI release supports, so there is no need for a deadsnakes PPA unless you have a specific version requirement.
sudo apt update
sudo apt install -y python3-venv python3-dev build-essential nginx libpq-devlibpq-dev is there so psycopg can compile against the PostgreSQL client library. If you use psycopg[binary] you can skip it, though the binary wheel is discouraged for production because it bundles its own libpq and libssl that you then have to patch separately from the system ones.
Deploying the code and its virtualenv
Put the application in /opt/myapi/app and its virtualenv in /opt/myapi/venv — keeping them siblings rather than nesting the venv inside the repo means a git clean or a fresh checkout never destroys your installed packages.
sudo -u myapi git clone https://github.com/you/myapi.git /opt/myapi/app
sudo -u myapi python3 -m venv /opt/myapi/venv
sudo -u myapi /opt/myapi/venv/bin/pip install --upgrade pip
sudo -u myapi /opt/myapi/venv/bin/pip install -r /opt/myapi/app/requirements.txt
sudo -u myapi /opt/myapi/venv/bin/pip install gunicorn uvicorn[standard]uvicorn[standard] pulls in uvloop and httptools, which are meaningfully faster than the pure-Python event loop and HTTP parser. This is not a micro-optimisation — uvloop alone typically buys 20 to 40 percent on request throughput for I/O-bound handlers, and it costs you one line.
Pin your dependencies. A requirements.txt full of unpinned package names means the version you tested and the version that lands on the server are decided by whatever was on PyPI the day you deployed. Use pip freeze, or better, a lockfile from uv or Poetry.
Configuration and secrets
Never commit a database password. The simplest arrangement that is genuinely safe on a single box is an environment file readable only by the service user, loaded by systemd rather than by the application.
sudo install -o myapi -g myapi -m 600 /dev/null /opt/myapi/.env
sudo -u myapi tee /opt/myapi/.env > /dev/null <<'EOF'
DATABASE_URL=postgresql+asyncpg://myapi:REPLACE_ME@localhost:5432/myapi
SECRET_KEY=REPLACE_ME
ENVIRONMENT=production
EOFMode 600 with the service user as owner means root and the app can read it and nobody else can, including any other unprivileged account on the machine. When you graduate past one instance, move these into AWS Secrets Manager or SSM Parameter Store and fetch them at boot — but do that when you need it, not before.
Gunicorn with Uvicorn workers
FastAPI is an ASGI framework, so it needs an ASGI server. Uvicorn is that server. Gunicorn is a process manager that knows how to spawn, monitor, and recycle worker processes — it does not speak ASGI itself, which is why you run Uvicorn's worker class inside it. Gunicorn handles the supervision, Uvicorn handles the protocol.
The reasonable worker count for a CPU-bound workload is (2 x cores) + 1. For an async I/O-bound API, which most FastAPI services are, that formula overshoots: each Uvicorn worker runs its own event loop and can hold thousands of concurrent connections, so you are usually better off with one worker per core and letting async concurrency do the rest.
/opt/myapi/venv/bin/gunicorn app.main:app \
--workers 2 \
--worker-class uvicorn.workers.UvicornWorker \
--bind unix:/run/myapi/gunicorn.sock \
--timeout 60 \
--graceful-timeout 30 \
--keep-alive 5 \
--access-logfile - \
--error-logfile -A few of these flags matter more than they look. Binding to a Unix socket instead of 127.0.0.1:8000 removes a whole class of problem: the port cannot be accidentally exposed, and socket permissions become filesystem permissions. Logging to stdout and stderr with a single dash hands log management to systemd's journal instead of leaving you to rotate files yourself.
The timeout is the one that bites people. Gunicorn's --timeout kills a worker that has not responded to the master's heartbeat, and the default is 30 seconds. If you have a legitimately slow endpoint — a report export, a large upload — you will see workers being killed mid-request and returning nothing to the client. Raise the timeout for those cases, or better, move the work to a background task.
| Flag | Default | Why you change it |
|---|---|---|
| --workers | 1 | One per CPU core for async apps; (2 x cores) + 1 for sync |
| --timeout | 30s | Raise it if any request legitimately runs longer |
| --graceful-timeout | 30s | How long a worker gets to finish in-flight requests on reload |
| --keep-alive | 2s | Raise to 5s when Nginx sits in front and reuses connections |
| --max-requests | off | Set to a few thousand to recycle workers and bound slow memory leaks |
The systemd unit
systemd is what makes this a service rather than a process someone started once. It starts the app at boot, restarts it if it crashes, captures its logs, and — through its sandboxing directives — constrains what a compromised process can reach.
Write the unit to /etc/systemd/system/myapi.service:
[Unit]
Description=MyAPI FastAPI service
After=network.target postgresql.service
Wants=postgresql.service
[Service]
Type=notify
User=myapi
Group=myapi
RuntimeDirectory=myapi
WorkingDirectory=/opt/myapi/app
EnvironmentFile=/opt/myapi/.env
ExecStart=/opt/myapi/venv/bin/gunicorn app.main:app \
--workers 2 \
--worker-class uvicorn.workers.UvicornWorker \
--bind unix:/run/myapi/gunicorn.sock \
--timeout 60 \
--graceful-timeout 30 \
--keep-alive 5 \
--access-logfile - \
--error-logfile -
ExecReload=/bin/kill -s HUP $MAINPID
Restart=always
RestartSec=3
# Sandboxing — a compromised worker should not be able to reach the rest of the box.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/opt/myapi
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
[Install]
WantedBy=multi-user.targetRuntimeDirectory=myapi is doing quiet but important work: systemd creates /run/myapi owned by the service user before ExecStart and removes it on stop. Without it your socket path does not exist on a fresh boot and the service fails to start — which is exactly the kind of bug that only shows up the first time the instance reboots, weeks after you deployed.
Type=notify lets Gunicorn tell systemd when it is genuinely ready to accept connections rather than systemd assuming readiness the moment the process forks. That matters for ordering and for zero-downtime restarts.
sudo systemctl daemon-reload
sudo systemctl enable --now myapi
sudo systemctl status myapi
sudo journalctl -u myapi -fProtectSystem=strict mounts the entire filesystem read-only except what you list in ReadWritePaths. If your app writes anywhere else — a cache directory, an upload staging path — add it explicitly or you will get permission errors that look nothing like permission errors.
Nginx in front
Nginx is not optional decoration. It terminates TLS, buffers slow clients so they do not tie up a Python worker for the duration of a mobile upload, serves static files without touching your app, and gives you a place to put rate limits and security headers.
Create /etc/nginx/sites-available/myapi:
upstream myapi {
server unix:/run/myapi/gunicorn.sock fail_timeout=0;
}
server {
listen 80;
server_name api.example.com;
# Certbot writes its challenge files here; everything else goes to HTTPS.
location /.well-known/acme-challenge/ {
root /var/www/html;
}
location / {
return 301 https://$host$request_uri;
}
}
server {
listen 443 ssl;
http2 on;
server_name api.example.com;
ssl_certificate /etc/letsencrypt/live/api.example.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/api.example.com/privkey.pem;
client_max_body_size 25m;
location /static/ {
alias /opt/myapi/app/static/;
expires 30d;
add_header Cache-Control "public, immutable";
}
location / {
proxy_pass http://myapi;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_redirect off;
proxy_read_timeout 60s;
}
}Enable it, drop Nginx's default site so it does not shadow yours, and reload:
sudo ln -s /etc/nginx/sites-available/myapi /etc/nginx/sites-enabled/
sudo rm -f /etc/nginx/sites-enabled/default
sudo nginx -t && sudo systemctl reload nginxAlways run nginx -t before reloading. A syntax error in a reload is harmless because Nginx refuses to apply it, but a syntax error discovered during a restart leaves you with no web server at all.
Telling FastAPI it is behind a proxy
This step is skipped constantly and causes bugs that look supernatural. Behind Nginx, your app sees every request as coming from 127.0.0.1 over plain HTTP. Redirects generated with request.url_for come back as http:// instead of https://, and any rate limiting or logging keyed on client IP sees one address for the entire internet.
The fix is to have Uvicorn honour the X-Forwarded-* headers Nginx is already sending. Add the flags to your Gunicorn command:
--forwarded-allow-ips="127.0.0.1" --proxy-headersOnly trust forwarded headers from an address you actually control. Setting forwarded-allow-ips to * on a server reachable from the internet lets any client spoof its own IP address, which quietly defeats the rate limiting you are about to configure.
If TLS terminates at an ALB rather than at Nginx, the same applies but the trusted address is the load balancer subnet, not localhost. Getting this wrong is how an API ends up trusting a header the client set.
TLS with Certbot
Point your DNS A record at the Elastic IP first and wait for it to resolve — Certbot validates by fetching a file over HTTP from the name you are requesting, so it fails if DNS has not propagated.
sudo apt install -y certbot python3-certbot-nginx
sudo certbot --nginx -d api.example.com
sudo systemctl list-timers | grep certbotThe Debian and Ubuntu packages install a systemd timer that renews twice a day, so there is no cron job to write. Verify the renewal path works before you forget about it: certbot renew --dry-run exercises the whole flow against the staging endpoint without burning rate limit quota.
Health checks and deploys
Add a health endpoint that checks the things that actually break, not one that returns a constant. A handler that returns 200 unconditionally will happily report health while the database connection pool is exhausted.
from fastapi import APIRouter, Response, status
from sqlalchemy import text
router = APIRouter()
@router.get("/healthz")
async def healthz(response: Response):
try:
async with engine.connect() as conn:
await conn.execute(text("SELECT 1"))
except Exception:
response.status_code = status.HTTP_503_SERVICE_UNAVAILABLE
return {"status": "degraded", "database": "unreachable"}
return {"status": "ok"}For deploys, a HUP signal makes Gunicorn start new workers with the new code and retire the old ones as they finish their in-flight requests. Combined with a health check, that gets you a restart nobody notices:
cd /opt/myapi/app && sudo -u myapi git pull --ff-only
sudo -u myapi /opt/myapi/venv/bin/pip install -r requirements.txt
sudo -u myapi /opt/myapi/venv/bin/alembic upgrade head
sudo systemctl reload myapi
curl -fsS https://api.example.com/healthzNote that reload, not restart. A restart drops in-flight requests; a reload does not. And run migrations before the reload, not after, so the new code never starts against an old schema.
When it does not work
Four failures account for most of the debugging time on a first deployment. Each has a distinctive signature.
| Symptom | Usual cause | Check |
|---|---|---|
| 502 Bad Gateway | Gunicorn is not running, or Nginx cannot reach the socket | systemctl status myapi and ls -l /run/myapi/ |
| 502 after a reboot only | RuntimeDirectory missing, socket path never created | journalctl -u myapi -b |
| 403 on the socket | www-data cannot traverse to /run/myapi | Check directory mode and group ownership |
| Redirects go to http:// | Proxy headers not trusted | Confirm --proxy-headers and --forwarded-allow-ips |
| Worker timeout in logs | A handler is blocking the event loop | Look for sync I/O inside an async def |
That last one deserves emphasis, because it is the most common genuine performance bug in FastAPI code. Calling a blocking library — requests, a synchronous database driver, time.sleep — inside an async def handler blocks the entire event loop, and therefore every other concurrent request that worker is serving. Either use an async client, or declare the handler as a plain def and let FastAPI run it in a threadpool. It will do the right thing automatically.
# Wrong — blocks the event loop for every concurrent request on this worker.
@app.get("/report")
async def report():
data = requests.get("https://slow.example.com/data").json()
return data
# Right — a sync handler, which FastAPI runs in a threadpool.
@app.get("/report")
def report():
data = requests.get("https://slow.example.com/data").json()
return dataWhat to do next
- Harden the instance itself — SSH keys only, UFW, fail2ban and unattended upgrades
- Move PostgreSQL onto its own instance or RDS once the database and app start competing for memory
- Add rate limiting and security headers in the Nginx layer rather than in application code
- Replace the manual git pull deploy with a pipeline that builds, tests, and reloads on merge
- Ship logs and metrics off the box, so a terminated instance does not take your only copy of them with it
Do I need Gunicorn at all, or can I run Uvicorn directly?
You can run Uvicorn directly with --workers, and for many deployments that is enough. Gunicorn buys you a more mature process manager: graceful worker recycling with --max-requests, better signal handling for zero-downtime reloads, and a large amount of production history. If you run Uvicorn alone, keep systemd as the supervisor and use Restart=always.
How many workers should I actually run?
Start with one Uvicorn worker per CPU core and measure. Async workers hold many concurrent connections each, so adding workers beyond core count mostly adds memory use and context switching rather than throughput. If your handlers are CPU-heavy rather than I/O-heavy, the classic (2 x cores) + 1 is closer to right.
Should I use a Unix socket or a TCP port between Nginx and Gunicorn?
A Unix socket, when both run on the same host. It is marginally faster, and more importantly it cannot be reached from the network at all, so a misconfigured security group cannot accidentally expose your app server. Use TCP only when Nginx and the app run on different machines.
Why does my app work over HTTP but break behind HTTPS?
Almost always the proxy headers. Nginx terminates TLS and forwards plain HTTP, so unless you pass --proxy-headers and set --forwarded-allow-ips, your app believes the request scheme is http and generates http:// URLs, which browsers then block as mixed content.
Is a single EC2 instance good enough for production?
For a great many real services, yes — one well-configured instance handles far more traffic than people assume. What a single box does not give you is redundancy: any reboot, kernel patch, or hardware failure is downtime. Move to two instances behind a load balancer when downtime starts costing more than the second instance.