Most deployment pipelines begin as a shell script someone runs over SSH, and that is a perfectly honest starting point. The problems arrive later: the script assumes the server is in a known state, there is no record of what was deployed, a failed deploy leaves the service half-updated, and the AWS credentials it uses have been sitting in a GitHub secret since 2023.

This builds a pipeline that fixes all four. It tests on every push, builds an image tagged with the commit SHA, authenticates to AWS with short-lived OIDC credentials rather than stored keys, deploys through Systems Manager rather than SSH, and verifies health before declaring success.

Stop storing AWS keys in GitHub

A long-lived AWS_SECRET_ACCESS_KEY in repository secrets is a credential that never expires, is readable by every workflow in the repo, and survives the departure of whoever created it. GitHub Actions can instead present an OIDC token that AWS trusts directly, exchanging it for credentials that live for the duration of the job.

Register GitHub as an identity provider in IAM once:

aws iam create-open-id-connect-provider \
  --url https://token.actions.githubusercontent.com \
  --client-id-list sts.amazonaws.com

Then create a role whose trust policy accepts tokens only from your repository, on your default branch:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": {
      "Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com"
    },
    "Action": "sts:AssumeRoleWithWebIdentity",
    "Condition": {
      "StringEquals": {
        "token.actions.githubusercontent.com:aud": "sts.amazonaws.com"
      },
      "StringLike": {
        "token.actions.githubusercontent.com:sub": "repo:your-org/your-repo:ref:refs/heads/main"
      }
    }
  }]
}

The sub condition is the whole security boundary. A wildcard like repo:your-org/*:* lets any workflow in any repository in the organisation — including one opened by a pull request from a fork, if you are careless with triggers — assume this role. Scope it to the exact repository and ref.

The test job

Run tests against a real PostgreSQL rather than SQLite. Service containers make this straightforward, and testing against a different database than you deploy to is how migration bugs reach production.

name: deploy

on:
  push:
    branches: [main]

concurrency:
  group: deploy-${{ github.ref }}
  cancel-in-progress: false

permissions:
  contents: read
  id-token: write
  packages: write

jobs:
  test:
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:17-alpine
        env:
          POSTGRES_PASSWORD: postgres
          POSTGRES_DB: test
        options: >-
          --health-cmd pg_isready
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5
        ports: ['5432:5432']
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.12'
          cache: pip
      - run: pip install -r requirements.txt -r requirements-dev.txt
      - run: ruff check .
      - run: mypy app
      - run: pytest -q --maxfail=1
        env:
          DATABASE_URL: postgresql://postgres:postgres@localhost:5432/test

The concurrency block is doing real work. Without it, two merges in quick succession run two deploy jobs in parallel and the older one can finish last, leaving the server running the previous commit. cancel-in-progress: false is deliberate — you want the second deploy to queue behind the first, not to kill a deploy halfway through.

Building and tagging the image

Tag with the commit SHA. A tag like latest means you cannot tell what is running, and cannot roll back to a specific known-good build.

  build:
    needs: test
    runs-on: ubuntu-latest
    outputs:
      tag: ${{ steps.meta.outputs.tag }}
    steps:
      - uses: actions/checkout@v4

      - id: meta
        run: echo "tag=sha-${GITHUB_SHA::7}" >> "$GITHUB_OUTPUT"

      - uses: docker/setup-buildx-action@v3

      - uses: docker/login-action@v3
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}

      - uses: docker/build-push-action@v6
        with:
          context: .
          push: true
          tags: |
            ghcr.io/${{ github.repository }}:${{ steps.meta.outputs.tag }}
            ghcr.io/${{ github.repository }}:latest
          cache-from: type=gha
          cache-to: type=gha,mode=max
          provenance: false

cache-from and cache-to with type=gha use GitHub's own layer cache, which typically cuts build times by more than half on repeat runs. Push both the SHA tag and latest — the SHA is what you deploy, latest is a convenience for humans.

Deploying without SSH

SSH deployment requires an inbound port open to GitHub's IP ranges and a private key stored as a secret. AWS Systems Manager Session Manager removes both: the instance makes an outbound connection to SSM, and your workflow sends commands through the AWS API using its OIDC credentials. Port 22 can be closed entirely.

The instance needs the SSM agent — preinstalled on Amazon Linux and Ubuntu AMIs — and an instance profile with AmazonSSMManagedInstanceCore attached.

  deploy:
    needs: build
    runs-on: ubuntu-latest
    environment: production
    steps:
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::111122223333:role/github-deploy
          aws-region: ap-south-1

      - name: Deploy via SSM
        env:
          TAG: ${{ needs.build.outputs.tag }}
        run: |
          set -euo pipefail
          CMD_ID=$(aws ssm send-command \
            --document-name "AWS-RunShellScript" \
            --targets "Key=tag:Role,Values=api" \
            --comment "deploy $TAG" \
            --parameters commands="[
              'set -euo pipefail',
              'cd /opt/myapi',
              'export APP_TAG=$TAG',
              'docker compose pull api',
              'docker compose run --rm migrate',
              'docker compose up -d --no-deps api',
              'sleep 5',
              'curl -fsS --retry 10 --retry-delay 3 http://localhost:8000/healthz'
            ]" \
            --query 'Command.CommandId' --output text)

          aws ssm wait command-executed \
            --command-id "$CMD_ID" \
            --instance-id "$(aws ec2 describe-instances \
              --filters 'Name=tag:Role,Values=api' 'Name=instance-state-name,Values=running' \
              --query 'Reservations[0].Instances[0].InstanceId' --output text)"

Targeting by tag rather than instance ID means the workflow keeps working when an instance is replaced. The curl with --retry at the end is the health gate: if the new container never becomes healthy, curl exits non-zero, the SSM command fails, and the workflow fails with it.

Use a GitHub environment named production with a required reviewer if you want a human approval gate, and with environment-scoped secrets so a workflow on another branch cannot reach production credentials at all.

Making the restart genuinely zero-downtime

docker compose up -d --no-deps api stops the old container before starting the new one, which is a gap of a few seconds. For a background service that is irrelevant; for a user-facing API it is a burst of 502s.

Three approaches, in increasing order of effort:

ApproachDowntimeComplexity
Recreate the container2 to 10 seconds of 502sNone — the default
Gunicorn reload inside a running containerNone, but code must be volume-mountedLow
Two containers, Nginx upstream swapNoneModerate
Compose scale up, health check, scale downNoneModerate

The scale approach is the most practical for a Compose stack. Start a second replica of the new version, wait for it to pass its healthcheck, then remove the old one. Nginx must resolve the upstream dynamically for this to work:

docker compose up -d --no-deps --scale api=2 --no-recreate api

# Wait for both replicas to report healthy
for i in $(seq 1 30); do
  UNHEALTHY=$(docker compose ps --format json api \
    | jq -r 'select(.Health != "healthy") | .Name' | wc -l)
  [ "$UNHEALTHY" -eq 0 ] && break
  sleep 2
done

docker compose up -d --no-deps --scale api=1 api

Also make sure your application shuts down gracefully. Docker sends SIGTERM and waits ten seconds before SIGKILL. Gunicorn handles SIGTERM correctly by default, finishing in-flight requests; application code that traps signals for cleanup needs to complete well inside that window, or raise it with stop_grace_period.

Rolling back

Because every build is tagged with a SHA, rollback is a deploy of an older tag. Make it a workflow you can trigger by hand rather than something requiring someone to remember the commands during an incident:

on:
  workflow_dispatch:
    inputs:
      tag:
        description: 'Image tag to roll back to (e.g. sha-4d81e02)'
        required: true

Rolling back code is easy; rolling back a migration is not. Write migrations so the previous version of the application still works against the new schema — add columns before using them, and drop them a release later. That expand-then-contract pattern is what makes rollback safe.

What to add once it works

  • A smoke test after deploy that exercises one real endpoint end to end, not just /healthz
  • Deployment notifications to Slack, including the tag and the person who triggered it
  • A scheduled workflow that rebuilds the image weekly so base image CVE patches land without a code change
  • Dependabot or Renovate for both application dependencies and Action versions
  • Pin third-party Actions to a commit SHA rather than a tag — tags are mutable and a compromised Action runs with your credentials

Why prefer SSM over SSH for deployment?

Because it removes the two riskiest parts of SSH-based deploys: an inbound port that must be open to a large range of GitHub IPs, and a long-lived private key stored in repository secrets. SSM works over an outbound connection from the instance, authenticates through IAM, and logs every command to CloudTrail.

Should I build the Docker image on the server instead?

No. Building on the server competes with your running application for CPU and memory, has no layer cache between deploys, and means a build failure happens after you have already started the deploy. Build in CI, push a tagged image, and let the server only pull.

How do I keep secrets out of the workflow logs?

Values from GitHub secrets are masked automatically, but anything you derive from them is not. Avoid echoing constructed connection strings, and never pass a secret as a command-line argument on a shared host where it is visible in the process list.

Do I need a staging environment?

You need somewhere that runs the same migration against the same database engine before production does. That can be a full staging stack or an ephemeral environment created per pull request. What matters is that a migration has run somewhere real before it runs against your production data.

What is a sensible healthcheck for the deploy gate?

One that checks the dependencies the new code needs — a database round trip and any critical external service — and returns non-200 when they are unavailable. A handler that returns a constant 200 will happily pass the gate on a build that cannot reach its database.