CI/CD for an enterprise MVP: GitHub to Cloud Build to Cloud Deploy
A complete delivery pipeline on Google Cloud: Cloud Build triggers from GitHub for pull requests and merges, images in Artifact Registry, Cloud Deploy releases promoted from dev to staging to production with an approval gate, database migrations as a predeploy hook, and one-command rollback.
Series · Part 6 of 7Enterprise MVP, end to end
- Enterprise MVP, end to end: from the first workshop to production
- Requirements gathering for an enterprise MVP: what to ask, what to write down
- Architecture design for an enterprise MVP: C4 diagrams, a modular monolith and the seams that matter
- Database design for an enterprise MVP: from domain model to the first PostgreSQL migration
- Developing an enterprise MVP with Flask and React: structure, workflow and tests
- CI/CD for an enterprise MVP: GitHub to Cloud Build to Cloud Deploy
- GKE or Cloud Run? Choosing where an enterprise MVP runs, and the go-live checklist
On this page · 8 sections
- What does the pipeline look like end to end?
- How do you connect GitHub to Cloud Build?
- What goes in cloudbuild.yaml?
- What does the Dockerfile look like?
- How is the Cloud Deploy pipeline defined?
- How do you run database migrations in a Cloud Deploy pipeline?
- How do promotion, approval and rollback work?
- How do you keep secrets and access safe in the pipeline?
The pipeline is the part of an MVP that clients never see and always feel. When it works, every Friday demo runs on code merged that morning, and go-live is just one more promotion. When it doesn’t, every release is a ceremony and every bug fix waits for one.
This is part 6 of a series following an illustrative supplier onboarding portal from requirements to production. It builds on the repository layout and tests from part 5. As I said in part 1, the pipeline is built in the first sprint, not the last.
What does the pipeline look like end to end?
The division of labour is the important design decision:
- Cloud Build does continuous integration. It checks out the commit, runs every test, builds one container image, and pushes it to Artifact Registry tagged with the commit SHA.
- Cloud Deploy does continuous delivery. It takes that exact image and promotes it through environments, recording who promoted what and when, with approvals and rollbacks built in.
The image is built once. Staging and production run byte-for-byte the same artefact that passed the tests, which is the property that makes a staging sign-off mean something.
How do you connect GitHub to Cloud Build?
Through a Cloud Build repository connection to GitHub (installed once as a GitHub App), and two triggers:
- Pull request trigger: on any PR targeting
main, runcloudbuild-pr.yaml: linting, type-checks, the Python unit and integration tests, and the React tests. Its status is a required check in GitHub’s branch protection, so nothing merges red. - Main trigger: on a push to
main, runcloudbuild.yaml: the same tests again on the merged code, then build, push and create a release.
Running the tests again after merge isn’t redundant. Two PRs can each pass on their own and fail together.
What goes in cloudbuild.yaml?
# cloudbuild.yaml (main branch)
substitutions:
_REGION: asia-south1
_IMAGE: asia-south1-docker.pkg.dev/supplier-portal-build/apps/app
steps:
# A throwaway PostgreSQL for integration tests, reachable as "pg" on the cloudbuild network.
- id: postgres
name: gcr.io/cloud-builders/docker
args: [run, -d, --name=pg, --network=cloudbuild, -e, POSTGRES_PASSWORD=test, postgres:17]
- id: api-tests
name: python:3.13-slim
dir: api
env: [DATABASE_URL=postgresql+psycopg://postgres:test@pg:5432/postgres]
entrypoint: bash
args: [-c, pip install -q -r requirements-dev.lock && flask db upgrade && pytest -q]
- id: web-tests
name: node:22
dir: web
entrypoint: bash
args: [-c, npm ci && npm run typecheck && npm test -- --run]
waitFor: ['-'] # run in parallel with the API tests
- id: build
name: gcr.io/cloud-builders/docker
args: [build, -t, '${_IMAGE}:$SHORT_SHA', .]
waitFor: [api-tests, web-tests]
- id: push
name: gcr.io/cloud-builders/docker
args: [push, '${_IMAGE}:$SHORT_SHA']
- id: release
name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
entrypoint: gcloud
args:
- deploy
- releases
- create
- rel-$SHORT_SHA
- --delivery-pipeline=supplier-portal
- --region=${_REGION}
- --source=deploy
- --images=app=${_IMAGE}:$SHORT_SHA
serviceAccount: projects/supplier-portal-build/serviceAccounts/cloudbuild-ci@supplier-portal-build.iam.gserviceaccount.com
options:
logging: CLOUD_LOGGING_ONLY
A few things worth pointing out:
- Locally, the integration tests use Testcontainers. In Cloud Build, a PostgreSQL container started as the first step plays the same role; every step runs on the
cloudbuildDocker network, so the tests reach it atpg:5432. - The Artifact Registry repository has immutable tags turned on, so
app:3f9c2abcan never be overwritten to point at something else. - The build runs as a dedicated service account that can push images and create releases, and nothing else. It has no access to any environment’s database or secrets.
What does the Dockerfile look like?
As decided in part 3, the React build ships inside the Flask image, so one image is the whole application:
# Stage 1: build the React app
FROM node:22-alpine AS web
WORKDIR /web
COPY web/package.json web/package-lock.json ./
RUN npm ci
COPY web/ ./
RUN npm run build
# Stage 2: the runtime
FROM python:3.13-slim
ENV PYTHONDONTWRITEBYTECODE=1 PYTHONUNBUFFERED=1
WORKDIR /srv
COPY api/requirements.lock ./
RUN pip install --no-cache-dir -r requirements.lock
COPY api/ ./
COPY --from=web /web/dist ./app/static
RUN useradd --uid 10001 app
USER app
CMD exec gunicorn --bind :$PORT --workers 1 --threads 8 --timeout 0 "app:create_app()"
Flask serves the fingerprinted assets with long cache headers (Cloud CDN does the rest) and returns index.html for any non-API path so client-side routing works. The image contains no Node.js, no build tools and no tests, and it runs as a non-root user.
How is the Cloud Deploy pipeline defined?
Three files in deploy/: the Skaffold config that says what to deploy, the Cloud Run service manifests for each environment, and the delivery pipeline that says where and in what order.
skaffold.yaml maps one profile to each environment’s manifest:
apiVersion: skaffold/v4beta7
kind: Config
metadata:
name: supplier-portal
profiles:
- name: dev
manifests: { rawYaml: [run/dev.yaml] }
- name: staging
manifests: { rawYaml: [run/staging.yaml] }
- name: prod
manifests: { rawYaml: [run/prod.yaml] }
deploy:
cloudrun: {}
run/prod.yaml is a standard Cloud Run service definition. The image is the placeholder app, which --images in the release step replaces with the exact image built from the commit:
apiVersion: serving.knative.dev/v1
kind: Service
metadata:
name: supplier-portal
annotations:
run.googleapis.com/ingress: internal-and-cloud-load-balancing
spec:
template:
metadata:
annotations:
autoscaling.knative.dev/minScale: '1'
autoscaling.knative.dev/maxScale: '10'
spec:
serviceAccountName: portal-runtime@supplier-portal-prod.iam.gserviceaccount.com
containerConcurrency: 40
timeoutSeconds: 60
containers:
- image: app
resources:
limits: { cpu: '1', memory: 1Gi }
env:
- name: APP_ENV
value: prod
- name: DB_INSTANCE # reached via the Cloud SQL Python Connector with IAM auth
value: supplier-portal-prod:asia-south1:portal-db
- name: OIDC_CLIENT_SECRET
valueFrom:
secretKeyRef: { name: oidc-client-secret, key: latest }
The dev and staging files differ in project, scaling (minScale: '0' in dev, so it costs nothing overnight) and secrets. Environment differences live in these three small files and nowhere else.
clouddeploy.yaml defines the pipeline, the targets, and the automation:
apiVersion: deploy.cloud.google.com/v1
kind: DeliveryPipeline
metadata:
name: supplier-portal
serialPipeline:
stages:
- targetId: dev
profiles: [dev]
strategy: &migrate-then-deploy
standard:
predeploy:
tasks:
- type: container
image: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
command: [/bin/bash]
args:
- -c
- |
set -euo pipefail
REGION="${CLOUD_RUN_LOCATION##*/}"
SHA="${CLOUD_DEPLOY_RELEASE#rel-}"
gcloud run jobs update db-migrate --project="$CLOUD_RUN_PROJECT" --region="$REGION" \
--image="asia-south1-docker.pkg.dev/supplier-portal-build/apps/app:$SHA"
gcloud run jobs execute db-migrate --project="$CLOUD_RUN_PROJECT" --region="$REGION" --wait
- targetId: staging
profiles: [staging]
strategy: *migrate-then-deploy
- targetId: prod
profiles: [prod]
strategy: *migrate-then-deploy
---
apiVersion: deploy.cloud.google.com/v1
kind: Target
metadata:
name: prod
requireApproval: true
run:
location: projects/supplier-portal-prod/locations/asia-south1
executionConfigs:
- usages: [RENDER, PREDEPLOY, DEPLOY]
serviceAccount: deployer@supplier-portal-prod.iam.gserviceaccount.com
---
# dev and staging targets look the same, in their own projects, without requireApproval.
apiVersion: deploy.cloud.google.com/v1
kind: Automation
metadata:
name: supplier-portal/promote-dev-to-staging
serviceAccount: deployer@supplier-portal-build.iam.gserviceaccount.com
selector:
targets:
- id: dev
rules:
- promoteReleaseRule:
id: promote-to-staging
wait: 10m
destinationTargetId: '@next'
gcloud deploy apply --file=deploy/clouddeploy.yaml --region=asia-south1 registers all of it. Like everything else in the pipeline, it’s versioned in the repository and changed through pull requests.
How do you run database migrations in a Cloud Deploy pipeline?
As a predeploy hook, the job that runs in each environment before the new version is deployed. The hook doesn’t run the migration itself; it points a Cloud Run job called db-migrate at the new image and executes it:
- The
db-migratejob is created once per environment by Terraform, with the same service account, database settings and secrets as the service. Its command isflask db upgrade. - The hook reads the release name (
rel-3f9c2ab) from Cloud Deploy’sCLOUD_DEPLOY_RELEASEvariable to find the image tag, updates the job’s image, and runs it with--wait. If the migration fails, the hook fails, and the deploy never happens; the old version keeps serving traffic on the old schema. - Migrations run once per environment, not once per instance, so there is no race between instances on start-up.
This only works because of the expand-then-contract rule from part 4: between the migration finishing and the new revision taking traffic, the old code is running against the new schema, and it must keep working.
How do promotion, approval and rollback work?
- Dev: every merged commit is deployed automatically, a few minutes after merge.
- Staging: the automation promotes from dev to staging 10 minutes after a successful dev rollout. A postdeploy hook runs the Playwright smoke tests from part 5 against staging, and Friday’s demo runs here.
- Production:
requireApproval: truemeans the promotion waits for someone with the Cloud Deploy approver role (the client’s release manager, not the developer who wrote the change) to approve it in the console or withgcloud deploy rollouts approve. The approval and the approver are recorded against the release. - Rollback:
gcloud deploy targets rollback prod --delivery-pipeline=supplier-portal --region=asia-south1redeploys the previous successful release. Because migrations are backward-compatible, rolling back the code doesn’t require rolling back the schema, which is exactly why we design them that way.
How do you keep secrets and access safe in the pipeline?
- Separate projects per environment (
-dev,-staging,-prod) plus a-buildproject for the registry and pipeline. Production data is in a project that developers can’t open. - A service account per role: the CI account can push images and create releases; each environment’s deployer account can deploy to that environment only; the runtime account can read that environment’s secrets and connect to its database, and nothing else.
- No keys anywhere. Nothing uses downloaded service-account key files. GitHub never holds Google Cloud credentials, because the build runs inside Google Cloud. The runtime reads secrets from Secret Manager, and Cloud SQL uses IAM database authentication.
- Everything is audited. Cloud Build logs every build, Cloud Deploy records every release, promotion, approval and rollback, and Cloud Audit Logs record who changed IAM. That trail is what makes a security review short.
Next: GKE or Cloud Run?, choosing where the application runs, and the go-live checklist.
A deploy should be so boring that nobody remembers it. Make the pipeline do the remembering.