GKE or Cloud Run? Choosing where an enterprise MVP runs, and the go-live checklist
How to decide between serverless Cloud Run and GKE workloads on Google Cloud: traffic shape, cost model, operational burden, team skills and hard limits, with a decision flowchart, both deployment topologies, how the Cloud Deploy pipeline changes, and the checklist I use before go-live.
Series · Part 7 of 7Enterprise MVP, end to end
- Enterprise MVP, end to end: from the first workshop to production
- Requirements gathering for an enterprise MVP: what to ask, what to write down
- Architecture design for an enterprise MVP: C4 diagrams, a modular monolith and the seams that matter
- Database design for an enterprise MVP: from domain model to the first PostgreSQL migration
- Developing an enterprise MVP with Flask and React: structure, workflow and tests
- CI/CD for an enterprise MVP: GitHub to Cloud Build to Cloud Deploy
- GKE or Cloud Run? Choosing where an enterprise MVP runs, and the go-live checklist
On this page · 8 sections
“Should we run this on Kubernetes?” comes up in almost every architecture review I’ve been part of. Often the honest answer is “not yet”. Sometimes it’s “yes, and here’s why”. The answer should come from the requirements, not from what the team wants on its CV.
This is the final part of a series following an illustrative supplier onboarding portal from requirements to production. The pipeline from part 6 deploys to Cloud Run. This post explains that choice, when I would make the other one, and what it takes to go live.
What are the options on Google Cloud?
For a containerised application there are three realistic choices:
- Cloud Run: serverless containers. You give it an image; it runs instances in response to requests and scales them, including to zero. It offers services (HTTP), jobs (run-to-completion tasks) and worker pools (always-on background workers).
- GKE Autopilot: managed Kubernetes where Google runs the nodes. You write Kubernetes manifests; you’re billed for what your pods request, not for nodes.
- GKE Standard: managed Kubernetes where you choose and pay for the node pools. It gives you the most control and the most responsibility.
All three run the same container image. That’s the most important fact in this whole decision, because it means the choice can be changed later.
How do Cloud Run and GKE compare?
| Criterion | Cloud Run | GKE (Autopilot / Standard) |
|---|---|---|
| Unit you manage | A service, job or worker pool | A cluster, plus deployments, services, ingress, autoscalers |
| Scaling | Per request, automatic, can scale to zero | Horizontal Pod Autoscaler; cluster or node autoscaling; rarely to zero |
| Billing model | Per request (CPU billed only during requests) or per instance | Autopilot: per pod resource request. Standard: per node. Both pay a cluster management fee of $0.10/hour |
| Protocols | HTTP/1.1, HTTP/2, gRPC, WebSockets | Anything, including raw TCP and UDP |
| Limits per instance | Up to 8 vCPU and 32 GiB; requests up to 60 minutes; up to 1,000 concurrent requests | Whatever the node type offers |
| Background work | Jobs (tasks up to 7 days), worker pools, Cloud Tasks | Deployments, Jobs, CronJobs, any long-running process |
| State | Stateless; state lives in managed services | StatefulSets and persistent volumes if you need them |
| Extensibility | Sidecars, GPUs (NVIDIA L4), Direct VPC egress | Operators, DaemonSets, service mesh, custom schedulers, the whole Kubernetes ecosystem |
| Operational burden | Very low: no nodes, no upgrades, no cluster security | Autopilot: moderate. Standard: high. Upgrades, policies, add-ons, capacity |
| Skills needed | Containers and Cloud Run configuration | Kubernetes, which takes real experience to operate well |
How do you decide?
I walk the requirements through a short set of questions. The first “yes” usually decides it.
Applied to this project, using the numbers captured in part 2:
- Traffic: ~50 internal users in business hours, ~500 suppliers a year, quarter-end peaks. That’s spiky and mostly idle, which is the shape Cloud Run bills best.
- Workload: a stateless HTTP API and a static frontend in one image. Background work (ERP sync, migrations) fits Cloud Tasks and Cloud Run jobs.
- Availability: 99.5% in business hours, regional. Cloud Run is regional by default, with no node failures to plan around.
- Team: the client’s IT operations team will own the system and has no Kubernetes experience. Handing them a cluster would hand them a new discipline.
Every answer points to Cloud Run, which is what ADR-004 in part 3 recorded.
What does the Cloud Run deployment look like?
The settings that matter in production:
- Minimum instances of 1 in production removes cold starts for the first user of the morning. It costs one idle instance; dev and staging scale to zero.
- Concurrency of 40 with 1 vCPU suits a Flask app on Gunicorn threads, which spends most of its time waiting on PostgreSQL. Load testing decides the real number.
- Ingress restricted to the load balancer, so the only way in is through Cloud Armor and the CDN.
- Request-based billing for the service: CPU is billed only while handling requests. Instance-based billing is the better choice if the service does background work between requests.
When should you choose GKE instead?
When the workload needs something Cloud Run deliberately doesn’t do. The cases I see in practice:
- Non-HTTP or long-lived connections at scale: MQTT brokers, game servers, custom TCP or UDP protocols.
- Stateful components you have to run yourself: a search cluster, a message broker or a vector database that has no suitable managed service, needing persistent volumes and stable identities.
- Kubernetes-native tooling: operators, DaemonSets for node-level agents, service mesh policies across dozens of services, custom schedulers, or batch systems built on Kubernetes primitives.
- Many services with steady, high utilisation. When dozens of services run busy around the clock, bin-packing them onto nodes (with committed-use discounts) is often cheaper than per-request pricing. Model it with real numbers; don’t assume it.
- An existing platform. If the client already runs a well-operated GKE platform with security policies, observability and an on-call team, deploying onto it is usually the right call, even for a small app. The operational cost is already paid.
If GKE is the answer, I start with Autopilot unless there’s a specific need for node-level control. It removes node management and bills by pod requests, which takes away most of what makes Kubernetes expensive for small teams.
For comparison, here is the same application on GKE:
It’s not dramatically more complex on a diagram. The difference is everything the diagram doesn’t show: cluster upgrades, network policies, Workload Identity bindings, pod security standards, resource requests that have to be tuned, and someone on call who understands all of it.
Does the CI/CD pipeline change if you move to GKE?
Very little, which is why the decision is safe to make early. Cloud Deploy supports both target types:
- Add a target with
gke: { cluster: projects/…/locations/asia-south1/clusters/portal }instead ofrun: { location: … }. - Add a Skaffold profile that renders Kubernetes manifests (a Deployment, a Service, an HTTPRoute) with
deploy: kubectl: {}instead ofcloudrun: {}. - The predeploy migration hook can stay as it is, or run as a Kubernetes Job on the cluster.
The image, the Cloud Build configuration, the approval gates and the rollback command stay the same. Moving runtimes becomes a planned change of one stage, not a re-platforming project.
What’s on the go-live checklist?
Runtime chosen and pipeline working, this is the list I walk through with the client’s IT operations and security teams in the final two weeks. Nothing goes live until every line has an owner and a tick:
Security
- Pen test completed; high and critical findings fixed and retested
- Cloud Armor policy on the load balancer (OWASP rules, rate limiting on login and magic-link endpoints)
- No service-account keys exist; IAM reviewed against least privilege
- Secrets only in Secret Manager; rotation procedure written
Data
- Cloud SQL high availability on, automated backups and point-in-time recovery enabled
- A restore actually tested into a scratch instance, with the time measured against the 4-hour RTO
- Spreadsheet data migrated, reconciled with the business, and signed off
Reliability
- Load test at 3× the expected quarter-end peak, with p95 latency within target
- Uptime checks and alerting on error rate, latency, Cloud Tasks backlog and outbox age
- Dashboards for the service, database and queues, and structured logs with request IDs
- Rollback rehearsed in staging with
gcloud deploy targets rollback
Operations
- Runbook: deploy, roll back, re-drive failed ERP syncs, restore from backup, rotate a secret
- On-call rota and escalation path agreed for the first month
- Hyper-care plan: daily check-in with the business for the first two weeks
- Hand-over sessions with the client’s team, using the ADRs and data dictionary
Business
- User acceptance testing signed off against the acceptance criteria from part 2
- Training done for procurement, compliance and finance; supplier help page live
- Success measure baseline recorded (current median onboarding time) so the improvement can be proven
The restore test is the line teams most often want to skip, and the one I never let go. A backup you haven’t restored is a hope, not a backup.
The series in one paragraph
Requirements written as numbers decided the architecture. The architecture’s seams decided the schema and the code structure. Expand-then-contract migrations made the pipeline safe to run on every merge, and the pipeline made the runtime choice reversible. Each phase left an artefact behind (scope, ADRs, ERD, contract, pipeline, checklist), so the system can be understood and owned by people who weren’t in the room. That’s what an enterprise MVP needs to be: small, and built to be owned.
If you’re starting from the beginning, part 1 has the map. If you’re weighing a cloud migration rather than a new build, migrating legacy systems to Google Cloud and cutting the bill by 40% covers the other side of the same decisions.
Choose the runtime your requirements need and your team can operate, and keep the door open to the other one.