Architecture design for an enterprise MVP: C4 diagrams, a modular monolith and the seams that matter
Turning signed-off requirements into an architecture: C4 context and container diagrams, a Flask modular monolith on Google Cloud, asynchronous ERP integration with an outbox, and architecture decision records that explain the why.
Series · Part 3 of 7Enterprise MVP, end to end
- Enterprise MVP, end to end: from the first workshop to production
- Requirements gathering for an enterprise MVP: what to ask, what to write down
- Architecture design for an enterprise MVP: C4 diagrams, a modular monolith and the seams that matter
- Database design for an enterprise MVP: from domain model to the first PostgreSQL migration
- Developing an enterprise MVP with Flask and React: structure, workflow and tests
- CI/CD for an enterprise MVP: GitHub to Cloud Build to Cloud Deploy
- GKE or Cloud Run? Choosing where an enterprise MVP runs, and the go-live checklist
On this page · 7 sections
Architecture for an MVP is the art of making the few decisions that are expensive to change, and deliberately deferring the rest. The input is the signed scope document from part 2; the output is two diagrams, a handful of decision records, and an empty application deployed to a real environment.
This is part 3 of a series following an illustrative supplier onboarding portal from requirements to production.
Which architecture diagrams does an MVP need?
I use the C4 model and stop at level two. The context diagram is for the client’s architecture board: what the system is, who uses it, and what it talks to.
The container diagram is for the team: the deployable units and the managed services, with the technology on each box.
Two details in that diagram are deliberate. There is one container of our own: the React app is built into static files and served by the same Flask service that serves /api/*, with Cloud CDN caching the fingerprinted assets. The browser therefore talks to one origin, so there is no CORS to configure, session cookies just work, and the UI and API are always deployed as the same version. Everything else is a managed service the client’s operations team doesn’t have to patch.
Should an enterprise MVP use microservices?
No. The answer for an MVP is almost always a modular monolith: one deployable API whose code is divided into modules with hard boundaries. In Flask, each module is a Blueprint with its own routes, services and tables:
api/
app/
__init__.py # create_app(): config, extensions, register blueprints
suppliers/ # supplier profiles, invitations
applications/ # the onboarding application and its workflow
documents/ # uploads, signed URLs, document review
reviews/ # reviewer assignments and decisions
integrations/erp/ # outbox dispatch, ERP adapter
audit/ # append-only audit log
auth/ # OIDC login, magic links, roles
common/ # db session, errors, pagination
The rules that make it modular rather than a monolith with folders:
- A module only touches its own tables.
reviewsdoesn’t querydocumentstables; it callsdocuments.service.get_documents(application_id). - Cross-module calls go through a module’s service functions, never its models or routes.
- Side effects that other modules care about are recorded as domain events (“application approved”) rather than direct calls into those modules.
Those rules cost almost nothing on day one and are exactly the seams you would cut along if the ERP integration ever needed to become its own service.
How should an MVP integrate with an ERP or other external system?
Never synchronously inside the user’s request. The ERP will be slow, will be down for maintenance on the evening you go live, and will reject records for reasons nobody documented. The requirement from discovery was “within 15 minutes”, which gives us room to be resilient.
The pattern is a transactional outbox:
- When an application is approved, the API updates the application and inserts an
outboxrow (event = 'application.approved') in the same database transaction. Either both happen or neither does. - Cloud Scheduler calls an internal dispatch endpoint every minute. It picks up unsent outbox rows and enqueues one Cloud Tasks task per row.
- Cloud Tasks calls the API’s ERP sync handler, with exponential backoff and a retry limit. The handler is idempotent: it uses the application ID as the ERP’s external reference, so a retry never creates a duplicate vendor.
- After the final failed attempt, the application moves to Sync failed and appears on an admin dashboard with the ERP’s error message.
The ERP call itself sits behind an adapter interface. That’s the seam for the day the client changes ERP, and it also means tests and the dev environment use a fake ERP.
How do authentication and authorisation work?
Two kinds of user, two login paths, one session model:
- Internal users sign in with the client’s identity provider over OIDC (in Flask, Authlib does the heavy lifting). Roles come from IdP group claims mapped to procurement, compliance, finance and admin.
- Suppliers don’t have accounts in the client’s directory, so they sign in with a single-use, time-limited email link.
Both end with the same thing: a server-side session referenced by an HttpOnly, Secure, SameSite=Lax cookie. No access tokens in the browser, no token refresh logic in React. Authorisation is checked in the API on every request: role checks for internal users, and row-level ownership checks so a supplier can only ever load their own application.
Where do documents go?
In Cloud Storage, never in the database. The API issues short-lived V4 signed URLs, and the browser uploads directly to the bucket. That keeps 20 MB uploads off the API instances and the database small. The bucket has uniform bucket-level access, no public access, and object versioning; a document’s metadata and review status live in PostgreSQL.
How do you record architecture decisions?
As architecture decision records (ADRs): one short Markdown file per decision, in the repository next to the code, numbered and never edited once accepted (a later ADR supersedes an earlier one). Here’s one from this project:
# ADR-004: Run the application on Cloud Run, not GKE
Status: Accepted
Date: Week 3
## Context
Load is ~50 internal users and ~500 suppliers a year, with quarter-end peaks.
Availability target is 99.5% in business hours. The client's IT operations team
has no Kubernetes experience and will own the system after hand-over.
## Decision
Deploy the application container (Flask API plus the React build) to Cloud Run
in asia-south1. Use Cloud SQL for PostgreSQL,
Cloud Tasks for retries, and Cloud Scheduler for periodic work.
## Consequences
+ No cluster to patch, upgrade or secure; scales to near zero out of hours.
+ Cloud Deploy supports Cloud Run targets, so the pipeline is unchanged if we move.
- Background work must fit the request/response model (Cloud Tasks, Cloud Run jobs).
- Revisit if we need long-running workers, non-HTTP protocols or a service mesh.
The full reasoning behind that decision gets its own post in part 7.
What did we deliberately not build?
This list goes into the architecture document too, because the next architect will ask:
- No microservices. One API, modular inside.
- No Kubernetes. Nothing in the requirements needs it yet.
- No message bus. Cloud Tasks covers the one asynchronous integration; Pub/Sub can come when there are multiple consumers.
- No cache. PostgreSQL with sensible indexes is fast enough for this load; Redis would be one more thing to run.
- No multi-region. 99.5% in business hours is met by a regional deployment with managed services and tested backups.
Every item on that list is a decision, not an omission, and each one is reversible because of the seams above.
Next: Database design, from the domain model to the first migration.
Good MVP architecture is mostly a list of things you chose not to build, and the seams that let you build them later.