All writing

Architecture4 min read

Developing an enterprise MVP with Flask and React: structure, workflow and tests

How the build phase of an enterprise MVP runs: a monorepo with a Flask API and a React frontend, contract-first endpoints with generated TypeScript types, the approval workflow as an explicit state machine, tests that hit a real PostgreSQL, and a working demo every week.

Series · Part 5 of 7Enterprise MVP, end to end
  1. Enterprise MVP, end to end: from the first workshop to production
  2. Requirements gathering for an enterprise MVP: what to ask, what to write down
  3. Architecture design for an enterprise MVP: C4 diagrams, a modular monolith and the seams that matter
  4. Database design for an enterprise MVP: from domain model to the first PostgreSQL migration
  5. Developing an enterprise MVP with Flask and React: structure, workflow and tests
  6. CI/CD for an enterprise MVP: GitHub to Cloud Build to Cloud Deploy
  7. GKE or Cloud Run? Choosing where an enterprise MVP runs, and the go-live checklist
On this page · 7 sections
  1. How should the repository be organised?
  2. How is the Flask app structured?
  3. What does contract-first API development look like?
  4. How do you implement an approval workflow?
  5. How is the React app structured?
  6. How do you test an enterprise MVP?
  7. How should an MVP team work week to week?

By the time development starts, the hard thinking should be done: the scope is signed, the architecture is decided, the schema exists as a migration. The build phase is about turning that into working software at a steady, visible pace, with structure that keeps week eight as productive as week two.

This is part 5 of a series following an illustrative supplier onboarding portal from requirements to production. The design it implements is in part 3 (architecture) and part 4 (database).

How should the repository be organised?

One repository for everything that ships together:

supplier-portal/
  api/                      # Flask API
    app/                    # one package per module (suppliers, applications, documents, ...)
    migrations/             # Alembic
    tests/
    pyproject.toml
  web/                      # React + TypeScript (Vite)
    src/
      features/             # applications/, documents/, reviews/, admin/
      api/                  # generated types + typed client
      components/           # shared UI
  Dockerfile                # multi-stage: build web/, then the Flask image that serves it
  deploy/                   # skaffold.yaml, clouddeploy.yaml, Cloud Run service manifests
  infra/                    # Terraform: Cloud SQL, buckets, queues, IAM
  docs/adr/                 # architecture decision records
  cloudbuild.yaml
  compose.yaml              # local PostgreSQL, fake ERP, fake GCS

A monorepo means one pull request can add a column, the endpoint that exposes it, the form field that edits it, and the Cloud Run setting it needs, reviewed together and deployed together. For a team of four to six, the coordination saved is worth far more than any build-time cost.

Locally, docker compose up starts PostgreSQL, a fake ERP and a fake Cloud Storage server, and every engineer has a working system in minutes. If onboarding a new developer takes more than an hour, that’s a bug.

How is the Flask app structured?

An application factory, one Blueprint per module, and three layers inside each module:

# api/app/__init__.py
from flask import Flask
from .common import db, migrate, api_spec, errors
from . import suppliers, applications, documents, reviews, audit, auth
from .integrations import erp


def create_app(config: str = "app.config.Production") -> Flask:
    app = Flask(__name__)
    app.config.from_object(config)

    db.init_app(app)
    migrate.init_app(app, db)
    api_spec.init_app(app)          # flask-smorest: OpenAPI from marshmallow schemas
    errors.register(app)            # one JSON error shape for every error

    for module in (suppliers, applications, documents, reviews, audit, auth, erp):
        api_spec.register_blueprint(module.blp)
    return app

Inside each module:

  • routes.py: HTTP only. It parses the request, checks authorisation, calls a service function and serialises the result. No business logic.
  • service.py: the business rules, and the only place that commits a transaction. Other modules call these functions.
  • models.py: SQLAlchemy models for the module’s own tables.

What does contract-first API development look like?

Before an endpoint is implemented, its request and response shapes are agreed in a small pull request that adds only the marshmallow schemas and a stub route. flask-smorest turns those into an OpenAPI document, which CI publishes, and the frontend generates TypeScript types from it with openapi-typescript.

The payoff is that the API and the UI can be built in parallel against the agreed contract, and any breaking change to the API fails the React type-check in the same pull request, not in a reviewer’s browser three days later.

How do you implement an approval workflow?

As an explicit state machine. The states come straight from the diagram in part 2, and the transitions are data, not scattered if statements:

# api/app/applications/workflow.py
from enum import StrEnum


class Status(StrEnum):
    DRAFT = "draft"
    SUBMITTED = "submitted"
    IN_REVIEW = "in_review"
    CHANGES_REQUESTED = "changes_requested"
    APPROVED = "approved"
    REJECTED = "rejected"
    SYNCED_TO_ERP = "synced_to_erp"
    SYNC_FAILED = "sync_failed"


TRANSITIONS: dict[tuple[Status, str], Status] = {
    (Status.DRAFT, "submit"): Status.SUBMITTED,
    (Status.SUBMITTED, "start_review"): Status.IN_REVIEW,
    (Status.IN_REVIEW, "request_changes"): Status.CHANGES_REQUESTED,
    (Status.CHANGES_REQUESTED, "submit"): Status.SUBMITTED,
    (Status.IN_REVIEW, "approve"): Status.APPROVED,
    (Status.IN_REVIEW, "reject"): Status.REJECTED,
    (Status.APPROVED, "erp_synced"): Status.SYNCED_TO_ERP,
    (Status.APPROVED, "erp_failed"): Status.SYNC_FAILED,
    (Status.SYNC_FAILED, "retry_sync"): Status.APPROVED,
}


class InvalidTransition(Exception):
    pass


def next_status(current: Status, event: str) -> Status:
    try:
        return TRANSITIONS[(current, event)]
    except KeyError:
        raise InvalidTransition(f"Cannot '{event}' an application that is {current}") from None

State changes happen in exactly one place. The service function records the reviewer’s decision, applies the transition, writes the audit row, and adds the outbox event, all in one transaction:

# api/app/reviews/service.py
EVENT_FOR = {Decision.APPROVE: "approve", Decision.REJECT: "reject",
             Decision.REQUEST_CHANGES: "request_changes"}


def record_decision(app_id, reviewer, team, decision, comment, expected_version):
    application = applications.service.get_for_update(app_id, expected_version)  # 409 if stale
    db.session.add(ReviewDecision(application_id=app_id, reviewer_id=reviewer.id,
                                  team=team, decision=decision, comment=comment))

    before = application.status
    # Approval needs both compliance and finance; anything else takes effect immediately.
    if decision is not Decision.APPROVE or other_team_approved(app_id, team):
        application.status = next_status(application.status, EVENT_FOR[decision])
        if application.status is Status.APPROVED:
            outbox.add("application.approved", aggregate_id=app_id)

    audit.record(actor=reviewer, entity=application, action=f"review.{decision}",
                 before={"status": before},
                 after={"status": application.status, "team": team, "comment": comment})
    db.session.commit()
    return application

Because the transition table is plain data, the workflow’s unit tests are a loop over every (status, event) pair: allowed pairs move to the right state, and every other pair raises InvalidTransition. That one test catches most workflow regressions.

Here is the whole approval path end to end, including the asynchronous ERP sync from part 3. Each write to Postgres is one transaction, and the ERP call is idempotent, keyed by the application ID:

Approval and ERP sync sequenceA supplier submits, compliance and finance reviewers approve, each step is written to Postgres with an audit entry, and a dispatcher queues an ERP sync task that creates the vendor and records it as synced.ERPTasksPostgresAPIReviewersSupplierERPTasksPostgresAPIReviewersSupplierdispatcher, every minutesubmit1submitted + audit2approve (compliance)3decision + audit4approve (finance)5approved + audit + outbox6enqueue erp-sync7erp-sync (retries)8create vendor9vendor ID10synced_to_erp + audit11

How is the React app structured?

Vite and TypeScript, organised by feature rather than by file type. Each feature folder owns its pages, components and data hooks. Three libraries do most of the work:

  • TanStack Query for server state: caching, refetching, and invalidating after a mutation.
  • React Hook Form with Zod for the long supplier questionnaire, with validation that mirrors the API’s.
  • openapi-fetch with the generated types, so every API call is typed end to end.

The reviewer’s decision buttons show how the pieces fit, including the conflict case when two reviewers act at once:

// web/src/features/reviews/DecisionButtons.tsx
export function DecisionButtons({ application }: { application: Application }) {
  const queryClient = useQueryClient();
  const [comment, setComment] = useState('');
  const decide = useMutation({
    mutationFn: async (decision: DecisionRequest['decision']) => {
      const { data, error, response } = await api.POST('/api/applications/{id}/decisions', {
        params: { path: { id: application.id } },
        body: { decision, comment, version: application.version },
      });
      if (response.status === 409) throw new StaleDataError();
      if (error) throw error;
      return data;
    },
    onSettled: () => queryClient.invalidateQueries({ queryKey: ['application', application.id] }),
  });

  if (decide.error instanceof StaleDataError) {
    return <Notice>Someone else just updated this application. It has been reloaded; please review again.</Notice>;
  }
  // ...comment field bound to setComment, and approve / request changes / reject buttons
  // that call decide.mutate(...), disabled while decide.isPending
}

Sessions are cookie-based (part 3), so the API client never handles tokens; it just sends credentials: 'include' to the same origin.

How do you test an enterprise MVP?

A test pyramid, weighted towards fast tests, with one firm rule: database tests use a real PostgreSQL.

  • Unit tests (pytest): the state machine, validation rules, the ERP payload mapper. Pure functions, milliseconds each.
  • Integration tests (pytest + Testcontainers): a PostgreSQL container starts once per test session, Alembic migrates it to head, and each test runs in a transaction that is rolled back. These cover services and routes through Flask’s test client, including authorisation: every endpoint has a test proving a supplier can’t read another supplier’s application.
  • Frontend tests (Vitest + React Testing Library): forms, conditional questionnaire logic, error states.
  • End-to-end smoke tests (Playwright): five or six critical journeys, run against staging after each deploy, not on every commit.

SQLite instead of PostgreSQL would be faster, but it doesn’t enforce the same constraints, has no enums, no partial indexes and different jsonb behaviour. A test suite that passes on a database you don’t run in production is telling you very little.

How should an MVP team work week to week?

  • Trunk-based development. Short-lived branches, merged within a day or two. Pull requests need one approving review and green checks from Cloud Build.
  • Feature flags for unfinished work, so incomplete features can merge without being visible. For an MVP, a flags table read at request time is enough; there’s no need for a vendor.
  • A demo every Friday, from staging. Not from a laptop, and not from slides. Demoing from staging proves the pipeline works every week and keeps “it works on my machine” out of the client relationship.
  • Scope conversations happen at the demo. When the client sees what exists, “not this phase” is an easy conversation, especially with the signed out-of-scope list from part 2.

Next: CI/CD with GitHub, Cloud Build and Cloud Deploy, getting every merged commit safely to production.

The best predictor of an on-time MVP is a working demo from a real environment every single week.