← All writing

Leadership  ·   ·  3 min read

Architecture before AI: a CTO’s checklist for putting LLMs into production

Most teams choose a model first and discover their constraints later. Here is the order I work in instead — and the patterns that make GenAI systems production-grade.

Every engineering leader I know is being asked some version of “what’s our AI strategy?” The most common failure mode I see isn’t choosing the wrong model. It’s choosing a model first, and discovering the real constraints — cost, latency, trust, failure modes — after a demo has already set expectations.

Here’s the order I work in, whether at FinMoon AI or when reviewing someone else’s system.

1. Start with the architectural characteristics

Before any model discussion, write down which qualities matter and rank them. For an AI feature, the list usually includes:

  • Accuracy and trust: what does a wrong answer cost? A mis-tagged photo and a wrong tax recommendation are not the same risk.
  • Latency: is this interactive, or can it run in the background?
  • Cost per unit of value: not cost per token — cost per useful outcome.
  • Explainability: does a user, auditor or regulator need to see why?
  • Data sensitivity: what can leave your boundary, and to whom?

If you can’t rank these, you aren’t ready to pick a model. You’re ready to talk to users.

2. Map the business constraints

Budget, team skills, compliance, and time-to-market shape the architecture as much as any technical property. A three-person startup and a regulated enterprise can adopt the same model and need completely different systems around it.

3. Decide where the model sits in the process

The most important design decision is the boundary between deterministic logic and model reasoning. My default: the process is code; the model works inside the steps. Workflow, validation, permissions and business rules stay deterministic and testable. The model handles what only it can — reading unstructured input, reasoning over context, generating language.

This one choice makes the whole system more predictable, cheaper, and far easier to evaluate.

4. Treat evaluation as a first-class component

You wouldn’t ship a service without tests. Don’t ship a model-backed feature without evals:

  • Golden datasets of real inputs with expected outcomes, versioned alongside code.
  • LLM-as-a-judge for scoring open-ended outputs against a rubric — calibrated against human judgement, never trusted blindly.
  • Regression gates in CI, so a prompt or model change can’t silently degrade quality.
  • Fitness functions that continuously check the characteristics you ranked in step one.

5. Put a gateway in front of every model

An AI gateway between your application and model providers pays for itself quickly: one place for routing, fallbacks, rate limiting, caching, cost attribution, redaction of sensitive data, and swapping providers without touching application code. Models change quarterly. Your architecture shouldn’t have to.

6. Design the human in the loop on purpose

Decide explicitly where humans review, override or approve, and design the interface for it. Good human-in-the-loop design isn’t a fallback for weak models; it’s how high-stakes systems earn trust while they gather the data to improve.

7. Productionise like it matters

Observability (traces through every prompt and tool call), governance (who can change prompts, and how are changes reviewed), and security (prompt injection, data leakage, tool permissions) belong in the first architecture diagram — not in a “phase two” that never comes.

The short version

AI is a component, not an architecture.

The teams that win with GenAI aren’t the ones with access to the best model — everyone has that. They’re the ones who did the unglamorous architectural thinking first, and built the system that makes any model useful.