Posts AI Engineering Interview Prep — Part 3: Production Trade-offs & System Design
Post
Cancel

AI Engineering Interview Prep — Part 3: Production Trade-offs & System Design

System architecture abstract — production trade-offs and reliability.

Series: ← Part 1: Foundations · ← Part 2: RAG & LLMOps · Part 3 of 3

Who this is for: Forward Deployed AI Engineers facing system design rounds — where interviewers ask you to reduce cost, handle failures, migrate embeddings safely, and justify why you need an LLM at all.

Part 3 closes the series with production trade-offs, a real-world migration scenario, and rapid-fire model answers. For quantization deep-dive, see LLM Quantization Explained. For multi-GPU serving, see Serving a large LLM with vLLM + Ray.


Q1: How do you reduce token usage?

Short answer (30 seconds)

I measure token usage by component first, then reduce unnecessary context: summarize conversation history instead of sending full chat, retrieve fewer but better documents via reranking, compress or extractively trim RAG context, set output limits, and cache embeddings and repeated responses.

Deep explanation

1. Reduce conversation history

1
2
BAD:  Send last 100 messages verbatim
GOOD: Recent 5 messages + rolling summary + retrieved memories

2. Retrieve fewer, better documents

1
Retrieve 30 → Rerank → Top 5   (not Top 20 straight to LLM)

3. Compress context

  • Contextual compression (LLM or extractive summarization per chunk)
  • Smaller chunks with parent-document references
  • Extractive retrieval (only relevant sentences, not full chunks)

4. Output limits

1
max_tokens = 512   (enforce at API level)

5. Cache aggressively

Cache targetBenefit
EmbeddingsSkip re-embedding identical text
Frequent queriesSkip full RAG + LLM pipeline
Prompt prefixesPrefix caching on supported APIs
Retrieval resultsSame query + same index version

FDE / production angle

Build a token breakdown dashboard: system prompt vs history vs RAG vs user input vs output. Most cost surprises come from unbounded history or over-retrieval. This pairs directly with Part 1’s token budgeting discussion.

Follow-up questions to expect

  • How much quality do you lose when summarizing history?
  • When is caching unsafe (stale answers)?
  • How do you cost-optimize multi-step agent workflows?

Interview takeaway

“Measure tokens by component, then cut waste — summarized history, rerank-then-trim retrieval, output caps, and caching — before switching to a smaller model.”


Q2: When should you quantize a model?

Short answer (30 seconds)

Quantize when VRAM is the bottleneck, you need high-volume inference at lower cost, and you have benchmarked that quality degradation is acceptable. Always compare FP16 vs INT8 vs INT4 on your actual tasks — not generic benchmarks.

Deep explanation

1
FP32 → FP16 → INT8 → INT4
BenefitTradeoff
Lower VRAMPotential quality loss
Lower memory bandwidthTask-dependent accuracy drop
Faster inference (often)Calibration complexity
Lower infra costNot all ops accelerate equally

Use quantization when:

  • Model does not fit on available GPU(s)
  • Serving high QPS at fixed hardware budget
  • Running large models on consumer GPUs locally
  • Batch inference where latency tolerance is higher

Always benchmark:

1
2
3
FP16: quality X, latency Y, VRAM Z
INT8: quality X', latency Y', VRAM Z'
INT4: quality X'', latency Y'', VRAM Z''

For the math, formats, and Python snippets, see LLM Quantization Explained.

FDE / production angle

FDE interviews want a decision framework, not “always use 4-bit.” Mention GPTQ vs AWQ vs GGUF trade-offs, eval on domain tasks, and that quantization affects fine-tuned LoRA compatibility.

Follow-up questions to expect

  • Can you fine-tune a quantized model?
  • When does INT4 quality become unacceptable?
  • Quantization for embedding models vs generation models?

Interview takeaway

“Quantize when memory or cost demands it — after benchmarking quality and latency on real workloads, not leaderboard scores.”


Q3: What is your batching and caching strategy?

Short answer (30 seconds)

I use dynamic or continuous batching for throughput — vLLM-style — with a short request queue to bound latency. I cache embeddings, KV cache during generation, and deterministic responses where safe. Excessive batching increases tail latency, so I tune batch size against SLA.

Deep explanation

Static batching:

1
[Req1, Req2, Req3] → GPU   (efficient, but Req3 waits for Req1+2)

Continuous batching (vLLM, TGI):

Requests join and leave the batch dynamically — better GPU utilization at scale. See vLLM + Ray serving guide.

Caching layers:

CacheWhat it storesWhen it helps
Embedding cacheVector for input textRepeated or overlapping docs
KV cacheKey/value attention statesAutoregressive generation — avoids recomputing prior tokens
Response cacheFull LLM outputIdentical deterministic queries

Production stack:

1
Short Request Queue → Dynamic Batching → Inference Engine → KV Cache

FDE / production angle

Quote p99 latency targets when discussing batching. Mention prefix caching (OpenAI, some open engines) for RAG where system prompt + retrieved docs repeat across users.

Follow-up questions to expect

  • How do you handle latency SLA vs throughput?
  • When does response caching produce stale or wrong answers?
  • Multi-GPU batching vs tensor parallelism?

Interview takeaway

“Continuous batching for throughput, KV cache for generation speed, embedding and response caches where inputs repeat — always bounded by latency SLA.”


Q4: When should you use hosted APIs vs open-source models?

Short answer (30 seconds)

Hosted APIs for fast time-to-market, minimal infra, and best frontier capability without GPU ops. Self-hosted open-source when you need data sovereignty, predictable high-volume economics, customization, fine-tuning control, or offline/air-gapped deployment.

Deep explanation

Decision framework:

1
2
3
4
5
6
Fastest prototype?              → Hosted API
Sensitive data / sovereignty? → Self-hosted
Very high sustained volume?     → Analyze unit economics
Need custom weights / LoRA?     → Open-source
Best capability, no ML ops?     → Hosted API
Regulatory air-gap?             → Self-hosted
FactorHosted APISelf-hosted
Time to productionDaysWeeks–months
Data controlProvider policyFull
Unit cost at low volumeOften cheaperHigh fixed cost
Unit cost at high volumeCan exceed self-hostCan win
Model updatesAutomaticManual
CustomizationLimitedFull (fine-tune, quantize)

FDE / production angle

FDE roles often deploy both — frontier model via API for complex tasks, smaller self-hosted model for high-volume classification or routing. Mention hybrid routing by query complexity.

Follow-up questions to expect

  • How do you calculate the break-even point for self-hosting?
  • How do you handle model deprecation on hosted APIs?
  • Multi-model routing architecture?

Interview takeaway

“Hosted for speed and capability; self-hosted for control, customization, and economics at scale — often both in a routed architecture.”


Q5: How do you make an AI system more deterministic and less brittle?

Short answer (30 seconds)

Wrap the LLM in deterministic constraints: input validation, intent classification, structured output schemas, post-generation validation, business rules, state machines, and retries. Do not ask the LLM to do what regular code does better.

Deep explanation

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
User Input
    ↓
Validation
    ↓
Intent Classification
    ↓
LLM / Workflow
    ↓
Structured Output (JSON schema)
    ↓
Validation
    ↓
Business Rules
    ↓
Final Response

Techniques:

  • Temperature near zero + structured output
  • Function calling with strict schemas
  • Constrained decoding where available
  • Schema validation (Pydantic, JSON Schema) on every response
  • Retries with backoff on validation failure
  • State machines for multi-step workflows
  • Deterministic business logic outside the LLM (pricing, eligibility, dates)

Key principle:

Don’t ask an LLM to perform deterministic computation if normal code can do it better.

FDE / production angle

Interviewers love hearing you separate probabilistic (LLM summarization) from deterministic (SQL lookup, rule engine) layers. Draw the validation sandwich diagram.

Follow-up questions to expect

  • What happens when JSON schema validation fails repeatedly?
  • How do you test brittle edge cases?
  • Agent vs workflow — when is each appropriate?

Interview takeaway

“The LLM is one component inside a controlled system — validation, schemas, state machines, and deterministic code handle everything that must not vary.”


Q6: What fallback do you use if the LLM fails mid-task?

Short answer (30 seconds)

Graceful degradation: retry with exponential backoff, fall back to a smaller or more reliable model, simplify the prompt, switch to a deterministic workflow, then escalate to human support. For multi-step agents, checkpoint state so you resume from the failed step — not from scratch.

Deep explanation

>flowchart TB Primary[Primary Model] -->|Success| Continue[Continue Workflow] Primary -->|Failure| Retry[Retry with Backoff] Retry -->|Still failing| Fallback[Fallback Model] Fallback -->|Still failing| Simple[Simplified Prompt] Simple -->|Still failing| Deterministic[Deterministic Workflow] Deterministic -->|Still failing| Human[Human Escalation]

Agent checkpoint example:

1
2
3
4
5
6
{
  "workflow_id": "123",
  "completed_steps": [1, 2, 3],
  "current_step": 4,
  "state": {"order_id": "ORD-456", "user_verified": true}
}

If step 4 fails, resume from step 4 with persisted state — do not re-run steps 1–3.

Failure modes to handle:

  • Timeout / rate limit → retry + backoff
  • Invalid JSON output → retry with stricter schema instruction
  • Model unavailable → fallback model or cached response
  • Tool call failure → deterministic fallback or human handoff

FDE / production angle

Define SLAs per degradation tier. Log which fallback tier triggered for every request — patterns indicate primary model or prompt issues.

Follow-up questions to expect

  • How many retries before giving up?
  • How do you avoid infinite retry loops?
  • Circuit breaker pattern for model endpoints?

Interview takeaway

“Fallback is a tiered strategy — retry, alternate model, simplify, deterministic path, human — with checkpointed state for long workflows.”


Q7: Can you solve this without an LLM or vector DB?

Short answer (30 seconds)

My first question is whether an LLM is necessary at all. If the input is structured, the knowledge base is small, or the output is deterministic — SQL, rules engines, keyword search, or templates may be simpler, cheaper, and more reliable.

Deep explanation

Decision checklist:

QuestionIf yes → consider
Is input structured?SQL, API lookup, rules engine
Is knowledge base small?Full-text search, BM25
Is output deterministic?Templates, business rules, code
Is semantic understanding required?Embeddings, RAG, LLM

Examples where LLM is overkill:

  • “What is order #12345 status?” → SQL
  • “Reset password” with fixed steps → workflow + templates
  • FAQ with 50 known questions → keyword match + canned answers
  • Policy lookup by document ID → direct fetch, no retrieval

When LLM + RAG is justified:

  • Unstructured natural language queries
  • Large, evolving document corpus
  • Answers require synthesis across multiple sources
  • Semantic similarity matters (“similar issues” not exact keyword match)

FDE / production angle

This is a senior/system design signal. Interviewers test whether you reach for LLM by default. Answer with the simplest architecture that meets requirements, then justify escalation.

Follow-up questions to expect

  • When would you add BM25 before vector search?
  • Hybrid search — why not keyword alone?
  • Cost comparison: rules engine vs RAG for 1000 FAQs?

Interview takeaway

“Start with the simplest architecture — SQL, search, rules — and add LLM/RAG only when semantic understanding or open-ended synthesis is genuinely required.”


Q8: What’s the right database — SQL, NoSQL, or Vector?

Short answer (30 seconds)

Usually the answer is all three in a polyglot setup. SQL for transactions and structured metadata, NoSQL for flexible documents and high-volume events, vector store for semantic retrieval — plus object storage for raw files and Redis for cache/sessions.

Deep explanation

StoreUse forExample
SQL (Postgres)Transactions, relationships, metadataUsers, orders, chunk metadata
NoSQLFlexible schema, scale, eventsSession state, app config, logs
Vector DBSemantic similarity, RAGDocument embeddings
Object storageRaw files, model artifactsPDFs, checkpoints
RedisCache, sessions, rate limitsEmbedding cache, chat state

Typical AI application stack:

1
2
3
4
PostgreSQL     → Users, transactions, document metadata, pgvector (optional)
Redis          → Session, cache, rate limiting
Vector Store   → Semantic retrieval (Pinecone, OpenSearch, etc.)
Object Storage → S3/GCS for raw documents and model weights

FDE / production angle

Mention pgvector when the team already runs Postgres — strong consistency between metadata and vectors, simpler ops. Dedicated vector DB when ANN scale or hybrid search features exceed pgvector comfort zone.

Follow-up questions to expect

  • When is pgvector enough vs dedicated vector DB?
  • How do you keep SQL metadata in sync with vector index?
  • Eventual consistency between stores — how do you handle it?

Interview takeaway

“Polyglot persistence — SQL for structure, vector for semantics, object storage for files, Redis for speed — not one database for everything.”


Scenario: Embedding model migration — how do you migrate safely?

This scenario appears frequently in FDE interviews. Part 2 introduced the dual-index pattern; here is the full playbook.

Problem statement

Interviewer: “Your production RAG uses embedding model A (1024-dim). You want to upgrade to model B (1536-dim). Walk me through migration with zero downtime.”

Short answer (2 minutes)

I build index v2 in parallel, backfill all historical documents with model B, dual-write new documents to both indexes, evaluate retrieval on a golden set, run shadow traffic, canary to production traffic, switch the read alias, and keep v1 for rollback until v2 is stable.

Migration architecture

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
                  Documents
                      │
                      ▼
                Processing
                  │       │
                  ▼       ▼
          Embedding A   Embedding B
                  │       │
                  ▼       ▼
               Index v1  Index v2
                  │       │
                  └───┬───┘
                      ▼
               Evaluation Layer
                      │
                      ▼
                Production Alias

Phase-by-phase playbook

PhaseActionGate to proceed
1. Create v2Stand up new index with model B schemaIndex healthy, empty
2. BackfillRe-embed all historical documents100% document coverage
3. Dual-writeNew docs → both v1 and v2No write failures >0.1%
4. EvaluateGolden set: Precision@K, Recall@K, NDCG, MRR, answer correctness, citation accuracyv2 ≥ v1 on all critical metrics
5. ShadowRoute duplicate queries to v2, compare metrics onlyNo regression over 1–2 weeks
6. Canary5% → 25% → 50% → 100% production trafficError rate and latency within SLA
7. Switch aliasretrieval_current → index_v2Canary stable
8. Rollback bufferKeep v1 for 2–4 weeksSoak period clean

Evaluation metrics table

Metricv1 baselinev2 targetRollback trigger
Precision@50.72≥ 0.72Drop > 5%
Recall@100.85≥ 0.85Drop > 5%
MRR0.68≥ 0.68Drop > 5%
Answer correctness (golden set)0.81≥ 0.81Drop > 3%
Citation accuracy0.76≥ 0.76Drop > 3%
p99 retrieval latency120ms≤ 150msExceed 200ms

Rollback plan

1
2
3
4
IF canary error_rate > 2× baseline OR Precision@5 drops > 5%:
  → Flip alias back to index_v1
  → Stop dual-write to v2
  → Root-cause before retry

Interview takeaway

“Embedding migration is a data migration problem — parallel indexes, dual-write, golden-set evaluation, shadow, canary, alias switch, and rollback buffer. Never mix incompatible vectors in one index.”


Cheat sheet — three rapid-fire model answers

“How do you design a production RAG system?”

“I separate ingestion, chunking, embedding, retrieval, reranking, generation, and evaluation into independently versioned components. Structure-aware chunking, metadata filtering, hybrid retrieval where appropriate, rerank before generation. Every request is traced with model, prompt, embedding, retrieval, and document versions. For embedding migrations I use parallel indexes, backfill, shadow evaluation, canary rollout, and alias switching.”

“How do you reduce AI cost?”

“Measure token usage and latency by component. Reduce unnecessary context — summarize history, retrieve 30 and rerank to 5, cap output tokens, cache embeddings and repeated responses. Route simple queries to smaller models. Batch inference where latency allows. Quantize self-hosted models after benchmarking quality.”

“How do you make an LLM application reliable?”

“Don’t rely on the LLM alone. Deterministic code for deterministic tasks, structured outputs, validation layers, retries, tiered fallbacks, versioned prompts and context, full observability, golden datasets in CI, and graceful degradation. The LLM is one component inside a controlled system.”


Study plan — how to rehearse

FDE interviews reward scenario-based answers, not flashcard definitions.

Week 1 — Foundations (Part 1):

  • Speak each short answer out loud in under 30 seconds
  • Draw the context window budget diagram from memory
  • Practice the LoRA vs RAG decision tree with a real customer example

Week 2 — RAG & Ops (Part 2):

  • Whiteboard the full RAG pipeline and trace schema
  • Walk through zero-downtime embedding migration end to end
  • Prepare one story from your experience about a retrieval quality bug

Week 3 — System design (Part 3):

  • Practice “solve without LLM” for 3 different prompts
  • Run through the embedding migration scenario with metrics and rollback
  • Rehearse the three cheat sheet answers until conversational

Day before interview:

  • Re-read only the Interview takeaway lines from all three parts
  • Prepare 2 production war stories: one RAG quality issue, one cost/latency win

Series complete

Full series:

  1. Part 1: LLM Fundamentals & Context Engineering
  2. Part 2: RAG Systems & LLMOps
  3. Part 3: Production Trade-offs & System Design (you are here)

Related deep dives on this blog:

Source: Adapted and expanded from interview prep notes, reframed for Forward Deployed / Applied AI Engineer interviews.

This post is licensed under CC BY 4.0 by the author.