Posts AI Engineering Interview Prep — Part 2: RAG Systems & LLMOps
Post
Cancel

AI Engineering Interview Prep — Part 2: RAG Systems & LLMOps

Data pipeline abstract — retrieval and operations at scale.

Series: ← Part 1: Foundations · Part 2 of 3 · Part 3: Production & System Design →

Who this is for: Forward Deployed AI Engineers building production RAG systems — where chunking, retrieval quality, zero-downtime migrations, and observability matter as much as model choice.

Part 2 covers RAG engineering and LLMOps. For foundational RAG concepts, see the RAG Comprehensive Guide. For safety and guardrails, see Guardrail LLM.


Production RAG pipeline overview

>flowchart TB subgraph ingestion [Ingestion] Docs[Documents] Chunk[Chunking] Embed[Embedding] end subgraph retrieval [Retrieval] VecSearch[Vector Search] Hybrid[Hybrid BM25] Rerank[Reranker] end subgraph generation [Generation] LLM[LLM] Validate[Output Validation] end subgraph ops [LLMOps] Trace[Tracing] Eval[Offline Eval] Monitor[Drift Monitor] end Docs --> Chunk --> Embed --> VecSearch Query[User Query] --> VecSearch VecSearch --> Hybrid --> Rerank --> LLM --> Validate LLM --> Trace Trace --> Eval Trace --> Monitor

Q1: What is your chunking strategy?

Short answer (30 seconds)

There is no universal chunk size. I use structure-aware chunking first — respect headings, sections, and document hierarchy — then semantic boundaries, with token limits as a hard constraint. For enterprise docs I target 300–800 tokens per chunk with 10–20% overlap where needed, plus rich metadata for filtering.

Deep explanation

Option 1: Fixed-size chunking

1
512 tokens, 50-token overlap
  • Simple and predictable
  • May split mid-sentence or mid-concept

Option 2: Semantic chunking

Split when topic meaning shifts. Better retrieval quality, more complex to implement and debug.

Option 3: Structure-aware chunking (enterprise default)

1
2
3
4
5
Document
├── Chapter
│   ├── Section
│   │   ├── Heading
│   │   └── Content

Example metadata:

1
2
3
4
5
6
{
  "document": "installation_guide",
  "chapter": "Security",
  "section": "Kerberos Authentication",
  "chunk_id": 47
}

Recommended enterprise recipe:

1
Heading + Parent Context + 300–800 token chunk + 10–20% overlap

FDE / production angle

Chunking version is a first-class artifact — changing it requires re-indexing. Pair chunks with metadata filters (product, version, region) so retrieval narrows before vector search. Interviewers often follow up: “How do you handle PDFs with tables?” — answer with layout-aware parsers (Unstructured, Docling) and table-as-single-chunk strategies.

Follow-up questions to expect

  • How do you chunk code vs prose vs tables?
  • What chunk size do you pick for legal or policy documents?
  • How do you evaluate whether chunking changed retrieval quality?

Interview takeaway

“Structure-aware chunking first, semantic boundaries second, token limits as hard constraint. Metadata makes chunks filterable before similarity search runs.”


Q2: How do you choose a vector database?

Short answer (30 seconds)

I do not pick a vector DB on ANN benchmark scores alone. I evaluate scale, metadata filtering, hybrid search support, operational model, multi-tenancy, cost, and fit with the existing platform — then prototype with real queries and documents.

Deep explanation

RequirementConsideration
ScaleNumber of vectors, growth rate
Latencyp50/p99 query SLA
FilteringMetadata pre-filter vs post-filter
Hybrid searchKeyword + vector fusion
OperationsManaged vs self-hosted
CostStorage + query volume at peak
EcosystemExisting OpenSearch, Postgres, cloud stack

Quick reference (not gospel):

ToolGood for
ChromaPrototypes, local dev, small deployments
PineconeManaged, fast time-to-production
OpenSearch / ElasticsearchHybrid search, existing search ops team
pgvectorPostgres-native, moderate scale, strong consistency needs

FDE / production angle

FDE interviews test requirements mapping, not fanboy picks. Mention multi-tenancy isolation, backup/restore, and whether you need strong consistency with transactional metadata in the same store (Postgres + pgvector) vs dedicated ANN engine.

Follow-up questions to expect

  • When would you use pgvector instead of a dedicated vector DB?
  • How do you implement hybrid search?
  • How do you handle multi-tenant isolation?

Interview takeaway

“Vector DB selection is a requirements exercise — scale, filtering, hybrid search, ops model, and platform fit — validated with real data, not leaderboard scores.”


Q3: Can you update or backfill embeddings with zero downtime?

Short answer (30 seconds)

Yes — use a versioned dual-index approach. Build index v2 in parallel, backfill historical documents, dual-write new docs, evaluate retrieval quality offline, run shadow traffic, canary rollout, then switch a read alias. Keep v1 for rollback.

Deep explanation

Problem: embeddings from model A are not compatible with model B. Never mix them in one index.

1
2
Documents → Embedding v1 → Index v1 (production)
Documents → Embedding v2 → Index v2 (building)

Zero-downtime migration flow:

>flowchart TB Docs[Documents] --> DualWrite[Dual Write] DualWrite --> IndexV1[Index v1 Production] DualWrite --> IndexV2[Index v2 Backfill] IndexV1 --> ProdRead[Production Reads] IndexV2 --> ShadowEval[Shadow Evaluation] ShadowEval --> Canary[Canary Rollout] Canary --> AliasSwitch[Switch Alias] AliasSwitch --> IndexV2Prod[Index v2 Production] IndexV1 --> Rollback[Rollback Buffer]

Steps:

  1. Create index v2
  2. Backfill all historical documents
  3. Dual-write new documents to both indexes
  4. Validate coverage and retrieval metrics
  5. Shadow traffic (v2 serves metrics only, not users)
  6. Canary (5% → 25% → 50% → 100%)
  7. Switch alias: retrieval_current → index_v2
  8. Keep v1 for rollback; retire later

Part 3 expands this into a full scenario with evaluation metrics and rollback criteria.

FDE / production angle

This is one of the highest-value FDE answers. Emphasize alias-based switching, dual-write during transition, and never deleting v1 until v2 is proven stable for a defined soak period.

Follow-up questions to expect

  • How do you dual-write without doubling ingestion latency?
  • What metrics gate the alias switch?
  • How long do you keep v1 for rollback?

Interview takeaway

“Embedding migrations use parallel indexes, dual-write, shadow evaluation, canary rollout, and alias switching — never in-place overwrites of incompatible vectors.”


Q4: How do you evaluate retrieval quality?

Short answer (30 seconds)

I separate retrieval evaluation from generation evaluation. For retrieval I measure Precision@K, Recall@K, MRR, and NDCG on a golden query set — before and after reranking. For generation I add citation accuracy and answer correctness against retrieved evidence.

Deep explanation

Retrieval metrics:

MetricQuestion it answers
Precision@KOf top-K retrieved, how many are relevant?
Recall@KOf all relevant docs, how many appear in top-K?
MRRHow high is the first relevant result ranked?
NDCGDoes ranking order respect relevance grades?

Example:

1
2
3
4
5
6
Relevant docs: A, B, C
Retrieved top 5: A, D, B, E, F

Precision@5 = 3/5 = 0.60
Recall@5    = 3/3 = 1.00  (all relevant found)
MRR         = 1/1 = 1.00  (A is rank 1)

Full RAG pipeline evaluation:

1
2
Query → Vector Top 50 → Keyword Top 50 → Fusion → Reranker → Top 5 → LLM
         ↑ evaluate here                    ↑ and here

Citation evaluation checklist:

  • Is a citation present?
  • Does the cited passage support the claim?
  • Is it the correct document?
  • Is the cited span actually relevant?

FDE / production angle

Build a golden dataset of 50–200 real user queries with human-labeled relevant documents. Run it in CI on every prompt, embedding, or chunking change. Track regression, not just absolute scores.

Follow-up questions to expect

  • How do you build a golden dataset without manual labeling for everything?
  • LLM-as-judge for retrieval — when is it valid?
  • How do you evaluate multi-hop or conversational retrieval?

Interview takeaway

“Retrieval and generation are evaluated separately. Golden datasets, pre/post rerank metrics, and citation checks run in CI — not just ad-hoc eyeballing.”


Q5: Sketch the pipeline — raw data to model to serving to feedback

Short answer (30 seconds)

Production LLM systems split into ingestion, training/RAG indexing, serving, monitoring, and feedback loops — each independently versioned. Data flows from raw sources through validation, into either a training pipeline or RAG document pipeline, then to model registry and vector index, then to serving with full tracing and evaluation feedback.

Deep explanation

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
                 DATA SOURCES
                      │
                      ▼
              Ingestion / ETL
                      │
                      ▼
                Raw Storage
                      │
                      ▼
              Validation / Quality
                      │
              ┌───────┴────────┐
              ▼                ▼
        Training Data      RAG Documents
              │                │
              ▼                ▼
        Training Pipeline   Chunking
              │                │
              ▼                ▼
           Evaluation       Embeddings
              │                │
              ▼                ▼
         Model Registry     Vector Index
              │                │
              └───────┬────────┘
                      ▼
                Model Serving
                      │
                      ▼
              Monitoring / Tracing
                      │
                      ▼
                  Feedback
                      │
                      ▼
              Evaluation Loop

Key FDE principle: each box is independently deployable and versioned. Changing chunking should not require redeploying the LLM. Changing the LLM should not require re-chunking unless eval shows retrieval regression.

FDE / production angle

Draw this diagram in interviews. Label version pins at each stage. Mention that FDE work often focuses on the RAG branch + serving + monitoring, not training from scratch.

Follow-up questions to expect

  • Where does human feedback enter the loop?
  • How do you handle document updates vs full re-index?
  • What triggers a retraining vs RAG-only update?

Interview takeaway

“I decompose the pipeline into versioned stages — ingestion, indexing, serving, monitoring, feedback — so changes are isolated and evaluable.”


Q6: How do you monitor performance drift or hallucinations?

Short answer (30 seconds)

LLM drift is multi-layered. I monitor input drift (query topics shifting), retrieval drift (similarity scores, no-result rate, citation coverage), and output quality (human feedback, automated evaluators, golden datasets). Hallucination signals include unsupported claims, citation mismatch, and low retrieval relevance.

Deep explanation

Traditional ML drift: feature distribution shift, prediction decay.

LLM-specific layers:

LayerWhat to monitor
Input driftEmbedding distribution of queries, topic/language shifts
Retrieval driftSimilarity scores, no-result rate, Precision@K, citation coverage
Output qualityHuman thumbs, LLM-as-judge, golden set regression
Hallucination signalsUnsupported claims, citation mismatch, contradiction with context

Production pattern:

1
Online Monitoring + Offline Evaluation + Human Feedback

Set alerts on: sudden drop in retrieval similarity, spike in “I don’t know” responses, increase in citation-free answers, user negative feedback rate.

FDE / production angle

Pair online dashboards with weekly offline golden-set runs. For regulated industries, log hallucination flags for audit. See Guardrail LLM for safety-specific patterns.

Follow-up questions to expect

  • How do you detect hallucinations without human review on every request?
  • What is your threshold for triggering a rollback?
  • LLM-as-judge — how do you calibrate it against human labels?

Interview takeaway

“Drift monitoring spans input, retrieval, and output layers. Hallucination detection combines retrieval relevance checks, citation validation, and automated evaluators — not hope.”


Q7: How do you log prompts and outputs for debugging and auditing?

Short answer (30 seconds)

Every request gets a trace with model, prompt, retrieval, and embedding versions, token counts, latency, retrieved documents, tool calls, and output. I redact PII before storage, enforce RBAC, and apply retention policies — never log raw passwords, API keys, or unredacted sensitive documents.

Deep explanation

Trace schema:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
{
  "trace_id": "abc123",
  "timestamp": "2026-08-24T10:15:00Z",
  "user_id_hash": "sha256:...",
  "model": "gpt-4o-2026-08",
  "prompt_version": "v12",
  "retrieval_version": "v5",
  "embedding_version": "bge-large-v2",
  "input_tokens": 1200,
  "output_tokens": 450,
  "latency_ms": 2100,
  "retrieved_documents": [
    {"doc_id": "policy_123", "chunk_id": 47, "score": 0.89}
  ],
  "tool_calls": [],
  "output": "...",
  "feedback": null
}

Security pipeline:

1
PII Detection → Redaction → Encrypted Storage → RBAC → Retention Policy

Never log: passwords, API keys, full credit card numbers, unredacted health records.

FDE / production angle

Traces are your debugging lifeline when a customer reports a wrong answer three days later. Hash user IDs, store document IDs not full document text when possible, and make traces searchable by trace_id, prompt_version, and time range.

Follow-up questions to expect

  • How do you balance audit requirements with privacy (GDPR)?
  • What do you log for tool-calling agents vs simple RAG?
  • How long do you retain traces?

Interview takeaway

“Structured traces with version pins enable debugging and audit. PII redaction and RBAC are non-negotiable — logging everything raw is a liability.”


Q8: CI/CD for LLM workflows — what’s different from traditional ML?

Short answer (30 seconds)

LLM CI/CD adds prompt regression tests, golden dataset evaluation, safety tests, retrieval metrics, cost and latency regression checks, and canary deployment — on top of standard unit tests. Version prompts, models, embeddings, indexes, chunking, eval datasets, and tools as first-class artifacts.

Deep explanation

Traditional CI/CD:

1
Code → Unit Test → Build → Deploy

LLM CI/CD:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
Prompt / Index / Model Change
      ↓
Unit Tests (parsers, validators, API contracts)
      ↓
Golden Dataset — retrieval + generation quality
      ↓
Safety Tests (injection, PII leakage, refusal behavior)
      ↓
Cost Regression (token count vs baseline)
      ↓
Latency Regression (p99 vs baseline)
      ↓
Canary Deployment
      ↓
Production Monitoring

Version as artifacts:

1
Prompts | Models | Embeddings | RAG Index | Chunking | Eval Datasets | Tools

Block deploy if: golden set Precision@5 drops >5%, average token cost rises >20%, p99 latency exceeds SLA, or safety test failures.

FDE / production angle

This is how you answer “how do you ship prompt changes safely?” — not “we test manually.” Mention feature flags for prompt versions and shadow mode for retrieval changes before full rollout.

Follow-up questions to expect

  • How do you prevent eval dataset contamination?
  • What runs in CI vs nightly vs pre-release?
  • How do you test non-deterministic outputs in CI?

Interview takeaway

“LLM CI/CD versions every behavior-affecting artifact and gates deploys on golden-set quality, safety, cost, and latency — not just green unit tests.”


What’s next

Part 3 covers production trade-offs and system design: token cost reduction, quantization, batching, hosted vs self-hosted, determinism, fallbacks, when to skip LLM entirely, database selection, a full embedding migration scenario, and a rapid-fire cheat sheet.

Series navigation: ← Part 1: Foundations · Part 2 of 3 · Part 3: Production & System Design →

This post is licensed under CC BY 4.0 by the author.