Practise testing, ingestion, and enterprise workflows. Prioritize the first four topics; spend about three minutes on each.
1. Testing AI applications
Question: “How do you test an application when LLM outputs are nondeterministic?”
“I test deterministic components with unit and integration tests, then evaluate model behavior against a versioned dataset with measurable acceptance criteria. I check correctness, grounding, security, latency, and cost. I avoid requiring exact wording unless the output contract demands it.”
Know the layers:
| Layer | Example |
|---|---|
| Unit tests | Parsing, metadata filters, tool argument validation |
| Integration tests | Retrieval against a test index; API and database interactions |
| Contract tests | Required fields, valid types, allowed tool names |
| AI evaluations | Correct answers, appropriate abstention, unsupported claims |
| Adversarial tests | Malicious documents, unauthorized requests, misleading inputs |
| Regression tests | Whether a prompt/model change harms previously successful cases |
Follow-up: “Would you use an LLM as a judge?”
Yes, as one evaluation signal. Calibrate it against human judgments, use explicit rubrics, and check for bias and inconsistent scoring. Judge scores are not ground truth.
2. Structured outputs and validation
Question: “How do you reliably extract information from documents into JSON?”
“I define a schema, use constrained structured output where supported, validate the result, and apply business rules separately. I preserve source evidence for extracted values and represent missing information explicitly.”
Three levels of correctness:
- Syntax: Is it valid JSON?
- Schema: Are required fields and types correct?
- Meaning: Are the values supported by the document and valid for the business?
For example, a valid numeric invoice total can still be incorrectly extracted. Check currency, line items, and relevant arithmetic through deterministic code.
Remember: Schema-valid output does not guarantee factual accuracy.
3. Document ingestion, OCR, and tables
Question: “Why does a RAG system struggle with PDFs?”
PDFs often encode visual layout rather than clean reading order. Scanned pages require OCR; multi-column layouts, tables, headers, and footnotes can break extraction.
“I classify document types and choose suitable parsing methods. I preserve headings, page references, and table structure, then evaluate extraction quality before tuning retrieval.”
For tables:
- Retain column headers, units, dates, and row labels.
- Avoid chunks that separate values from their meaning.
- Keep enough provenance to inspect the original page.
- Use structured queries or code when exact aggregation is needed.
Follow-up: “Should you embed every document as plain text?”
No universal rule. Text, tables, and images may need different representations and retrieval paths.
4. Versioning and incremental indexing
Question: “What happens when a source document changes or is deleted?”
“I maintain stable document identifiers, content versions, and chunk-to-source mappings. Changes trigger reprocessing and replacement of affected chunks. Deletions remove retrievable content, and related caches are invalidated.”
Important decisions:
| Concern | What to explain |
|---|---|
| Duplicate ingestion | Detect repeated content or source versions |
| Failed processing | Retry safely and track processing status |
| Partial updates | Avoid exposing an inconsistent mixture of versions |
| Model changes | Track which embedding model generated each index |
| Rollback | Retain a usable previous version where requirements allow |
Embedding migration trap: Even two models with the same embedding dimension can produce incompatible vector spaces. Use compatible query/document embeddings; typically build a separate index and re-embed documents before switching.
5. Vector search internals
Question: “Why use approximate nearest-neighbor search?”
Exact search compares a query against all indexed vectors. Approximate search trades some retrieval accuracy for lower latency and improved scalability.
Know these terms:
- Cosine similarity: Compares vector direction.
- Dot product: Depends on direction and magnitude.
- Euclidean distance: Measures straight-line distance.
- HNSW: A graph-based approximate search method.
- ANN recall: How well approximate search reproduces exact nearest neighbors.
For unit-normalized vectors, dot product equals cosine similarity, and Euclidean distance gives the same ranking.
Follow-up: “Does higher embedding dimension mean better retrieval?”
No. Task fit and measured retrieval quality matter. More dimensions also increase storage and computation.
Distinguish ANN recall from relevance recall: reproducing exact vector neighbors does not prove those neighbors answer the user’s question.
6. Stateful agents and conversation memory
Question: “How do you manage state across a long-running AI workflow?”
“I keep durable workflow state outside the model, persist checkpoints, and define explicit transitions. After interruption, the workflow resumes from recorded state, with safeguards against repeating completed actions.”
Separate:
- Conversation history: Previous messages.
- Workflow state: Current step, intermediate results, completed actions.
- Long-term memory: Selected information retained across sessions.
Do not store everything indefinitely. Consider relevance, user isolation, consent requirements, retention, and deletion.
Subtle point: Summarizing conversation history saves tokens but may lose important details. Preserve critical facts and action records explicitly.
7. Text-to-SQL
Question: “How would you build an assistant that answers questions from a database?”
“I provide relevant schema context, generate a proposed query, validate it, and execute through a restricted database interface. I apply authorization and query limits, then ground the response in returned results.”
Cover:
- Read-only credentials for analytical use.
- Authorized tables, rows, and columns.
- Parameterized values where applicable.
- Query timeouts and result-size limits.
- Ambiguous definitions such as “active customer.”
- Joins, aggregation grain, and duplicate counting.
Common trap: A query can execute successfully and still answer the wrong business question. Evaluate semantic correctness, not just execution success.
8. Data pipelines and delivery guarantees
Question: “How would you process a large volume of documents reliably?”
“I decouple ingestion from processing with a queue, use workers with bounded concurrency, and track document status. Processing is idempotent, retries are limited, and persistent failures go to a dead-letter queue for investigation.”
Be ready to explain:
- At-least-once delivery: Messages may be delivered more than once.
- Idempotent processing: Repeated processing does not duplicate effects.
- Backpressure: Limit intake or concurrency when downstream systems cannot keep up.
- Dead-letter queue: Holds messages that repeatedly fail.
Avoid casually promising “exactly once.” Explain the boundaries and how duplicate effects are prevented.
9. Business value and deciding whether AI is appropriate
Question: “How do you decide whether a use case needs GenAI?”
“I first define the task and compare a simple baseline against an AI approach. I consider input variability, error tolerance, operational cost, and whether generated output provides measurable value.”
Examples:
| Task | Useful baseline |
|---|---|
| Fixed calculations | Deterministic code |
| Known structured rules | Rules engine |
| Predicting a category from labeled examples | Conventional ML |
| Searching documents | Search without generation |
| Interpreting varied language and synthesizing evidence | Consider an LLM |
Measure business outcomes such as handling time, correction rate, completion rate, and cost per successful task. A technically impressive model may still fail to improve the workflow.
10. Five difficult follow-ups to practise now
Answer these aloud before reading the suggested direction:
-
“Your evaluation score improved, but users complain more. Why?”
The dataset may be unrepresentative, metrics may miss important errors, or latency and workflow usability may have worsened. Investigate actual failure cases. -
“Would you use multiple agents?”
Only when decomposition provides measured value. Consider coordination overhead, compounded errors, latency, and harder debugging. -
“What if the model says it is 95% confident?”
Self-reported confidence is not automatically calibrated. Validate uncertainty signals empirically and define escalation criteria. -
“How do you reproduce an incorrect answer?”
Capture relevant versions, authorized context identifiers, retrieval results, generation settings, and tool traces—subject to data handling constraints. Exact replay may still vary. -
“What do you do when two authoritative documents disagree?”
Use documented precedence and effective-date rules. If the conflict remains unresolved, surface it with sources and escalate.
Finish with one sentence you can defend:
“I choose the approach from the requirements, verify it against a baseline, and make its failures observable.”