Skip to content
StudyHubA place to keep learning

Explore

  • Browse Topics
  • Study Packs
  • Library Topics
    • AI engineering interviews

My Workspace

  • My Notes
StudyHub · Testing and enterprise workflows
Browse topics
AI engineering interviews

Testing and enterprise workflows

StudyHub7 min readUpdated Oct 3, 2026

Practise testing, ingestion, and enterprise workflows. Prioritize the first four topics; spend about three minutes on each.

1. Testing AI applications

Question: “How do you test an application when LLM outputs are nondeterministic?”

“I test deterministic components with unit and integration tests, then evaluate model behavior against a versioned dataset with measurable acceptance criteria. I check correctness, grounding, security, latency, and cost. I avoid requiring exact wording unless the output contract demands it.”

Know the layers:

Layer Example
Unit tests Parsing, metadata filters, tool argument validation
Integration tests Retrieval against a test index; API and database interactions
Contract tests Required fields, valid types, allowed tool names
AI evaluations Correct answers, appropriate abstention, unsupported claims
Adversarial tests Malicious documents, unauthorized requests, misleading inputs
Regression tests Whether a prompt/model change harms previously successful cases
Scroll across to read all columns.

Follow-up: “Would you use an LLM as a judge?”

Yes, as one evaluation signal. Calibrate it against human judgments, use explicit rubrics, and check for bias and inconsistent scoring. Judge scores are not ground truth.

2. Structured outputs and validation

Question: “How do you reliably extract information from documents into JSON?”

“I define a schema, use constrained structured output where supported, validate the result, and apply business rules separately. I preserve source evidence for extracted values and represent missing information explicitly.”

Three levels of correctness:

  • Syntax: Is it valid JSON?
  • Schema: Are required fields and types correct?
  • Meaning: Are the values supported by the document and valid for the business?

For example, a valid numeric invoice total can still be incorrectly extracted. Check currency, line items, and relevant arithmetic through deterministic code.

Remember: Schema-valid output does not guarantee factual accuracy.

3. Document ingestion, OCR, and tables

Question: “Why does a RAG system struggle with PDFs?”

PDFs often encode visual layout rather than clean reading order. Scanned pages require OCR; multi-column layouts, tables, headers, and footnotes can break extraction.

“I classify document types and choose suitable parsing methods. I preserve headings, page references, and table structure, then evaluate extraction quality before tuning retrieval.”

For tables:

  • Retain column headers, units, dates, and row labels.
  • Avoid chunks that separate values from their meaning.
  • Keep enough provenance to inspect the original page.
  • Use structured queries or code when exact aggregation is needed.

Follow-up: “Should you embed every document as plain text?”

No universal rule. Text, tables, and images may need different representations and retrieval paths.

4. Versioning and incremental indexing

Question: “What happens when a source document changes or is deleted?”

“I maintain stable document identifiers, content versions, and chunk-to-source mappings. Changes trigger reprocessing and replacement of affected chunks. Deletions remove retrievable content, and related caches are invalidated.”

Important decisions:

Concern What to explain
Duplicate ingestion Detect repeated content or source versions
Failed processing Retry safely and track processing status
Partial updates Avoid exposing an inconsistent mixture of versions
Model changes Track which embedding model generated each index
Rollback Retain a usable previous version where requirements allow
Scroll across to read all columns.

Embedding migration trap: Even two models with the same embedding dimension can produce incompatible vector spaces. Use compatible query/document embeddings; typically build a separate index and re-embed documents before switching.

5. Vector search internals

Question: “Why use approximate nearest-neighbor search?”

Exact search compares a query against all indexed vectors. Approximate search trades some retrieval accuracy for lower latency and improved scalability.

Know these terms:

  • Cosine similarity: Compares vector direction.
  • Dot product: Depends on direction and magnitude.
  • Euclidean distance: Measures straight-line distance.
  • HNSW: A graph-based approximate search method.
  • ANN recall: How well approximate search reproduces exact nearest neighbors.

For unit-normalized vectors, dot product equals cosine similarity, and Euclidean distance gives the same ranking.

Follow-up: “Does higher embedding dimension mean better retrieval?”

No. Task fit and measured retrieval quality matter. More dimensions also increase storage and computation.

Distinguish ANN recall from relevance recall: reproducing exact vector neighbors does not prove those neighbors answer the user’s question.

6. Stateful agents and conversation memory

Question: “How do you manage state across a long-running AI workflow?”

“I keep durable workflow state outside the model, persist checkpoints, and define explicit transitions. After interruption, the workflow resumes from recorded state, with safeguards against repeating completed actions.”

Separate:

  • Conversation history: Previous messages.
  • Workflow state: Current step, intermediate results, completed actions.
  • Long-term memory: Selected information retained across sessions.

Do not store everything indefinitely. Consider relevance, user isolation, consent requirements, retention, and deletion.

Subtle point: Summarizing conversation history saves tokens but may lose important details. Preserve critical facts and action records explicitly.

7. Text-to-SQL

Question: “How would you build an assistant that answers questions from a database?”

“I provide relevant schema context, generate a proposed query, validate it, and execute through a restricted database interface. I apply authorization and query limits, then ground the response in returned results.”

Cover:

  • Read-only credentials for analytical use.
  • Authorized tables, rows, and columns.
  • Parameterized values where applicable.
  • Query timeouts and result-size limits.
  • Ambiguous definitions such as “active customer.”
  • Joins, aggregation grain, and duplicate counting.

Common trap: A query can execute successfully and still answer the wrong business question. Evaluate semantic correctness, not just execution success.

8. Data pipelines and delivery guarantees

Question: “How would you process a large volume of documents reliably?”

“I decouple ingestion from processing with a queue, use workers with bounded concurrency, and track document status. Processing is idempotent, retries are limited, and persistent failures go to a dead-letter queue for investigation.”

Be ready to explain:

  • At-least-once delivery: Messages may be delivered more than once.
  • Idempotent processing: Repeated processing does not duplicate effects.
  • Backpressure: Limit intake or concurrency when downstream systems cannot keep up.
  • Dead-letter queue: Holds messages that repeatedly fail.

Avoid casually promising “exactly once.” Explain the boundaries and how duplicate effects are prevented.

9. Business value and deciding whether AI is appropriate

Question: “How do you decide whether a use case needs GenAI?”

“I first define the task and compare a simple baseline against an AI approach. I consider input variability, error tolerance, operational cost, and whether generated output provides measurable value.”

Examples:

Task Useful baseline
Fixed calculations Deterministic code
Known structured rules Rules engine
Predicting a category from labeled examples Conventional ML
Searching documents Search without generation
Interpreting varied language and synthesizing evidence Consider an LLM
Scroll across to read all columns.

Measure business outcomes such as handling time, correction rate, completion rate, and cost per successful task. A technically impressive model may still fail to improve the workflow.

10. Five difficult follow-ups to practise now

Answer these aloud before reading the suggested direction:

  1. “Your evaluation score improved, but users complain more. Why?”
    The dataset may be unrepresentative, metrics may miss important errors, or latency and workflow usability may have worsened. Investigate actual failure cases.

  2. “Would you use multiple agents?”
    Only when decomposition provides measured value. Consider coordination overhead, compounded errors, latency, and harder debugging.

  3. “What if the model says it is 95% confident?”
    Self-reported confidence is not automatically calibrated. Validate uncertainty signals empirically and define escalation criteria.

  4. “How do you reproduce an incorrect answer?”
    Capture relevant versions, authorized context identifiers, retrieval results, generation settings, and tool traces—subject to data handling constraints. Exact replay may still vary.

  5. “What do you do when two authoritative documents disagree?”
    Use documented precedence and effective-date rules. If the conflict remains unresolved, surface it with sources and escalate.

Finish with one sentence you can defend:
“I choose the approach from the requirements, verify it against a baseline, and make its failures observable.”

Take a moment to recall

A short quiz is ready when you want to check your understanding.

Continue exploringModel serving and ML fundamentals
BrowseAccount