A 60-minute crash session for a senior AI engineering interview. Set a timer and follow the sections in order. Read each answer, then explain it aloud without looking.
Your goal is to sound clear on architecture, tradeoffs, evaluation, production reliability, and your own contribution.
Minutes 0–5: Prepare your introduction and project story
For “Tell me about yourself,” use this structure:
“I have over eight years of IT experience, with experience in data science and GenAI. My strongest areas are [your actual strengths]. In a recent project, I worked on [business problem], where I personally owned [specific components]. We used [architecture] and measured success through [actual metrics]. I’m looking for an AI engineering role where I can contribute to building and operating enterprise AI applications.”
Keep it under 90 seconds. Explain how your experience progressed; avoid listing every tool.
Prepare one project through these seven points:
- Problem: Who needed it, and what was inefficient?
- Architecture: How did data move through the system?
- Ownership: What did you personally design or implement?
- Decision: Why this approach over an alternative?
- Evaluation: How did you establish that it worked?
- Failure: What broke, and how did you investigate?
- Outcome: What changed for the business?
Use real numbers where available. If it was a prototype, describe it accurately.
Practise aloud: “Explain your strongest GenAI project in two minutes.”
Minutes 5–17: RAG—the highest-value technical topic
What is RAG?
Retrieval-augmented generation retrieves external information and supplies it as context to a model. It is useful when answers depend on private, changing, or source-specific knowledge.
Explain its two pipelines:
| Indexing pipeline | Query pipeline |
|---|---|
| Ingest documents | Authenticate user |
| Parse text, tables, and structure | Apply authorization constraints |
| Clean and split into chunks | Retrieve relevant chunks |
| Attach source, version, and access metadata | Optionally rerank |
| Create embeddings and index | Construct prompt and generate |
| Refresh or remove outdated content | Return citations or abstain |
Why RAG instead of fine-tuning?
“I’d start with RAG when the requirement is access to changing knowledge, citations, or document-level permissions. Fine-tuning is more appropriate for adapting behavior, style, or task performance. They can be combined.”
Fine-tuning does not provide reliable document retrieval or guarantee factual correctness.
How do you choose chunk size?
“I preserve meaningful document boundaries where possible, then evaluate chunk sizes and overlap against representative questions. Small chunks can lose context; large chunks can dilute relevance and increase prompt cost.”
Preserve headings, table context, source identifiers, and relevant metadata. There is no universal optimal chunk size.
Dense retrieval versus keyword retrieval?
- Dense retrieval captures semantic similarity.
- Keyword retrieval, such as BM25, helps with exact terms, identifiers, and uncommon names.
- Hybrid retrieval combines signals, often through rank fusion.
What does a reranker do?
A reranker scores query–document relevance after initial retrieval. It can improve context selection, with added latency and compute cost.
How would you diagnose poor RAG answers?
Separate the stages:
| Symptom | Investigate |
|---|---|
| Relevant content never appears | Parsing, chunking, indexing, filters, retrieval |
| Correct content ranks low | Retrieval method, embeddings, reranking |
| Correct context is supplied but answer is wrong | Prompt, model behavior, conflicting context |
| Answer uses outdated information | Document versioning and index refresh |
| Citation does not support the claim | Citation alignment and answer evaluation |
A strong answer is:
“First, I check whether the required evidence was retrieved. If not, I investigate retrieval. If it was, I investigate context construction and generation. I change one component at a time and compare against a fixed evaluation set.”
How do you evaluate RAG?
Evaluate retrieval and generation separately.
- Recall@k: Fraction of relevant items found among the top k results.
- Ranking quality: Whether useful results appear early.
- Answer correctness: Whether the answer resolves the question accurately.
- Groundedness: Whether claims are supported by supplied evidence.
- Citation accuracy: Whether cited sources support the associated claims.
- Operational metrics: Latency, cost, errors, and abstention behavior.
Build a representative evaluation set including unanswerable questions, conflicting documents, access restrictions, and difficult tables. Human review helps calibrate automated evaluators.
How do you reduce hallucinations?
Supply relevant evidence, instruct the model to acknowledge insufficient evidence, validate outputs where possible, evaluate unsupported claims, and escalate high-impact cases. Low temperature and RAG alone do not guarantee correctness.
Practise aloud: “Design a RAG assistant and explain how you would diagnose and measure its failures.”
Minutes 17–25: LLM fundamentals and adaptation
| Question | Interview-ready answer |
|---|---|
| What is a token? | A unit of text representation used by the model; it need not correspond to a whole word. |
| What is an embedding? | A learned numerical representation used to compare inputs or support downstream tasks. |
| What is self-attention? | A mechanism that combines information from token positions using query, key, and value representations. |
| Why scale attention scores? | Dividing by √dₖ helps prevent dot products from growing excessively with dimension and saturating softmax. |
| What is temperature? | A parameter that changes the sharpness of the token probability distribution during sampling. |
| What is a context window? | The maximum sequence length the model supports; a large window does not ensure reliable use of every detail. |
| What is LoRA? | Parameter-efficient adaptation using trainable low-rank updates while keeping base weights frozen. |
| What is quantization? | Representing weights or computations with lower precision to reduce memory and potentially improve efficiency, with possible quality tradeoffs. |
Recognize the attention equation:
Explain it simply: queries compare against keys; the resulting weights combine values. Decoder models use causal masking to prevent attending to future tokens.
How do you select an LLM?
“I benchmark candidate models on representative tasks, then compare quality, latency, cost, context needs, deployment constraints, and data handling requirements.”
Choose the smallest or least expensive model that meets requirements, with escalation where useful.
Prompting versus RAG versus fine-tuning
| Requirement | Starting approach |
|---|---|
| Instructions, examples, output format | Prompting |
| Private or changing factual knowledge | RAG |
| Repeated task behavior that prompting cannot adequately achieve | Consider fine-tuning |
| Exact calculations or authoritative system data | Tools or deterministic code |
Hosted versus self-hosted models
Hosted models reduce infrastructure work. Self-hosting offers deployment control but requires capacity planning, serving, patching, monitoring, and operational expertise. Compare total cost and constraints rather than assuming either is inherently cheaper.
Practise aloud: “When would you fine-tune instead of improving retrieval or prompting?”
Minutes 25–35: Agents, tools, and security
What is an agent?
“An agent uses a model to choose actions or tools, observes the results, and continues toward a goal within defined limits.”
A workflow follows prescribed transitions. Agentic execution allows more model-driven decisions. A practical system can combine both.
When would you use an agent?
Use one when the next action depends on intermediate results and flexible tool selection provides value. Use a deterministic workflow when the sequence is known and predictable execution matters.
What is tool calling?
The model produces a structured request for a tool. Application code validates it, checks authorization, executes the operation, and returns the result. The model’s request is not permission to execute.
How do you make agents reliable?
- Define narrow tools with validated input schemas.
- Enforce permissions in application code.
- Set limits on steps, elapsed time, and spending.
- Handle timeouts and tool errors.
- Use idempotency for actions that might be retried.
- Require approval where the business process demands it.
- Trace decisions and tool results.
- Evaluate complete task outcomes.
What is idempotency?
Repeated execution of the same logical request should not create additional effects. For example, a repeated payment request with the same idempotency key should not create another payment.
How do you prevent agent loops?
Use explicit termination conditions, step limits, repeated-action detection, progress checks, and escalation paths.
What is prompt injection?
Untrusted text attempts to redirect model behavior—for example, a retrieved document instructing the assistant to reveal confidential information.
Controls include treating retrieved content as data, restricting tool privileges, validating actions, isolating secrets, and testing adversarial inputs. Prompt instructions alone are insufficient.
How do you enforce access control in RAG?
“I enforce document authorization during retrieval and ensure unauthorized content cannot enter the model context. I also account for permission changes, cache isolation, and auditability.”
Fetching confidential content and asking the model to hide it is inadequate.
Banking scenario: an assistant recommends updating a customer record.
A strong design separates:
- Retrieving authorized information.
- Producing a proposed change.
- Validating the proposal against business rules.
- Obtaining required approval.
- Executing through an authorized API.
- Recording an audit trail.
Practise aloud: “How would you build an agent that can read customer data and safely perform permitted actions?”
Minutes 35–45: System design and production engineering
Use this sequence for almost any design question:
Requirements → scale → architecture → evaluation → security → failure handling → tradeoffs.
Example: “Design an internal policy assistant for a bank.”
Start by clarifying:
- Who uses it, and which documents can each user access?
- Does it only answer questions or also take actions?
- How often do documents change?
- What are expected traffic, latency, and quality requirements?
- What happens when evidence is missing or contradictory?
Then describe ingestion, authorized retrieval, generation, citations, monitoring, and index updates.
Common production questions
| Question | Strong points to cover |
|---|---|
| How do you reduce latency? | Measure stage timings; optimize retrieval, context size, model choice, and unnecessary sequential calls. |
| How do you handle provider failures? | Timeouts, bounded retries with backoff and jitter, circuit breakers, and tested fallback behavior. |
| How do you control cost? | Track tokens and tool usage; reduce unnecessary context; route tasks appropriately; set budgets. |
| What do you monitor? | Errors, p50/p95 latency, retrieval behavior, quality samples, token usage, costs, and tool failures. |
| How do you deploy changes safely? | Version prompts, models, indexes, and code; run evaluation gates; use gradual rollout and rollback. |
| How do you scale? | Stateless API replicas where appropriate, asynchronous ingestion, backpressure, and capacity planning. |
Caching nuance: Cache keys must account for authorization, relevant versions, and freshness. A cache must not expose one user’s data to another.
Fallback nuance: A second model may have different quality, safety, or data handling characteristics. Benchmark fallback behavior before relying on it.
Async nuance: Async helps overlap waiting for network I/O. CPU-heavy work can still block the event loop and may need separate workers or processes.
How would you investigate a latency incident?
“I compare latency percentiles and trace time spent in retrieval, reranking, model generation, and tools. I check recent changes and downstream errors, mitigate the dominant bottleneck, then verify recovery.”
Practise aloud: Give a three-minute policy-assistant design, including one failure path and one tradeoff.
Minutes 45–52: Python, SQL, and ML refresh
Know these Python distinctions:
==compares equality;iscompares identity.- Generators produce values lazily and can reduce memory use.
- Mutable default arguments are created once, so they can retain state between calls.
- Catch specific exceptions and preserve useful error context.
- Threads can help with I/O; processes can help with CPU-bound Python work. Runtime and library behavior matter.
Fix the mutable default problem:
def add_item(item, items=None):
if items is None:
items = []
items.append(item)
return items
Mini coding exercise: Deduplicate strings while preserving order.
def unique_in_order(values):
seen = set()
result = []
for value in values:
if value not in seen:
seen.add(value)
result.append(value)
return result
Explain: expected O(n) time, O(n) extra space, and the approach requires hashable values.
SQL exercise: Get the latest record per customer.
WITH ranked AS (
SELECT
customer_id,
event_id,
event_time,
status,
ROW_NUMBER() OVER (
PARTITION BY customer_id
ORDER BY event_time DESC, event_id DESC
) AS rn
FROM customer_events
)
SELECT customer_id, event_id, event_time, status
FROM ranked
WHERE rn = 1;
The secondary sort makes equal timestamps deterministic, assuming event_id distinguishes records.
Refresh INNER JOIN versus LEFT JOIN, WHERE versus HAVING, window functions versus aggregation, and how indexes affect reads and writes.
ML fundamentals you should answer quickly
| Topic | Core answer |
|---|---|
| Data leakage | Training or evaluation uses information unavailable at prediction time. Split appropriately and fit preprocessing only on training data. |
| Imbalanced classification | Accuracy can mislead. Choose metrics and thresholds based on false-positive and false-negative costs. |
| Precision versus recall | Precision measures how many positive predictions are correct; recall measures how many actual positives are found. |
| Drift | Inputs or relationships can change. Monitor data and outcomes; investigate before automatically retraining. |
| Overfitting | Performance is strong on training data but generalizes poorly. Examine validation design, complexity, regularization, and data. |
Minutes 52–57: Seniority and behavioral answers
Prepare three real stories:
- A difficult technical tradeoff.
- A failed experiment or production problem.
- A stakeholder disagreement or delivery constraint.
Use situation → your responsibility → actions → result → lesson.
For “Why did you choose this architecture?”:
“The requirements were [X]. We compared [A] and [B] using [criteria]. I chose [A] because [evidence]. Its main downside was [tradeoff], which we managed through [mitigation].”
For “What was your contribution?”:
“I owned [components and decisions]. I collaborated with [teams] on [dependencies]. The broader team delivered [outcome].”
For an unfamiliar question:
“I haven’t implemented that directly. My understanding is [what you know]. I would validate [uncertainty] and compare [options] against the requirements.”
Show reasoning without inventing experience.
Minutes 57–60: Final recall drill
Answer each aloud in one or two sentences:
- When would you choose RAG over fine-tuning?
- How do you distinguish retrieval failure from generation failure?
- Which metrics establish whether a RAG system works?
- When is an agent justified?
- Where are authorization and tool permissions enforced?
- How do retries avoid duplicate side effects?
- How do you investigate high p95 latency?
- What did you personally own in your strongest project?
- Which tradeoff did you make, and what evidence supported it?
- What would you improve if you rebuilt that system?
If you remember one answer pattern, use this:
Requirement → decision → alternative → tradeoff → measurement.
That pattern helps turn a technically correct answer into a convincing senior engineering answer.