Skip to content
StudyHubA place to keep learning

Explore

  • Browse Topics
  • Study Packs
  • Library Topics
    • AI engineering interviews

My Workspace

  • My Notes
StudyHub · The 60-minute crash session
Browse topics
AI engineering interviews

The 60-minute crash session

StudyHub12 min readUpdated Oct 3, 2026

A 60-minute crash session for a senior AI engineering interview. Set a timer and follow the sections in order. Read each answer, then explain it aloud without looking.

Your goal is to sound clear on architecture, tradeoffs, evaluation, production reliability, and your own contribution.

Minutes 0–5: Prepare your introduction and project story

For “Tell me about yourself,” use this structure:

“I have over eight years of IT experience, with experience in data science and GenAI. My strongest areas are [your actual strengths]. In a recent project, I worked on [business problem], where I personally owned [specific components]. We used [architecture] and measured success through [actual metrics]. I’m looking for an AI engineering role where I can contribute to building and operating enterprise AI applications.”

Keep it under 90 seconds. Explain how your experience progressed; avoid listing every tool.

Prepare one project through these seven points:

  1. Problem: Who needed it, and what was inefficient?
  2. Architecture: How did data move through the system?
  3. Ownership: What did you personally design or implement?
  4. Decision: Why this approach over an alternative?
  5. Evaluation: How did you establish that it worked?
  6. Failure: What broke, and how did you investigate?
  7. Outcome: What changed for the business?

Use real numbers where available. If it was a prototype, describe it accurately.

Practise aloud: “Explain your strongest GenAI project in two minutes.”

Minutes 5–17: RAG—the highest-value technical topic

What is RAG?

Retrieval-augmented generation retrieves external information and supplies it as context to a model. It is useful when answers depend on private, changing, or source-specific knowledge.

Explain its two pipelines:

Indexing pipeline Query pipeline
Ingest documents Authenticate user
Parse text, tables, and structure Apply authorization constraints
Clean and split into chunks Retrieve relevant chunks
Attach source, version, and access metadata Optionally rerank
Create embeddings and index Construct prompt and generate
Refresh or remove outdated content Return citations or abstain
Scroll across to read all columns.

Why RAG instead of fine-tuning?

“I’d start with RAG when the requirement is access to changing knowledge, citations, or document-level permissions. Fine-tuning is more appropriate for adapting behavior, style, or task performance. They can be combined.”

Fine-tuning does not provide reliable document retrieval or guarantee factual correctness.

How do you choose chunk size?

“I preserve meaningful document boundaries where possible, then evaluate chunk sizes and overlap against representative questions. Small chunks can lose context; large chunks can dilute relevance and increase prompt cost.”

Preserve headings, table context, source identifiers, and relevant metadata. There is no universal optimal chunk size.

Dense retrieval versus keyword retrieval?

  • Dense retrieval captures semantic similarity.
  • Keyword retrieval, such as BM25, helps with exact terms, identifiers, and uncommon names.
  • Hybrid retrieval combines signals, often through rank fusion.

What does a reranker do?

A reranker scores query–document relevance after initial retrieval. It can improve context selection, with added latency and compute cost.

How would you diagnose poor RAG answers?

Separate the stages:

Symptom Investigate
Relevant content never appears Parsing, chunking, indexing, filters, retrieval
Correct content ranks low Retrieval method, embeddings, reranking
Correct context is supplied but answer is wrong Prompt, model behavior, conflicting context
Answer uses outdated information Document versioning and index refresh
Citation does not support the claim Citation alignment and answer evaluation
Scroll across to read all columns.

A strong answer is:

“First, I check whether the required evidence was retrieved. If not, I investigate retrieval. If it was, I investigate context construction and generation. I change one component at a time and compare against a fixed evaluation set.”

How do you evaluate RAG?

Evaluate retrieval and generation separately.

  • Recall@k: Fraction of relevant items found among the top k results.
  • Ranking quality: Whether useful results appear early.
  • Answer correctness: Whether the answer resolves the question accurately.
  • Groundedness: Whether claims are supported by supplied evidence.
  • Citation accuracy: Whether cited sources support the associated claims.
  • Operational metrics: Latency, cost, errors, and abstention behavior.

Build a representative evaluation set including unanswerable questions, conflicting documents, access restrictions, and difficult tables. Human review helps calibrate automated evaluators.

How do you reduce hallucinations?

Supply relevant evidence, instruct the model to acknowledge insufficient evidence, validate outputs where possible, evaluate unsupported claims, and escalate high-impact cases. Low temperature and RAG alone do not guarantee correctness.

Practise aloud: “Design a RAG assistant and explain how you would diagnose and measure its failures.”

Minutes 17–25: LLM fundamentals and adaptation

Question Interview-ready answer
What is a token? A unit of text representation used by the model; it need not correspond to a whole word.
What is an embedding? A learned numerical representation used to compare inputs or support downstream tasks.
What is self-attention? A mechanism that combines information from token positions using query, key, and value representations.
Why scale attention scores? Dividing by √dₖ helps prevent dot products from growing excessively with dimension and saturating softmax.
What is temperature? A parameter that changes the sharpness of the token probability distribution during sampling.
What is a context window? The maximum sequence length the model supports; a large window does not ensure reliable use of every detail.
What is LoRA? Parameter-efficient adaptation using trainable low-rank updates while keeping base weights frozen.
What is quantization? Representing weights or computations with lower precision to reduce memory and potentially improve efficiency, with possible quality tradeoffs.
Scroll across to read all columns.

Recognize the attention equation:

Attention⁡(Q,K,V)=softmax⁡(QK⊤dk)V\operatorname{Attention}(Q,K,V) =\operatorname{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right)V

Explain it simply: queries compare against keys; the resulting weights combine values. Decoder models use causal masking to prevent attending to future tokens.

How do you select an LLM?

“I benchmark candidate models on representative tasks, then compare quality, latency, cost, context needs, deployment constraints, and data handling requirements.”

Choose the smallest or least expensive model that meets requirements, with escalation where useful.

Prompting versus RAG versus fine-tuning

Requirement Starting approach
Instructions, examples, output format Prompting
Private or changing factual knowledge RAG
Repeated task behavior that prompting cannot adequately achieve Consider fine-tuning
Exact calculations or authoritative system data Tools or deterministic code
Scroll across to read all columns.

Hosted versus self-hosted models

Hosted models reduce infrastructure work. Self-hosting offers deployment control but requires capacity planning, serving, patching, monitoring, and operational expertise. Compare total cost and constraints rather than assuming either is inherently cheaper.

Practise aloud: “When would you fine-tune instead of improving retrieval or prompting?”

Minutes 25–35: Agents, tools, and security

What is an agent?

“An agent uses a model to choose actions or tools, observes the results, and continues toward a goal within defined limits.”

A workflow follows prescribed transitions. Agentic execution allows more model-driven decisions. A practical system can combine both.

When would you use an agent?

Use one when the next action depends on intermediate results and flexible tool selection provides value. Use a deterministic workflow when the sequence is known and predictable execution matters.

What is tool calling?

The model produces a structured request for a tool. Application code validates it, checks authorization, executes the operation, and returns the result. The model’s request is not permission to execute.

How do you make agents reliable?

  • Define narrow tools with validated input schemas.
  • Enforce permissions in application code.
  • Set limits on steps, elapsed time, and spending.
  • Handle timeouts and tool errors.
  • Use idempotency for actions that might be retried.
  • Require approval where the business process demands it.
  • Trace decisions and tool results.
  • Evaluate complete task outcomes.

What is idempotency?

Repeated execution of the same logical request should not create additional effects. For example, a repeated payment request with the same idempotency key should not create another payment.

How do you prevent agent loops?

Use explicit termination conditions, step limits, repeated-action detection, progress checks, and escalation paths.

What is prompt injection?

Untrusted text attempts to redirect model behavior—for example, a retrieved document instructing the assistant to reveal confidential information.

Controls include treating retrieved content as data, restricting tool privileges, validating actions, isolating secrets, and testing adversarial inputs. Prompt instructions alone are insufficient.

How do you enforce access control in RAG?

“I enforce document authorization during retrieval and ensure unauthorized content cannot enter the model context. I also account for permission changes, cache isolation, and auditability.”

Fetching confidential content and asking the model to hide it is inadequate.

Banking scenario: an assistant recommends updating a customer record.

A strong design separates:

  1. Retrieving authorized information.
  2. Producing a proposed change.
  3. Validating the proposal against business rules.
  4. Obtaining required approval.
  5. Executing through an authorized API.
  6. Recording an audit trail.

Practise aloud: “How would you build an agent that can read customer data and safely perform permitted actions?”

Minutes 35–45: System design and production engineering

Use this sequence for almost any design question:

Requirements → scale → architecture → evaluation → security → failure handling → tradeoffs.

Example: “Design an internal policy assistant for a bank.”

Start by clarifying:

  • Who uses it, and which documents can each user access?
  • Does it only answer questions or also take actions?
  • How often do documents change?
  • What are expected traffic, latency, and quality requirements?
  • What happens when evidence is missing or contradictory?

Then describe ingestion, authorized retrieval, generation, citations, monitoring, and index updates.

Common production questions

Question Strong points to cover
How do you reduce latency? Measure stage timings; optimize retrieval, context size, model choice, and unnecessary sequential calls.
How do you handle provider failures? Timeouts, bounded retries with backoff and jitter, circuit breakers, and tested fallback behavior.
How do you control cost? Track tokens and tool usage; reduce unnecessary context; route tasks appropriately; set budgets.
What do you monitor? Errors, p50/p95 latency, retrieval behavior, quality samples, token usage, costs, and tool failures.
How do you deploy changes safely? Version prompts, models, indexes, and code; run evaluation gates; use gradual rollout and rollback.
How do you scale? Stateless API replicas where appropriate, asynchronous ingestion, backpressure, and capacity planning.
Scroll across to read all columns.

Caching nuance: Cache keys must account for authorization, relevant versions, and freshness. A cache must not expose one user’s data to another.

Fallback nuance: A second model may have different quality, safety, or data handling characteristics. Benchmark fallback behavior before relying on it.

Async nuance: Async helps overlap waiting for network I/O. CPU-heavy work can still block the event loop and may need separate workers or processes.

How would you investigate a latency incident?

“I compare latency percentiles and trace time spent in retrieval, reranking, model generation, and tools. I check recent changes and downstream errors, mitigate the dominant bottleneck, then verify recovery.”

Practise aloud: Give a three-minute policy-assistant design, including one failure path and one tradeoff.

Minutes 45–52: Python, SQL, and ML refresh

Know these Python distinctions:

  • == compares equality; is compares identity.
  • Generators produce values lazily and can reduce memory use.
  • Mutable default arguments are created once, so they can retain state between calls.
  • Catch specific exceptions and preserve useful error context.
  • Threads can help with I/O; processes can help with CPU-bound Python work. Runtime and library behavior matter.

Fix the mutable default problem:

Python
def add_item(item, items=None):
    if items is None:
        items = []
    items.append(item)
    return items

Mini coding exercise: Deduplicate strings while preserving order.

Python
def unique_in_order(values):
    seen = set()
    result = []

    for value in values:
        if value not in seen:
            seen.add(value)
            result.append(value)

    return result

Explain: expected O(n) time, O(n) extra space, and the approach requires hashable values.

SQL exercise: Get the latest record per customer.

SQL
WITH ranked AS (
    SELECT
        customer_id,
        event_id,
        event_time,
        status,
        ROW_NUMBER() OVER (
            PARTITION BY customer_id
            ORDER BY event_time DESC, event_id DESC
        ) AS rn
    FROM customer_events
)
SELECT customer_id, event_id, event_time, status
FROM ranked
WHERE rn = 1;

The secondary sort makes equal timestamps deterministic, assuming event_id distinguishes records.

Refresh INNER JOIN versus LEFT JOIN, WHERE versus HAVING, window functions versus aggregation, and how indexes affect reads and writes.

ML fundamentals you should answer quickly

Topic Core answer
Data leakage Training or evaluation uses information unavailable at prediction time. Split appropriately and fit preprocessing only on training data.
Imbalanced classification Accuracy can mislead. Choose metrics and thresholds based on false-positive and false-negative costs.
Precision versus recall Precision measures how many positive predictions are correct; recall measures how many actual positives are found.
Drift Inputs or relationships can change. Monitor data and outcomes; investigate before automatically retraining.
Overfitting Performance is strong on training data but generalizes poorly. Examine validation design, complexity, regularization, and data.
Scroll across to read all columns.

Minutes 52–57: Seniority and behavioral answers

Prepare three real stories:

  • A difficult technical tradeoff.
  • A failed experiment or production problem.
  • A stakeholder disagreement or delivery constraint.

Use situation → your responsibility → actions → result → lesson.

For “Why did you choose this architecture?”:

“The requirements were [X]. We compared [A] and [B] using [criteria]. I chose [A] because [evidence]. Its main downside was [tradeoff], which we managed through [mitigation].”

For “What was your contribution?”:

“I owned [components and decisions]. I collaborated with [teams] on [dependencies]. The broader team delivered [outcome].”

For an unfamiliar question:

“I haven’t implemented that directly. My understanding is [what you know]. I would validate [uncertainty] and compare [options] against the requirements.”

Show reasoning without inventing experience.

Minutes 57–60: Final recall drill

Answer each aloud in one or two sentences:

  1. When would you choose RAG over fine-tuning?
  2. How do you distinguish retrieval failure from generation failure?
  3. Which metrics establish whether a RAG system works?
  4. When is an agent justified?
  5. Where are authorization and tool permissions enforced?
  6. How do retries avoid duplicate side effects?
  7. How do you investigate high p95 latency?
  8. What did you personally own in your strongest project?
  9. Which tradeoff did you make, and what evidence supported it?
  10. What would you improve if you rebuilt that system?

If you remember one answer pattern, use this:

Requirement → decision → alternative → tradeoff → measurement.

That pattern helps turn a technically correct answer into a convincing senior engineering answer.

Take a moment to recall

A short quiz is ready when you want to check your understanding.

Continue exploringTesting and enterprise workflows
BrowseAccount