Skip to content
StudyHubA place to keep learning

Explore

  • Browse Topics
  • Study Packs
  • Library Topics
    • AI engineering interviews

My Workspace

  • My Notes
StudyHub · Final rehearsal
Browse topics
AI engineering interviews

Final rehearsal

StudyHub9 min readUpdated Oct 3, 2026

Practise training loops, tuning, retrieval, evaluation, and supply-chain security, then complete the independent rehearsal. Prioritize the first four topics.

1. Write and explain a basic training loop

For a PyTorch model, recognize this sequence:

Python
model.train()

for inputs, targets in loader:
    optimizer.zero_grad()

    logits = model(inputs)
    loss = loss_fn(logits, targets)

    loss.backward()
    optimizer.step()

Explain each step:

  • Clear gradients from the previous update.
  • Run the forward pass.
  • Calculate the loss.
  • Compute gradients through backpropagation.
  • Update parameters using the optimizer.

Important traps:

  • Gradients accumulate unless cleared deliberately.
  • model.eval() changes behaviors such as dropout and batch normalization; it does not itself disable gradient tracking.
  • Use torch.no_grad() during ordinary evaluation when gradients are unnecessary.
  • For standard multiclass classification with CrossEntropyLoss, pass logits rather than applying softmax first. [10]

“The loss is not decreasing. What would you check?”

“I check target correctness, tensor shapes, trainable parameters, gradient flow, learning rate, preprocessing, and whether the optimizer contains the intended parameters. I also try overfitting a tiny batch to verify the training path.”

Failure to fit a tiny batch often points toward an implementation or optimization problem.

2. Hyperparameter tuning without contaminating evaluation

Understand how tuning interacts with train/test splitting.

“Why shouldn’t you repeatedly select models using the test set?”

Because your decisions gradually adapt to that set. Its score becomes an optimistic estimate of generalization.

Use:

  • Training data to fit parameters.
  • Validation or cross-validation to choose configurations.
  • An independent test set for the final assessment.

“Why put preprocessing inside a pipeline?”

“During cross-validation, preprocessing must be fitted on each training fold and then applied to its validation fold. A pipeline helps keep these operations together and prevents leakage.”

This applies to imputation, scaling, feature selection, and other learned transformations. [11, 12]

Search method Useful distinction
Grid search Exhaustively evaluates specified combinations
Random search Samples configurations within a budget
Bayesian optimization Uses previous trials to guide subsequent choices
Scroll across to read all columns.

Nested cross-validation: An inner loop selects configurations; an outer loop estimates the performance of that selection procedure.

For time-dependent data, preserve the temporal structure in the tuning process.

3. SLI, SLO, SLA, and error budgets

These move your production answers beyond “we monitored latency.”

Term Meaning
SLI The measured indicator, such as successful-request ratio
SLO The target for that indicator over a defined window
SLA An agreement specifying service commitments and consequences
Error budget The amount of failure permitted by the SLO
Scroll across to read all columns.

For a request-success SLO of 99.9% over one million eligible requests, the error budget is 1,000 unsuccessful requests. Define what counts as an eligible request and a failure. [13, 14]

“What should an AI service’s reliability target include?”

“I distinguish operational success from answer quality. A request can return HTTP 200 while producing an unusable answer. I define latency and availability objectives alongside task-quality acceptance criteria.”

Subtle trap: You generally cannot obtain end-to-end p95 latency by adding each stage’s p95. Measure the distribution of complete requests.

4. Software design: make the AI system replaceable and testable

“How would you structure the codebase?”

Describe clear responsibilities:

Component Responsibility
API layer Requests, identity, validation, response contracts
Application service Coordinates the business workflow
Model adapter Handles model-provider interaction
Retrieval interface Searches and returns evidence
Tool interfaces Perform authorized external operations
Persistence layer Stores durable application state
Evaluation suite Measures behavior against representative cases
Scroll across to read all columns.

“What is dependency injection?”

“The application receives its dependencies rather than creating them internally. I can supply a real model client in production and a controlled fake during testing.”

Why it matters: It helps you test failure behavior, replace providers, and separate business rules from framework details.

“Microservices or a modular monolith?”

“I start with the simplest architecture that meets the requirements. Separate services are justified when components need independent scaling, ownership, deployment, or isolation.”

More services also introduce network failures and operational work.

5. Reproducibility means more than a random seed

“What would you record for an experiment?”

  • Dataset version and split definition.
  • Preprocessing and feature definitions.
  • Code version and configuration.
  • Model or checkpoint version.
  • Hyperparameters and random seeds.
  • Evaluation procedure and results.
  • Relevant environment and dependency versions.

For an AI application, also record prompt versions, retrieval settings, index versions, and tool configuration.

“Will a seed guarantee identical results?”

No. Different hardware, kernels, distributed execution, or nondeterministic operations can affect results.

Strong answer:

“I preserve enough provenance to explain what changed and reproduce the experiment as closely as practical. I distinguish statistical repeatability from exact numerical equality.”

6. Graph-based retrieval versus vector retrieval

“When would a knowledge graph help?”

When the question depends on explicit relationships, such as:

“Which suppliers are connected to this subsidiary through contracts expiring this quarter?”

A graph can represent entities and relationships directly. Vector retrieval finds semantically similar content.

Requirement Useful representation
Find passages discussing a topic Vector or keyword search
Follow explicit entity relationships Graph traversal
Filter and aggregate exact records Relational queries
Combine narrative evidence with relationships A hybrid design
Scroll across to read all columns.

Tradeoffs: Building a graph requires entity resolution, relationship extraction, provenance, and updates. Incorrect relationships can produce confident-looking errors.

Trap: Calling a system “GraphRAG” does not establish that its graph is accurate or that it improves the task.

7. Three adaptation terms that are easy to confuse

Term What changes?
Prompt engineering Human-written instructions and examples
Prompt tuning Learned continuous prompt representations
Fine-tuning Some or all model parameters
Scroll across to read all columns.

LoRA and QLoRA are parameter-efficient fine-tuning approaches.

Training detail: loss masking

For assistant-style supervised training, you may train on desired assistant outputs while excluding prompt or padding positions from the loss. The choice depends on the training objective.

“Why can incorrect masking hurt?”

The model can spend training effort optimizing positions that do not represent the desired prediction task, or learn from padding incorrectly.

Also ensure that training examples use the appropriate conversation formatting for the model.

8. Embedding similarity is not calibrated confidence

“The retrieved chunk has cosine similarity 0.85. Is the answer 85% likely to be correct?”

No.

“Similarity measures a relationship in the embedding space. It is not automatically a probability of relevance or answer correctness.”

A threshold needs empirical validation for the specific model, corpus, and task.

Similarity scores can change after:

  • Switching embedding models.
  • Changing document representation.
  • Changing the corpus.
  • Changing query distributions.

Another trap: Embeddings should not automatically be treated as anonymized data. They remain representations derived from source data and need appropriate access and handling controls.

9. Security beyond the prompt

Alongside prompt injection and authorization, recognize the software supply chain.

“What would you check before deploying a downloaded model or component?”

“I establish its provenance, review the loading and execution path, manage dependencies, and restrict runtime permissions. I treat executable model-loading code and serialized artifacts according to their trust level.”

Important distinctions:

  • A model artifact may include more than numerical weights.
  • Some serialization formats can execute code during loading.
  • A checksum confirms a match to an expected artifact; it does not prove the artifact is safe.
  • Credentials belong in controlled runtime configuration, not prompts, source code, or container images.

“How do you handle sensitive logs?”

Log the information needed for diagnosis, with deliberate redaction, access controls, retention, and sampling. Avoid making unrestricted prompt logging the default.

10. Estimation and answering under uncertainty

“How long will this project take?”

“I break it into data readiness, integration, evaluation, deployment, and operational acceptance. I identify dependencies and uncertainties, then provide an estimate with assumptions and checkpoints.”

“Can you guarantee 95% accuracy?”

“I first define accuracy, the target population, and the evaluation procedure. I establish a baseline and assess whether the target is achievable. I also examine failure severity and performance on important cases.”

“Which framework or model is best?”

“Best depends on the requirements. I compare a small set against representative tasks and operational constraints, then choose from the evidence.”

At senior level, uncertainty should lead to a validation plan, rather than an unsupported number.

Finish with this closed-book rehearsal

Answer each question before checking the expected points.

Interview question Your answer should contain
“Explain your project in two minutes.” Business problem, architecture, your contribution, evidence, outcome
“Why that architecture?” Requirements, alternative, tradeoff, measurement
“How do you know retrieval is working?” Representative queries, relevance judgments, retrieval metrics, failure analysis
“How do you safely execute an agent action?” Validated inputs, application authorization, bounded execution, idempotency, auditability
“What happens during a provider outage?” Deadlines, bounded retries, circuit breaking, tested fallback or graceful failure
“How do you prevent ML leakage?” Deployment-relevant split, feature availability, preprocessing within training partitions
“How do you release a change?” Versioned artifacts, evaluation gates, gradual rollout, monitoring, rollback
“What would you improve?” A specific weakness, supporting evidence, proposed change, success measure
Scroll across to read all columns.

Then write one Python solution and one SQL query from memory. Check their edge cases and explain their complexity or execution behavior.

Use this answer structure during the interview:

“The requirement is ___. I would choose ___ because ___. The main alternative is ___. The tradeoff is ___. I would validate it using ___ and handle failure through ___.”

If you have implemented the approach, replace “I would” with what you actually did and measured. If you have not, say so plainly.

Your final ten minutes: rehearse your introduction, strongest project, hardest technical decision, and one failure story aloud. Clear, defensible answers to those four give the interviewer a concrete basis for assessing your experience.

Take a moment to recall

A short quiz is ready when you want to check your understanding.

Continue exploringReferences
BrowseAccount