Develop core ML depth, practical backend engineering, coding patterns, and senior ownership. These topics connect model decisions to dependable production systems.
Read these once, then spend your remaining time answering aloud and writing code without looking.
1. Choosing and explaining traditional ML models
Be ready to explain why you would use a simpler model before an LLM.
| Model | Why choose it? | Main limitation |
|---|---|---|
| Linear regression | Continuous prediction with a simple, interpretable baseline | Linear relationship assumptions; sensitivity to outliers |
| Logistic regression | Classification with interpretable coefficients | Linear decision boundary in the supplied feature space |
| Decision tree | Nonlinear rules and feature interactions | Can overfit and change substantially with small data changes |
| Random forest | Robust tabular baseline; reduces variance through averaging | Less transparent and potentially expensive |
| Gradient boosting | Strong tabular performance through sequential improvement | Requires careful tuning and validation |
“Random forest versus boosting?”
“Random forests train randomized trees and aggregate their predictions, primarily reducing variance. Boosting builds models sequentially, with later models improving the ensemble’s objective. I compare them using deployment-relevant validation, latency, and interpretability requirements.”
“When is feature scaling necessary?”
Scaling matters for distance-based methods, many gradient-based methods, and regularized linear models. Standard decision-tree splits generally do not require scaling.
“Can logistic regression capture nonlinear relationships?”
Yes, through suitable nonlinear features and interactions. It remains linear in its supplied feature representation.
2. Loss functions, optimization, and regularization
“Loss function versus evaluation metric?”
“The loss guides optimization during training. The evaluation metric measures performance against the task or business objective. They need not be identical.”
Examples:
- Regression: mean squared error penalizes large errors strongly; mean absolute error is less sensitive to extreme residuals.
- Classification: cross-entropy penalizes assigning low probability to the correct class.
- Retrieval: contrastive objectives help distinguish relevant pairs from negatives.
“What happens if the learning rate is too high or too low?”
Too high can cause unstable updates or failure to converge. Too low can make training slow or stall progress within the available budget.
“L1 versus L2 regularization?”
- L1 penalizes absolute coefficient values and can produce sparse coefficients.
- L2 penalizes squared coefficient values and tends to shrink them smoothly.
“What is backpropagation?”
“Backpropagation applies the chain rule to compute loss gradients with respect to model parameters. The optimizer uses those gradients to update the parameters.”
Also recognize:
- Early stopping: Stop when validation performance no longer improves.
- Dropout: Randomly suppress activations during training.
- Gradient clipping: Bound gradients to help control unstable updates.
- Training versus evaluation mode: Layers such as dropout and batch normalization behave differently.
3. Fine-tuning: explain the actual engineering process
Be ready to explain both when to fine-tune and how to implement it.
“Walk me through a fine-tuning project.”
“I establish a baseline and measurable objective, curate representative training examples, remove duplicates and leakage, and create independent validation and test sets. I choose an adaptation method based on compute and task needs, train while monitoring validation performance, then compare against the baseline on quality, regressions, latency, and cost.”
Know the failure modes:
| Failure | What to investigate |
|---|---|
| Training improves; validation worsens | Overfitting, poor split, insufficient diversity |
| Model reproduces inconsistent behavior | Conflicting or low-quality training examples |
| Target task improves; other tasks worsen | Narrow adaptation and capability regression |
| Evaluation looks unusually strong | Duplicate examples or evaluation contamination |
“How much training data do you need?”
There is no universal count. It depends on task complexity, example quality, diversity, and the starting model. Use learning curves and held-out evaluation.
Important distinction: Fine-tuning changes model parameters. In-context learning uses examples in the prompt without updating parameters.
4. Backend engineering: APIs, authentication, and authorization
“Design an API for an AI application.”
Explain the request contract, validation, authentication, authorization, execution, response contract, and error handling.
| Situation | Typical HTTP response |
|---|---|
| Successful synchronous request | 200 |
| Resource created | 201 |
| Long-running task accepted | 202, with a job identifier |
| Missing or invalid authentication | 401 |
| Authenticated but insufficient permission | 403 |
| Rate limit exceeded | 429 |
| Invalid input | An appropriate 4xx, according to the API contract |
“Authentication versus authorization?”
Authentication establishes identity. Authorization determines what that identity may access or do.
“Is decoding a JWT enough?”
No. Decoding reads its contents. Verification must check its signature and relevant claims, such as expiry, issuer, and audience.
“Why might an async API perform poorly?”
“An async handler can still block its event loop if it performs synchronous network calls or CPU-heavy work. I inspect dependencies and move suitable work to asynchronous clients or separate workers.”
For long-running ingestion, a job API is often more suitable than holding an HTTP request open.
5. Database depth: indexes, transactions, and race conditions
“Why is this SQL query slow?”
“I inspect the execution plan and actual workload. I look for large scans, expensive joins or sorts, poor cardinality estimates, and missing or unsuitable indexes. I verify the improvement against representative data.”
Indexes speed some reads but consume storage and add write overhead. Composite index usefulness depends on query predicates, ordering, and the database.
“What does ACID mean?”
- Atomicity: A transaction completes as a unit or is rolled back.
- Consistency: Transactions preserve defined integrity rules.
- Isolation: Concurrent transactions interact according to the isolation level.
- Durability: Committed changes survive the failures covered by the database’s guarantees.
“Two workers both see a job as unprocessed. What prevents duplicate work?”
A read followed by an update can race. Use an atomic claim operation, suitable locking, or a database constraint and transaction.
“I put concurrency protection at the shared source of truth. An in-memory lock in one API process cannot coordinate separate replicas.”
6. Coding patterns worth recognizing immediately
Do not memorize isolated solutions. Identify the pattern.
| Problem | Useful pattern | Typical complexity |
|---|---|---|
| Frequency counting or two-sum lookup | Hash map | Expected O(n) |
| Longest substring without repetition | Sliding window | O(n) |
| Top k items | Heap | O(n log k) |
| Shortest path in an unweighted graph | BFS | O(V + E) |
| Dependency ordering | Topological sort | O(V + E) |
| Repeated calculations with overlapping subproblems | Dynamic programming | Depends on states and transitions |
Practise this now: longest substring without repeating characters.
def longest_unique_length(text):
last_seen = {}
left = 0
best = 0
for right, character in enumerate(text):
previous = last_seen.get(character)
if previous is not None and previous >= left:
left = previous + 1
last_seen[character] = right
best = max(best, right - left + 1)
return best
Explain:
“I maintain a window containing unique characters. When a character repeats inside the window, I move the left boundary past its previous occurrence.”
Expected O(n) time; space depends on the number of distinct characters.
Check mentally: "" → 0, "aaaa" → 1, "abba" → 2.
During coding: clarify inputs, give an approach, implement, check edge cases, then state complexity.
7. Distributed systems: the timeout ambiguity
“An external tool times out. Can you retry?”
A timeout means the caller did not receive a timely result. It does not prove the operation failed. The remote system may have completed it.
“For reads, retries are usually straightforward. For writes, I use an idempotency key or check operation status before repeating an action.”
“How do you keep a database update and an emitted event consistent?”
A transaction cannot automatically cover an unrelated database and messaging service.
One approach is the transactional outbox:
- Update business data and insert an outbox record in the same database transaction.
- A publisher sends the outbox event.
- Consumers tolerate duplicate delivery.
“What is a circuit breaker?”
It temporarily stops calls to a repeatedly failing dependency, allows limited recovery probes, and resumes normal calls when recovery is established.
Distinguish it from a retry: retries make another attempt; a circuit breaker limits attempts during sustained failure.
8. Statistics and proving that an improvement is real
“Model B scores 84%; model A scores 82%. Is B better?”
“I need the sample size, evaluation design, uncertainty, and error distribution. I compare models on the same representative cases and examine whether the improvement is practically meaningful.”
Remember:
- A small evaluation set can produce unstable estimates.
- An aggregate gain can hide deterioration for important user groups.
- Repeatedly tuning against the test set makes it less independent.
- Statistical significance does not establish business value.
- Offline improvement does not guarantee better user outcomes.
For a controlled online experiment, define the assignment unit, success metric, guardrails, and analysis plan before examining results.
“How would you prove productivity improvement?”
Measure comparable tasks, completion quality, total time including corrections, and adoption. A faster initial answer can still increase overall work if users must repair it.
9. Handling senior-level scenario questions
These assess judgment more than terminology.
“The business wants to launch, but quality is below target.”
“I identify which failure categories matter, quantify their impact, and propose a scoped release where acceptance criteria are met. I define monitoring and rollback conditions and make unresolved risks explicit.”
“A stakeholder asks for an agent, but a workflow is sufficient.”
“I translate the request into required capabilities, compare approaches through a small benchmark, and recommend the one that meets the requirements with acceptable complexity and operating cost.”
“A production change causes wrong answers.”
“I contain the impact, roll back or disable the affected path, preserve relevant evidence, and communicate the scope. Then I investigate the cause and add a regression check addressing the failure.”
“You inherited a poorly documented AI system.”
“I establish the data flow, dependencies, access controls, deployment process, and current performance. I reproduce key behavior and prioritize changes based on observed failure and business impact.”
Use “I” for your actions and “we” for shared decisions and outcomes.
10. Questions to ask the interviewer
Choose two or three:
- “What would successful delivery look like in the first three months?”
- “Is the immediate priority experimentation, production delivery, or improving an existing platform?”
- “How are AI quality criteria defined and approved before deployment?”
- “Which responsibilities belong to this role versus platform, data, and security teams?”
- “What is the hardest engineering problem the team is currently facing?”
These help you understand the role and discuss where your experience fits.
Your final preparation task
Close the notes and complete this 12-minute rehearsal:
| Time | Task |
|---|---|
| 2 minutes | Introduce yourself and describe your strongest project |
| 3 minutes | Defend one architectural decision, including its alternative and evidence |
| 3 minutes | Write the sliding-window solution from memory |
| 2 minutes | Explain how you would investigate a production failure |
| 2 minutes | Describe a disagreement, your actions, and the outcome |
Carry these five habits into the interview:
- Clarify the requirement before choosing technology.
- Explain mechanisms, not just framework names.
- Connect choices to evidence and measurable outcomes.
- Cover failure behavior and operational constraints.
- Be precise about what you personally implemented.
Make this knowledge easy to retrieve and explain under questioning.