Skip to content
StudyHubA place to keep learning

Explore

  • Browse Topics
  • Study Packs
  • Library Topics
    • AI engineering interviews

My Workspace

  • My Notes
StudyHub · Backend engineering and senior ownership
Browse topics
AI engineering interviews

Backend engineering and senior ownership

StudyHub10 min readUpdated Oct 3, 2026

Develop core ML depth, practical backend engineering, coding patterns, and senior ownership. These topics connect model decisions to dependable production systems.

Read these once, then spend your remaining time answering aloud and writing code without looking.

1. Choosing and explaining traditional ML models

Be ready to explain why you would use a simpler model before an LLM.

Model Why choose it? Main limitation
Linear regression Continuous prediction with a simple, interpretable baseline Linear relationship assumptions; sensitivity to outliers
Logistic regression Classification with interpretable coefficients Linear decision boundary in the supplied feature space
Decision tree Nonlinear rules and feature interactions Can overfit and change substantially with small data changes
Random forest Robust tabular baseline; reduces variance through averaging Less transparent and potentially expensive
Gradient boosting Strong tabular performance through sequential improvement Requires careful tuning and validation
Scroll across to read all columns.

“Random forest versus boosting?”

“Random forests train randomized trees and aggregate their predictions, primarily reducing variance. Boosting builds models sequentially, with later models improving the ensemble’s objective. I compare them using deployment-relevant validation, latency, and interpretability requirements.”

“When is feature scaling necessary?”

Scaling matters for distance-based methods, many gradient-based methods, and regularized linear models. Standard decision-tree splits generally do not require scaling.

“Can logistic regression capture nonlinear relationships?”

Yes, through suitable nonlinear features and interactions. It remains linear in its supplied feature representation.

2. Loss functions, optimization, and regularization

“Loss function versus evaluation metric?”

“The loss guides optimization during training. The evaluation metric measures performance against the task or business objective. They need not be identical.”

Examples:

  • Regression: mean squared error penalizes large errors strongly; mean absolute error is less sensitive to extreme residuals.
  • Classification: cross-entropy penalizes assigning low probability to the correct class.
  • Retrieval: contrastive objectives help distinguish relevant pairs from negatives.

“What happens if the learning rate is too high or too low?”

Too high can cause unstable updates or failure to converge. Too low can make training slow or stall progress within the available budget.

“L1 versus L2 regularization?”

  • L1 penalizes absolute coefficient values and can produce sparse coefficients.
  • L2 penalizes squared coefficient values and tends to shrink them smoothly.

“What is backpropagation?”

“Backpropagation applies the chain rule to compute loss gradients with respect to model parameters. The optimizer uses those gradients to update the parameters.”

Also recognize:

  • Early stopping: Stop when validation performance no longer improves.
  • Dropout: Randomly suppress activations during training.
  • Gradient clipping: Bound gradients to help control unstable updates.
  • Training versus evaluation mode: Layers such as dropout and batch normalization behave differently.

3. Fine-tuning: explain the actual engineering process

Be ready to explain both when to fine-tune and how to implement it.

“Walk me through a fine-tuning project.”

“I establish a baseline and measurable objective, curate representative training examples, remove duplicates and leakage, and create independent validation and test sets. I choose an adaptation method based on compute and task needs, train while monitoring validation performance, then compare against the baseline on quality, regressions, latency, and cost.”

Know the failure modes:

Failure What to investigate
Training improves; validation worsens Overfitting, poor split, insufficient diversity
Model reproduces inconsistent behavior Conflicting or low-quality training examples
Target task improves; other tasks worsen Narrow adaptation and capability regression
Evaluation looks unusually strong Duplicate examples or evaluation contamination
Scroll across to read all columns.

“How much training data do you need?”

There is no universal count. It depends on task complexity, example quality, diversity, and the starting model. Use learning curves and held-out evaluation.

Important distinction: Fine-tuning changes model parameters. In-context learning uses examples in the prompt without updating parameters.

4. Backend engineering: APIs, authentication, and authorization

“Design an API for an AI application.”

Explain the request contract, validation, authentication, authorization, execution, response contract, and error handling.

Situation Typical HTTP response
Successful synchronous request 200
Resource created 201
Long-running task accepted 202, with a job identifier
Missing or invalid authentication 401
Authenticated but insufficient permission 403
Rate limit exceeded 429
Invalid input An appropriate 4xx, according to the API contract
Scroll across to read all columns.

“Authentication versus authorization?”

Authentication establishes identity. Authorization determines what that identity may access or do.

“Is decoding a JWT enough?”

No. Decoding reads its contents. Verification must check its signature and relevant claims, such as expiry, issuer, and audience.

“Why might an async API perform poorly?”

“An async handler can still block its event loop if it performs synchronous network calls or CPU-heavy work. I inspect dependencies and move suitable work to asynchronous clients or separate workers.”

For long-running ingestion, a job API is often more suitable than holding an HTTP request open.

5. Database depth: indexes, transactions, and race conditions

“Why is this SQL query slow?”

“I inspect the execution plan and actual workload. I look for large scans, expensive joins or sorts, poor cardinality estimates, and missing or unsuitable indexes. I verify the improvement against representative data.”

Indexes speed some reads but consume storage and add write overhead. Composite index usefulness depends on query predicates, ordering, and the database.

“What does ACID mean?”

  • Atomicity: A transaction completes as a unit or is rolled back.
  • Consistency: Transactions preserve defined integrity rules.
  • Isolation: Concurrent transactions interact according to the isolation level.
  • Durability: Committed changes survive the failures covered by the database’s guarantees.

“Two workers both see a job as unprocessed. What prevents duplicate work?”

A read followed by an update can race. Use an atomic claim operation, suitable locking, or a database constraint and transaction.

“I put concurrency protection at the shared source of truth. An in-memory lock in one API process cannot coordinate separate replicas.”

6. Coding patterns worth recognizing immediately

Do not memorize isolated solutions. Identify the pattern.

Problem Useful pattern Typical complexity
Frequency counting or two-sum lookup Hash map Expected O(n)
Longest substring without repetition Sliding window O(n)
Top k items Heap O(n log k)
Shortest path in an unweighted graph BFS O(V + E)
Dependency ordering Topological sort O(V + E)
Repeated calculations with overlapping subproblems Dynamic programming Depends on states and transitions
Scroll across to read all columns.

Practise this now: longest substring without repeating characters.

Python
def longest_unique_length(text):
    last_seen = {}
    left = 0
    best = 0

    for right, character in enumerate(text):
        previous = last_seen.get(character)

        if previous is not None and previous >= left:
            left = previous + 1

        last_seen[character] = right
        best = max(best, right - left + 1)

    return best

Explain:

“I maintain a window containing unique characters. When a character repeats inside the window, I move the left boundary past its previous occurrence.”

Expected O(n) time; space depends on the number of distinct characters.

Check mentally: "" → 0, "aaaa" → 1, "abba" → 2.

During coding: clarify inputs, give an approach, implement, check edge cases, then state complexity.

7. Distributed systems: the timeout ambiguity

“An external tool times out. Can you retry?”

A timeout means the caller did not receive a timely result. It does not prove the operation failed. The remote system may have completed it.

“For reads, retries are usually straightforward. For writes, I use an idempotency key or check operation status before repeating an action.”

“How do you keep a database update and an emitted event consistent?”

A transaction cannot automatically cover an unrelated database and messaging service.

One approach is the transactional outbox:

  1. Update business data and insert an outbox record in the same database transaction.
  2. A publisher sends the outbox event.
  3. Consumers tolerate duplicate delivery.

“What is a circuit breaker?”

It temporarily stops calls to a repeatedly failing dependency, allows limited recovery probes, and resumes normal calls when recovery is established.

Distinguish it from a retry: retries make another attempt; a circuit breaker limits attempts during sustained failure.

8. Statistics and proving that an improvement is real

“Model B scores 84%; model A scores 82%. Is B better?”

“I need the sample size, evaluation design, uncertainty, and error distribution. I compare models on the same representative cases and examine whether the improvement is practically meaningful.”

Remember:

  • A small evaluation set can produce unstable estimates.
  • An aggregate gain can hide deterioration for important user groups.
  • Repeatedly tuning against the test set makes it less independent.
  • Statistical significance does not establish business value.
  • Offline improvement does not guarantee better user outcomes.

For a controlled online experiment, define the assignment unit, success metric, guardrails, and analysis plan before examining results.

“How would you prove productivity improvement?”

Measure comparable tasks, completion quality, total time including corrections, and adoption. A faster initial answer can still increase overall work if users must repair it.

9. Handling senior-level scenario questions

These assess judgment more than terminology.

“The business wants to launch, but quality is below target.”

“I identify which failure categories matter, quantify their impact, and propose a scoped release where acceptance criteria are met. I define monitoring and rollback conditions and make unresolved risks explicit.”

“A stakeholder asks for an agent, but a workflow is sufficient.”

“I translate the request into required capabilities, compare approaches through a small benchmark, and recommend the one that meets the requirements with acceptable complexity and operating cost.”

“A production change causes wrong answers.”

“I contain the impact, roll back or disable the affected path, preserve relevant evidence, and communicate the scope. Then I investigate the cause and add a regression check addressing the failure.”

“You inherited a poorly documented AI system.”

“I establish the data flow, dependencies, access controls, deployment process, and current performance. I reproduce key behavior and prioritize changes based on observed failure and business impact.”

Use “I” for your actions and “we” for shared decisions and outcomes.

10. Questions to ask the interviewer

Choose two or three:

  • “What would successful delivery look like in the first three months?”
  • “Is the immediate priority experimentation, production delivery, or improving an existing platform?”
  • “How are AI quality criteria defined and approved before deployment?”
  • “Which responsibilities belong to this role versus platform, data, and security teams?”
  • “What is the hardest engineering problem the team is currently facing?”

These help you understand the role and discuss where your experience fits.

Your final preparation task

Close the notes and complete this 12-minute rehearsal:

Time Task
2 minutes Introduce yourself and describe your strongest project
3 minutes Defend one architectural decision, including its alternative and evidence
3 minutes Write the sliding-window solution from memory
2 minutes Explain how you would investigate a production failure
2 minutes Describe a disagreement, your actions, and the outcome
Scroll across to read all columns.

Carry these five habits into the interview:

  • Clarify the requirement before choosing technology.
  • Explain mechanisms, not just framework names.
  • Connect choices to evidence and measurable outcomes.
  • Cover failure behavior and operational constraints.
  • Be precise about what you personally implemented.

Make this knowledge easy to retrieve and explain under questioning.

Take a moment to recall

A short quiz is ready when you want to check your understanding.

Continue exploringTransformers, frameworks, and deployment
BrowseAccount