Skip to content
StudyHubA place to keep learning

Explore

  • Browse Topics
  • Study Packs
  • Library Topics
    • AI engineering interviews

My Workspace

  • My Notes
StudyHub · Model serving and ML fundamentals
Browse topics
AI engineering interviews

Model serving and ML fundamentals

StudyHub8 min readUpdated Oct 3, 2026

Build depth in ML concepts, model serving, coding traps, and quantitative reasoning alongside application architecture.

1. Training objectives: pretraining, SFT, and preference optimization

Question: “How does a language model become an assistant?”

“Pretraining learns language patterns through a predictive objective, commonly next-token prediction. Supervised fine-tuning trains on examples of desired responses. Preference optimization then adjusts behavior using preferred versus less-preferred outputs.”

Know the distinctions:

Stage Main purpose
Pretraining Learn broad representations and predictive capabilities
Supervised fine-tuning, or SFT Learn desired task and instruction-following behavior
RLHF Use human preferences, often through a reward model and reinforcement learning
DPO Optimize directly from preference pairs without the conventional RLHF training loop
Scroll across to read all columns.

Follow-up: “Does alignment make the model factually reliable?”

No. Desired behavior and factual reliability are related but different. An aligned model can still generate unsupported statements.

2. Decoder-only models versus encoder models

Question: “Why use different models for generation and embeddings?”

“Generation and retrieval have different objectives. A causal language model predicts successive tokens, while an embedding model is trained to produce representations useful for similarity or retrieval.”

  • Encoder models: Commonly use bidirectional context; useful for classification and representation tasks.
  • Decoder-only models: Use causal attention; commonly used for autoregressive generation.
  • Encoder–decoder models: Encode an input and generate an output; useful for sequence-to-sequence tasks.

Architecture alone does not guarantee suitability. Training objective and task evaluation matter.

Trap: Averaging an arbitrary LLM’s hidden states does not automatically produce a good retrieval embedding.

3. Contrastive learning and hard negatives

Question: “How are retrieval embedding models trained?”

A common approach brings relevant query–document pairs closer and pushes irrelevant pairs apart.

“The training objective distinguishes positive examples from negatives. Hard negatives are particularly useful because they resemble relevant documents but do not satisfy the query.”

Example:

  • Query: “What is the cancellation policy for premium accounts?”
  • Positive: Premium-account cancellation policy.
  • Hard negative: Standard-account cancellation policy.

Why this matters: Similar vocabulary does not necessarily imply the correct answer.

Trap: Incorrectly labeling a genuinely relevant document as a negative can harm training.

4. Model serving: prefill, decoding, and KV cache

Question: “Why do long prompts and long responses affect latency differently?”

LLM inference has two important phases:

  • Prefill: Process the input tokens and construct attention state.
  • Decode: Generate output tokens sequentially.

A long input increases prefill work. A long response requires more decoding steps.

What is a KV cache?

“It stores attention keys and values from previously processed tokens so the model can reuse them during generation.”

It avoids recomputing those representations, but consumes memory. Cache size grows with sequence length, concurrency, and model architecture.

Metric Meaning
Time to first token Delay before the first output token arrives
Time per output token Speed of subsequent generation
End-to-end latency Time until the response finishes
Throughput Requests or tokens processed per unit time
Scroll across to read all columns.

Follow-up: “Does streaming make inference faster?”

Streaming improves perceived responsiveness by showing output earlier. It does not inherently reduce total computation or completion time.

5. GPU memory and batching

Question: “Why might a model fit on a GPU but fail under concurrent traffic?”

“Weights are only part of memory use. Concurrent inference also requires KV caches, activations, temporary buffers, and runtime overhead.”

Approximate weight memory only:

Weight memory≈parameter count×bytes per parameter\text{Weight memory} \approx \text{parameter count}\times\text{bytes per parameter}

A 7-billion-parameter model at two bytes per parameter needs roughly 14 GB for weights alone, before inference overhead.

Batching tradeoff:

  • Larger batches can improve throughput and hardware utilization.
  • Waiting to form batches can increase latency.
  • More active sequences increase memory demand.

Continuous batching lets a serving system admit and remove sequences as requests progress.

Trap: Quantizing weights does not automatically shrink every other memory component.

6. Classifier thresholds and calibration

Question: “Your fraud model outputs 0.8. Does that mean an 80% chance of fraud?”

Only if the probabilities are appropriately calibrated for the relevant population.

“Calibration checks whether predicted probabilities match observed frequencies. Discrimination checks whether the model ranks positive cases above negative cases.”

These are different properties.

How do you choose a threshold?

Use validation data, operational capacity, and the cost of mistakes.

For fraud detection:

  • False positives inconvenience legitimate customers.
  • False negatives allow fraudulent activity.
  • Review capacity may limit how many alerts can be handled.

Trap: A threshold of 0.5 is not automatically optimal. A well-ranked model can still need calibration and a different threshold.

7. Cross-validation for time and grouped data

Question: “When is random train/test splitting wrong?”

When it breaks the structure of how the model will be used.

Situation Appropriate consideration
Forecasting future events Train on earlier periods and evaluate on later periods
Multiple records per customer Split by customer if evaluating generalization to unseen customers
Documents derived from the same source Keep related documents together to avoid leakage
Repeated experiments or measurements Respect the grouping that creates dependence
Scroll across to read all columns.

“I choose the split to reproduce the deployment scenario and prevent information crossing from evaluation into training.”

Follow-up: “Can time-based splitting still leak?”

Yes. Features may contain information recorded after the prediction time. Check feature availability, not just row timestamps.

8. Class imbalance and misleading metrics

Suppose only 1% of transactions are fraudulent. Predicting “not fraud” for every transaction produces 99% accuracy, while detecting no fraud.

Know these formulas:

Precision=TPTP+FP,Recall=TPTP+FN\text{Precision}=\frac{TP}{TP+FP}, \qquad \text{Recall}=\frac{TP}{TP+FN}
F1=2Precision×RecallPrecision+RecallF_1=2\frac{\text{Precision}\times\text{Recall}} {\text{Precision}+\text{Recall}}

Question: “PR-AUC or ROC-AUC?”

“For rare positive events, the precision–recall curve is often more informative about positive prediction quality. I still evaluate performance at operating thresholds and connect it to business costs.”

Trap: F1 gives precision and recall equal emphasis; your business may not.

9. Python traps interviewers can turn into coding questions

Late binding in closures

Python
functions = [lambda: i for i in range(3)]
print([f() for f in functions])
# [2, 2, 2]

The functions look up i when called.

Capture its value at creation:

Python
functions = [lambda i=i: i for i in range(3)]
# Calling them produces [0, 1, 2]

Shallow versus deep copying

A shallow copy creates a new outer container but retains references to nested objects.

Python
original = [[1], [2]]
copied = original.copy()
copied[0].append(9)

print(original)
# [[1, 9], [2]]

Generator exhaustion

Python
values = (x * x for x in range(3))

print(list(values))  # [0, 1, 4]
print(list(values))  # []

A generator is consumed as you iterate.

Practise explaining: What happens, why it happens, and how you would change the code.

10. SQL traps: NULLs, joins, and counting

Question: “Why can a LEFT JOIN unexpectedly drop rows?”

SQL
SELECT c.customer_id, o.order_id
FROM customers c
LEFT JOIN orders o
    ON c.customer_id = o.customer_id
WHERE o.status = 'completed';

The WHERE condition excludes unmatched rows because their order fields are NULL.

To retain all customers while matching completed orders:

SQL
SELECT c.customer_id, o.order_id
FROM customers c
LEFT JOIN orders o
    ON c.customer_id = o.customer_id
   AND o.status = 'completed';

Also remember:

  • COUNT(*) counts rows.
  • COUNT(column) counts non-NULL values.
  • NULL = NULL does not evaluate to true; use IS NULL.
  • One-to-many joins can multiply rows and inflate aggregates.
  • NOT IN can behave unexpectedly if the subquery contains NULL; consider NOT EXISTS for anti-joins.

Senior-level habit: State the intended grain—one row per customer, order, or transaction—before writing the query.

11. Quick capacity and cost calculations

Question: “How many concurrent requests should we expect?”

For a stable system, Little’s Law gives:

L=λWL=\lambda W

If arrival rate is 10 requests/second and mean time in the system is 3 seconds, mean requests in the system is approximately 30.

That is an average, not a safe capacity limit. Bursts and latency variation require headroom.

Question: “How do you estimate model API cost?”

Cost=input tokens106Pinput+output tokens106Poutput\text{Cost} = \frac{\text{input tokens}}{10^6}P_{\text{input}} + \frac{\text{output tokens}}{10^6}P_{\text{output}}

Use the applicable pricing units and add embedding, retrieval, tool, and infrastructure costs where relevant.

Strong follow-up: “I would track cost per successfully completed task, since retries and failed answers can make cost per request misleading.”

Final challenge: answer these without notes

  • Why can a model have strong ROC-AUC but poor business performance?
  • Why does increasing concurrent LLM requests increase memory pressure?
  • Why is a random split unsafe for some customer datasets?
  • Why can a syntactically correct SQL query double-count revenue?
  • Why does a shallow copy allow changes to affect the original?
  • Why might good semantic similarity still retrieve the wrong policy?

For each, aim for a mechanism, a concrete example, and a practical fix.

Take a moment to recall

A short quiz is ready when you want to check your understanding.

Continue exploringBackend engineering and senior ownership
BrowseAccount