Understand transformer internals, adaptation, frameworks, and deployment. Use the mechanisms and diagnostic examples to prepare for deeper technical questions.
1. Transformer internals beyond the attention equation
Know what each component contributes:
| Component | Purpose |
|---|---|
| Multi-head attention | Uses different learned projections to attend to different relationships and representation subspaces |
| Positional information | Gives the model information about token order |
| Feed-forward network | Applies a learned nonlinear transformation independently at each token position |
| Residual connections | Preserve a direct information and gradient path across layers |
| Layer normalization | Normalizes representations within a token’s feature dimensions |
| Causal mask | Prevents a position from attending to future positions |
Standard full self-attention has an attention computation cost of approximately O(n²d), where n is sequence length and d is representation dimension. Longer sequences therefore become expensive. [3]
“If generation is sequential, why can training run in parallel?”
“During training, the complete target sequence is available. We compute predictions for many positions simultaneously while masking future information. During autoregressive inference, the next token depends on previously generated tokens.”
Trap: An attention head is not necessarily an interpretable specialist such as “the grammar head.” Its role is learned.
2. LoRA versus QLoRA—and training memory
Be ready to explain LoRA’s mechanism.
Instead of training an entire weight matrix, LoRA learns a low-rank update:
Here, the original is frozen, and smaller matrices and are trained. A scaling factor is commonly included.
For a weight matrix with dimensions , the two adapter matrices contain approximately:
parameters, where is the chosen rank. [4]
QLoRA: The original approach trains low-rank adapters through a frozen, 4-bit quantized base model, reducing memory requirements further. The trainable adapters are not simply all trained in four-bit precision. [5]
“Why does training need more memory than inference?”
Training additionally needs gradients, optimizer state, and saved activations for backpropagation.
Recognize these memory techniques:
- Gradient accumulation: Accumulate gradients across microbatches before updating parameters.
- Activation checkpointing: Save fewer activations and recompute some during backpropagation.
- Mixed precision: Use appropriate lower-precision computations to reduce memory and potentially improve speed.
Tradeoff: Activation checkpointing saves memory by adding computation.
3. Framework knowledge: LangChain, LangGraph, and tracing
Framework names alone are weak evidence. Explain the capability you used.
| Capability | What to understand |
|---|---|
| LangChain | Model/tool integrations and higher-level agent abstractions |
| LangGraph | Explicit orchestration of stateful workflows and agents |
| LangSmith | Tracing, evaluation, and related development/operations capabilities |
LangGraph can be used without LangChain. Its value includes mixing deterministic steps with model-driven decisions, persistence, and human intervention. [6]
For LangGraph, know:
- State: Shared information describing the current execution.
- Node: A function that performs work and returns state updates.
- Edge: Defines the next step.
- Conditional edge: Chooses a route based on state.
- Reducer: Determines how an update combines with an existing state field.
For example, a reducer can append retrieved results rather than replacing the existing list. Without an explicitly defined reducer, updates normally replace that field’s value. [7]
“Why use graph orchestration rather than a simple function?”
“I use it when explicit branching, persistent execution, inspection, and intervention justify the extra machinery. A short, predictable sequence can remain ordinary application code.”
Trap: Persistence supports recovery; it does not independently guarantee that external side effects occur exactly once.
4. MCP: how it differs from ordinary tool calling
“What is Model Context Protocol?”
“MCP standardizes how AI applications discover and interact with external tools and context.”
Know its main roles and capabilities:
| Term | Meaning |
|---|---|
| Host | The AI application |
| Client | The connector within that application |
| Server | The service exposing capabilities |
| Tools | Callable functions |
| Resources | Readable context or data |
| Prompts | Reusable message/workflow templates |
An MCP server can wrap an existing API. Tool calling is the model’s structured request to use a function; MCP provides a standardized integration protocol around exposed capabilities. [8]
“Does MCP make an integration safe?”
“I still enforce authorization, validate inputs, restrict privileges, and treat returned content according to its trust level.”
5. Deployment details you should explain concretely
“How would you deploy your AI application?”
Walk through a real deployment:
- Package the application and dependencies into a reproducible image.
- Publish the image to a registry.
- Configure runtime identity, networking, secrets, and resource limits.
- Keep durable data in appropriate external storage.
- Configure health checks, scaling, logging, and rollout behavior.
- Verify the deployed application against representative requests.
Image versus container: An image is the packaged template; a container is a running instance of it.
For Kubernetes, distinguish:
| Probe | Question it answers |
|---|---|
| Startup | Has initialization completed? |
| Readiness | Should this instance receive traffic? |
| Liveness | Is this instance unhealthy enough to need restarting? |
A startup probe can protect a slow-starting application while it loads a model. Readiness failure removes the instance from normal service traffic; liveness failure can trigger a restart after the configured threshold. [9]
“The model provider is unavailable. Should liveness fail?”
Usually, restarting an otherwise healthy API process will not repair a remote outage. Design dependency handling and readiness deliberately; avoid creating a restart storm.
Practical trap: Multiple application workers may each load their own model copy. Account for memory ownership before increasing worker count.
6. Retrieval depth: bi-encoders, cross-encoders, and ranking metrics
“Why not use a cross-encoder for every document?”
- A bi-encoder represents the query and document separately. Document embeddings can be precomputed.
- A cross-encoder processes a query–document pair jointly and scores relevance.
A common design uses efficient initial retrieval, then a cross-encoder on a smaller candidate set. Joint scoring across the entire corpus would usually be much more expensive.
Know these metrics beyond Recall@k:
| Metric | What it captures |
|---|---|
| Precision@k | Fraction of top-k results that are relevant |
| MRR | How early the first relevant result appears |
| nDCG | Ranking quality using graded relevance and position discounts |
MRR example: If the first relevant result appears at rank 4, that query’s reciprocal rank is . Average across queries.
Follow-up: “Which metric would you select?”
“It depends on whether one supporting document is sufficient, several documents are required, or relevance has multiple grades.”
7. Questions that require more than similarity search
Consider:
“Which branches exceeded their budget in each of the last three quarters?”
Finding a similar paragraph is insufficient. The task requires exact definitions, records, grouping, and calculations.
“How would you handle this?”
“I identify the required data and computation, retrieve or query authorized records, perform the aggregation deterministically, and use the model to explain the results with provenance.”
Other difficult cases:
| Question type | Useful approach |
|---|---|
| Exact identifier lookup | Preserve identifiers and use suitable exact matching |
| Multiple dependent facts | Decompose the task and validate intermediate evidence |
| Calculations over many records | Database queries or deterministic computation |
| Ambiguous wording | Clarify the missing business definition |
| Time-sensitive question | Apply effective dates and source precedence |
Query rewriting trap: Rewriting can improve retrieval while accidentally changing an identifier, date, or constraint. Preserve the original intent and evaluate the rewrite.
8. Feature stores and training–serving skew
This matters when your interview includes conventional ML engineering.
“What is training–serving skew?”
“The features used during training differ from what the model receives in production, because of inconsistent transformations, timing, defaults, or source behavior.”
Examples:
- Training uses a completed monthly total; production predicts halfway through the month.
- Training and serving apply different category mappings.
- A production feature is stale or unavailable.
- Offline features contain information recorded after the prediction time.
“What does a feature store help with?”
It can support reusable feature definitions, historical training data, online feature access, and lineage. Its design must address freshness and point-in-time correctness.
Point-in-time correctness: Construct each training example using only feature information available at its prediction time.
Trap: Using a feature store does not automatically eliminate leakage or inconsistent business definitions.
9. Know where evaluation metrics stop being useful
“What is perplexity?”
Perplexity is the exponential of average token-level negative log-likelihood. Lower values indicate better prediction of the evaluated token sequence.
It does not directly establish:
- Factual correctness.
- Useful tool selection.
- Successful completion of a business task.
- Reliable abstention.
- Good retrieval.
Compare it carefully: different tokenizers, datasets, and evaluation procedures can make scores incomparable.
Strong answer:
“I use metrics aligned with the capability being measured. Language-model prediction quality and application task success require different evaluations.”
Also recognize selective prediction: allowing a system to abstain can improve accuracy among answered cases, but reduces coverage. Measure both rather than celebrating accuracy alone.
10. Debugging scenarios with deeper reasoning
| Scenario | First hypotheses to investigate |
|---|---|
| Quality drops after changing the embedding model | Query/document embedding compatibility, incomplete reindexing, changed ranking |
| Adding more retrieved chunks worsens answers | Irrelevant evidence, contradictions, truncation, context ordering |
| Average latency is stable but p95 worsens | Slow request classes, queuing, retries, downstream tail latency |
| Offline performance is strong; production is weak | Distribution differences, feature availability, evaluation mismatch |
| GPU memory grows during inference | Retained tensors, caches, concurrency, or unnecessary gradient tracking |
| Aggregate accuracy improves but complaints rise | Important subgroups, failure severity, or task distribution changed |
How to answer without guessing:
“My leading hypothesis is X. I would inspect Y to distinguish it from Z. Then I would run a controlled check and verify the fix against the affected cases.”
That is stronger than presenting an unverified diagnosis as certainty.
Topic coverage is useful only when you can apply it. Demonstrate three things from memory: a project walkthrough, a working coding solution, and an architecture decision defended through follow-up questions.