Study sampling, agent execution patterns, practical Python concurrency, and advanced ML. Practise explaining the mechanisms through concrete examples.
1. Sampling: temperature, top-k, top-p, and greedy decoding
Distinguish temperature from the other sampling controls.
| Method | What it does |
|---|---|
| Greedy decoding | Selects the highest-probability token at each step |
| Temperature | Changes how concentrated the probability distribution is |
| Top-k sampling | Restricts sampling to the k highest-probability tokens |
| Top-p sampling | Restricts sampling to a set whose cumulative probability reaches a chosen threshold |
“Why use top-p instead of top-k?”
“Top-p adapts the candidate set to the probability distribution. A confident prediction can use a small set; an uncertain prediction can retain more candidates.”
“Does greedy decoding produce the best answer?”
It chooses the locally most probable token. That does not guarantee the best complete sequence or the most factually correct answer.
“Why can outputs differ even with low temperature?”
Sampling settings are only part of reproducibility. Model versions, execution details, and changing input or retrieved context also matter.
2. Tokenization and special tokens
“Why does tokenization matter to an AI engineer?”
“It affects context usage, cost, truncation, and how different languages or identifiers are represented.”
Know these details:
- A token can represent a word, word fragment, punctuation, or another text unit.
- Different tokenizers split the same text differently.
- Character count is an unreliable universal estimate of token count.
- Special tokens can mark boundaries, roles, or sequence termination.
- Truncation can remove a critical instruction, source passage, or conversational detail.
“Why might the same application cost more in another language?”
The tokenizer may require more tokens for equivalent content. Measure actual token usage rather than assuming a fixed words-to-tokens ratio.
Practical answer: Budget space for instructions, conversation, retrieved evidence, and generated output.
3. Agent patterns: recognize the architecture
| Pattern | How it works | Main concern |
|---|---|---|
| ReAct-style loop | Select an action, observe its result, then choose the next action | Repeated steps, loops, and growing cost |
| Plan-and-execute | Create a plan, then execute its steps | Plans become stale when observations change |
| Router | Select a specialized model, tool, or workflow | Incorrect routing |
| Critic/reviewer | Review an output and request improvements | Reviewer may share the original model’s mistakes |
| Supervisor with specialists | Coordinate specialized workers | Coordination overhead and compounded errors |
“Does adding a reviewer agent guarantee better quality?”
“No. I measure whether review improves outcomes enough to justify its latency and cost. Independent checks, deterministic validation, or human review may be more reliable for particular errors.”
“How do you handle a plan becoming invalid?”
Validate intermediate results and allow bounded replanning. Record which actions have already completed so replanning does not repeat their effects.
4. Python concurrency: be able to write it
“Call an async service for several items, with at most five calls active.”
import asyncio
async def process_items(items, call, concurrency=5):
if concurrency < 1:
raise ValueError("concurrency must be positive")
semaphore = asyncio.Semaphore(concurrency)
async def process_one(item):
async with semaphore:
return await asyncio.wait_for(
call(item),
timeout=10,
)
return await asyncio.gather(
*(process_one(item) for item in items)
)
Explain the mechanism:
“The semaphore bounds active calls. Each call has a timeout. Gather returns results in input order.”
Then explain the limitations:
- This creates a coroutine/task for every item. For very large inputs, use a bounded queue and a fixed worker pool.
- Define an error policy: fail the batch, retry selected failures, or return individual outcomes.
- By default, an ordinary exception propagated by
gatherdoes not automatically cancel every other submitted task.
“Semaphore versus rate limiter?”
A semaphore limits concurrent work. A rate limiter controls work admitted over time. Five concurrent requests can still produce hundreds of requests per second if they finish quickly.
5. Rate limiting and retry budgets
Go beyond “retry with exponential backoff.”
“What happens when every layer retries?”
If three layers each make up to three attempts, one request can trigger up to 27 downstream attempts.
“I assign retry responsibility deliberately, retry only suitable failures, and enforce an overall deadline and retry budget.”
Also distinguish:
- Request limits: Number of requests admitted.
- Token limits: Model workload admitted.
- Concurrency limits: Number of operations active.
- Tenant quotas: Fair allocation between customers or teams.
“Why might requests per second be an inadequate LLM limit?”
A short classification request and a long document analysis request consume very different resources. Consider tokens, execution time, and active sequences.
6. Distillation and model compression
“Distillation versus quantization?”
“Distillation trains a student model to reproduce useful behavior from a teacher. Quantization reduces numerical precision in a model’s representation or computation.”
| Technique | Primary change |
|---|---|
| Distillation | Trains a student using teacher outputs or signals |
| Quantization | Uses lower numerical precision |
| Pruning | Removes selected weights or structures |
“When would you distill?”
When a narrower task needs lower latency or cost, and you can create representative training data and evaluate the student independently.
Important follow-up: Teacher-generated data can contain errors. Filter and validate it, and evaluate on an independent set rather than merely measuring agreement with the teacher.
A student can imitate a teacher’s mistakes exceptionally well.
7. Explainability: SHAP and causal claims
“What does SHAP tell you?”
“SHAP attributes a model’s prediction to input features relative to a defined reference. It helps explain the model’s behavior.”
Distinguish:
- Local explanation: Why the model produced one prediction.
- Global analysis: Patterns across many predictions.
- Causal explanation: What would happen if you intervened in the real world.
Feature attribution does not establish causation.
“A feature has high importance. Should we change it to improve the outcome?”
“Not automatically. It may be a proxy, correlate with other features, or be outside our control. A causal or experimental analysis is needed to support an intervention.”
Correlated features and the reference dataset can affect interpretation. Also ensure that explanations do not reveal information the recipient should not access.
8. Bias, subgroup performance, and fairness
“Your overall accuracy is good. How could the system still be problematic?”
Performance may differ substantially across user groups, languages, document types, or operating conditions.
Evaluate:
- False-positive and false-negative rates by relevant subgroup.
- Calibration where probabilities inform decisions.
- Sample sizes and uncertainty.
- Data representation and label quality.
- Consequences of different errors.
“Does removing a sensitive attribute remove bias?”
No. Other variables can act as proxies, and historical labels can encode unequal treatment.
Strong answer:
“I define the relevant groups and decision consequences, inspect subgroup performance, investigate causes, and evaluate mitigations against explicitly chosen objectives.”
Avoid claiming that one fairness metric proves the entire system is fair.
9. Unsupervised learning and anomaly detection
These are useful if the interviewer draws on your data science background.
| Method | Core idea | Important limitation |
|---|---|---|
| K-means | Groups points around centroids | Sensitive to scaling, initialization, outliers, and cluster shape |
| PCA | Finds orthogonal directions of maximum variance | High variance need not mean high predictive value |
| Isolation Forest | Detects points that are relatively easy to isolate through random partitions | Anomalous does not necessarily mean fraudulent or incorrect |
“How do you evaluate clustering without labels?”
Use internal measures, stability, and domain usefulness. A strong mathematical score does not guarantee actionable groups.
“Why not automatically reject every anomaly?”
An anomaly may be a legitimate rare event. The score needs a threshold and a handling process appropriate to the consequences.
10. Missing values and resampling
“How would you handle missing values?”
“I first investigate why values are missing and whether production will exhibit the same pattern. Then I choose an appropriate treatment, fit it on training data, and validate it.”
Possible approaches include simple imputation, model-based imputation, explicit missing categories, or missingness indicators.
“Can you oversample before splitting the data?”
That can leak duplicate or synthetic information into evaluation. Split first; apply resampling only within training partitions or training folds.
“Does oversampling fix the whole imbalance problem?”
No. You still need appropriate evaluation, threshold selection, and attention to probability calibration.
One final exercise that combines these topics
Consider this interview scenario:
“Your AI service handles long and short requests. Under load, short requests become slow, retries multiply, and one customer consumes most of the capacity. What would you change?”
A strong answer:
“I would inspect queueing and workload distribution, enforce bounded concurrency, and separate or schedule request classes where justified. I would apply tenant quotas and workload-aware limits, consolidate retries under an overall deadline, and measure latency by request class alongside throughput and successful completion.”
Practise execution: explain mechanisms, write working code, and connect decisions to evidence. Reading provides coverage; answering follow-ups from memory demonstrates depth.