Tailr
← All posts

Top 30 AI Engineer Interview Questions and Answers

· updated

An AI engineer interview loop shown as a browser window with cards for the five question areas: fundamentals, retrieval, agents, evaluation and production

AI engineer interviews are new enough that the format varies a lot between companies, but the questions underneath have settled. This is what comes up, grouped by round, with answers you can adapt. Where a question has a “gotcha”, the answer says what it is.

The short answer: AI engineer interviews test five things:

  • How language models actually behave: tokens, context, temperature, structured output, tool calling.
  • Retrieval-augmented generation: chunking, embeddings, hybrid search, reranking.
  • Agents and tool use: loops, state, limits, failure handling.
  • Evaluation: building eval sets, metrics, model graders, regressions.
  • Production concerns: latency, cost, observability, prompt injection, provider outages.

Most loops also include a coding round, a system design round for an AI feature, and questions about a project of yours that went wrong. Evaluation is the topic that decides most offers.

What the interview loop looks like

A typical 2026 loop for an AI engineer role:

  1. Recruiter screen, 30 minutes: your background, what you’ve built, salary range.
  2. Technical phone screen, 45–60 minutes: fundamentals conversation plus a short coding exercise.
  3. Take-home or practical, 2–4 hours: build a small RAG system or agent, or debug one that’s broken. Some companies replace this with a live pairing session.
  4. On-site or virtual loop, 3–4 hours: a coding round, an AI system design round, a deep dive on your project, and a behavioural round.

The questions below are grouped roughly by where they appear. If you’re new to the field, read what does an AI engineer do first so the role is clear; the rest of this post assumes it is.

Fundamentals: how models behave

1. What is a token, and why does it matter? A token is the unit a language model reads and writes: roughly three-quarters of an English word, though code and other languages tokenise differently. It matters because cost, latency and context limits are all measured in tokens. A practical answer mentions that you estimate token counts before sending long documents, and that output tokens usually cost more and take longer than input tokens.

2. What is a context window, and what happens when you exceed it? The context window is the maximum number of tokens the model can consider at once: the system prompt, conversation history, retrieved documents and the output all share it. Exceed it and the request fails or the oldest content is truncated, depending on the API. Good answer: bigger context isn’t free. Models attend less reliably to content in the middle of very long contexts, cost scales with length, and the engineering skill is deciding what to include, not including everything.

3. Explain temperature. Temperature controls how much randomness is in the sampling of each token. Low temperature makes the model pick the most likely tokens, giving consistent, deterministic-ish output; high temperature gives more varied output. Use low values for extraction, classification and anything with a schema; higher values for brainstorming or creative generation. Gotcha: temperature 0 doesn’t guarantee identical outputs across calls.

4. What is structured output, and how do you make it reliable? Getting the model to return data in a defined shape, usually JSON matching a schema, so a program can consume it. Reliable approaches: use the provider’s native structured output or tool-calling features rather than asking nicely in the prompt, validate the result against the schema in code, and retry with the validation error included when it fails. Mention that you keep schemas small and flat, because deeply nested optional fields are where models go wrong.

5. What is tool calling (function calling)? You describe functions to the model with a name, description and parameter schema. The model can respond with a request to call one, with arguments. Your code runs the function and sends the result back, and the model continues. Good answer: the description is a prompt, so write it carefully; the model chooses tools based on it. Also mention that you validate arguments before executing anything, because the model can produce plausible but wrong ones.

6. What is an embedding? A vector of numbers that represents the meaning of a piece of text, produced by an embedding model, so that texts with similar meaning have vectors that are close together. Used for semantic search, clustering and retrieval. Strong answer: mention that embeddings from different models aren’t comparable, that you have to re-embed everything if you change models, and that embedding similarity captures topic better than it captures exact facts, which is why hybrid search exists.

7. What is hallucination, and how do you reduce it? The model producing confident output that isn’t grounded in its input or in reality. You reduce it with retrieval (give it the facts), with instructions to cite sources and say “I don’t know”, with structured output that constrains what it can say, with a verification step, and with evals that measure it. You never eliminate it. The honest answer includes that last sentence.

8. What is prompt injection? An attack where content the model reads (a web page, an email, a document, a tool result) contains instructions that hijack its behaviour. Defences: treat all retrieved and user-supplied content as untrusted data, separate instructions from data clearly in the prompt, limit what tools can do without confirmation, filter or sandbox risky actions, and test with adversarial inputs. Gotcha: there’s no complete fix, so the design has to assume the model can be tricked.

Retrieval-augmented generation

New to the topic? What is RAG explains the basics before you tackle these questions.

9. Walk me through a RAG pipeline. Ingest: load documents, split into chunks, embed each chunk, store in a vector index with metadata. Query: embed the question, retrieve the top-k similar chunks (often combined with keyword search), optionally rerank, assemble the best ones into the prompt with instructions to answer from them and cite. Generate, then post-process and log. Strong candidates mention evaluation as part of the pipeline, not an afterthought.

10. How do you choose chunk size? Trade-off: small chunks are precise but lose context; large chunks keep context but dilute the signal and use more tokens. Start around 300–800 tokens with some overlap, split on natural boundaries (headings, paragraphs) rather than fixed character counts, and then measure. The right answer is “I’d test three sizes against my eval set”, not a magic number.

11. What is hybrid search, and why use it? Combining vector (semantic) search with keyword search such as BM25, then merging the results. Vector search finds meaning; keyword search finds exact terms, product codes, names and error messages that embeddings handle poorly. Most production RAG systems use both, because the failure cases of each are different.

12. What is reranking? A second-stage model that takes the top 20–50 retrieved chunks and re-scores them for relevance to the specific query, more accurately than the first-stage embedding similarity. It’s slower per document, which is why it runs on a shortlist. Typically one of the biggest single quality improvements in a RAG system; mention the latency cost.

13. How would you debug a RAG system that gives wrong answers? Separate retrieval failures from generation failures. Log the retrieved chunks for each bad answer and check: was the right information retrieved at all? If not, it’s a retrieval problem (chunking, embedding model, missing hybrid search, bad metadata filters). If it was retrieved and the model still answered wrongly, it’s a generation problem (prompt, context assembly, model choice). Most failures are retrieval failures.

14. RAG or fine-tuning? RAG when knowledge changes often, must be cited, or is too large to train in. Fine-tuning for consistent style or format, a narrow task at high volume, or to get a small fast model to do what a big one does. Often both: fine-tune for behaviour, retrieve for knowledge. The gotcha is thinking fine-tuning teaches facts reliably; it doesn’t.

Agents and tool use

15. What is an agent, and when would you not use one? A system where the model decides which tools to call and in what order, in a loop, until a task is done. Don’t use one when the steps are known in advance: a fixed pipeline is cheaper, faster, more predictable and easier to test. Use an agent when the path genuinely depends on intermediate results. Interviewers like candidates who reach for the simpler option first.

16. How do you stop an agent from looping forever or doing damage? Hard limits on iterations, tokens and cost per task. Tool permissions: read-only by default, confirmation for anything destructive or external. Timeouts. Clear stop conditions. Logging of every step so you can replay a run. Idempotent tools where possible, so a retry doesn’t double-charge a customer.

17. How do you design a good tool for an agent? Narrow purpose, clear name, a description written as if for a junior colleague, a small typed parameter schema, and an output that’s concise and useful to the model (not a raw 50 KB JSON dump). Return errors as informative text the model can act on. Fewer, better tools beat many overlapping ones.

18. How do you evaluate an agent? Harder than evaluating a single response, because there are many valid paths. Evaluate the end state (did the task get done?), the trajectory (did it take a reasonable path, within budget?), and safety (did it do anything it shouldn’t have?). Build a set of tasks with checkable outcomes, run them repeatedly because results vary, and report success rate with a confidence interval, not a single number.

Evaluation

19. How would you build an evaluation set from scratch? Start with real inputs: logs, support tickets, actual user questions, not invented ones. Aim for 50 to 200 to begin with. For each, define the expected output or a grading rubric. Cover the distribution: common cases, edge cases, known failure modes, adversarial inputs. Keep a held-out set you never tune against. Version it, and grow it whenever you find a new failure in production.

20. What metrics would you use? Depends on the task. Classification and extraction: accuracy, precision, recall, F1 against labels. Retrieval: recall@k, MRR. Generation: faithfulness to sources, answer relevance, and task-specific rubric scores, usually graded by a model. Always also track cost per request and latency percentiles, because a 2-point quality gain at 3× the cost is a business decision, not an engineering one.

21. What is “LLM as a judge”, and what are its problems? Using a strong model to grade another model’s outputs against a rubric, so you can evaluate thousands of open-ended responses without human labelling. Problems: judges have biases (they prefer longer answers, they prefer their own style, they can be inconsistent), so you calibrate them against a human-labelled sample, use precise rubrics, and never let the judge be the same model you’re evaluating without checking for self-preference.

22. A model provider releases a new version. What do you do? Run the full eval suite against it before switching anything. Compare quality, cost and latency. Look specifically at the cases that changed, not just the aggregate score, because a new model can fix ten things and break five others. Roll out behind a flag, watch production metrics, keep the rollback ready. The gotcha: “newer” isn’t “better for your task” until your eval says so.

System design

23. Design a support assistant that answers from a company’s help centre. Walk it through: ingestion of articles with metadata (product, version, date); chunking on headings; hybrid search plus reranking; a prompt that answers only from sources, cites them and escalates when unsure; a structured output with answer, sources and confidence; an eval set from real tickets; streaming for perceived speed; caching for common questions; logging of retrieved chunks per request; a filter so account data doesn’t leave; and a human handoff path. Then talk about what you’d measure after launch.

24. Design an agent that processes incoming invoices. Emphasise the failure cost: money moves. Extraction with a strict schema and validation; confidence thresholds that route low-confidence cases to a human; idempotency so a retry never pays twice; an audit log of every decision; evaluation on a labelled set of real invoices including the weird ones; and a gradual rollout starting with read-only suggestions before any automatic action.

25. How would you cut the cost of an AI feature by half without hurting quality? Measure first. Then, in order of usual impact: cache repeated or near-duplicate requests; shorten prompts and retrieved context; use a smaller model for easy cases with a router that escalates hard ones; batch offline work; reduce output length with tighter schemas; use prompt caching features from the provider. Re-run evals after each change so “without hurting quality” is a measured claim.

Production and operations

26. What do you log for every model call, and why? The full prompt (or a reference to it), the model and version, the parameters, the retrieved context, the output, token counts, cost, latency, and a request ID that links to the user session. Because when an answer is wrong, you need to replay exactly what the model saw. Also mention privacy: redact or hash sensitive fields, and know your provider’s retention policy.

27. How do you handle a provider outage? Timeouts and retries with backoff for transient errors. A fallback to a second provider or a smaller model for degraded service, tested regularly so it actually works. Graceful degradation in the product: a clear message, a cached answer, or a non-AI path. Circuit breakers so a slow provider doesn’t take down your whole service.

28. How do you keep sensitive data out of the model? Classify data before it reaches the prompt. Redact or tokenise personal and financial data. Use provider tiers with no training on your data and appropriate retention. For the most sensitive cases, run open-weight models in your own environment. And log what you sent, so you can prove it later.

Behavioural and project questions

29. Tell me about an AI feature you built that didn’t work at first. This is the most important question in the loop. Structure: what you built, how you measured it, what the number was, what you tried, what actually fixed it, and what you’d do differently. Be specific (“retrieval was returning the right article 61% of the time; adding BM25 took it to 79%; the remaining failures were date-sensitive questions, so we added a metadata filter”). Vagueness here costs offers.

30. How do you decide whether an AI feature should ship? Against pre-agreed thresholds on your eval, plus cost and latency budgets, plus a review of the worst failures rather than just the average score. Then a limited rollout with real users, watching both metrics and qualitative feedback. And a clear answer to “what happens when it’s wrong?”, because that determines how good it needs to be.

How to prepare in the last week

  • Re-run your project’s evals and refresh the numbers so they’re on the tip of your tongue.
  • Practise the two design questions above out loud, with a timer, drawing as you go.
  • Revise the fundamentals list (questions 1–8) until you can answer each in under a minute.
  • Prepare three failure stories, not one.
  • Read the company’s product and guess what their hardest AI problem is; ask about it.

For general interview technique, our guides on behavioural interview questions and how to ace your technical interview cover the parts that aren’t AI-specific.

Getting the interview in the first place

AI engineer listings vary a lot: some are retrieval-heavy enterprise roles, some are agent-and-product roles, some lean towards fine-tuning and inference. Your resume should lead with whichever the specific listing names, with your project quantified in quality, cost and latency. Tailr is a browser extension that tailors your resume to the job listing you’re viewing, generates a matching cover letter and tracks the application, so each version leads with what that team actually asked for. Try Tailr on the next AI engineer listing you open.

Conclusion

AI engineer interviews reward people who’ve built something real and measured it. Know how models behave, be able to design a RAG system and an agent and defend every trade-off, and above all be fluent in evaluation: how you’d build the set, what you’d measure, how you’d catch a regression. Bring specific numbers from your own project and honest stories about what broke. That combination is rarer than you’d think, and it’s what these thirty questions are really testing for.

Frequently asked questions

01What is asked in an AI engineer interview?

A typical loop has four parts: a coding round (usually ordinary data structures and API work in Python), a fundamentals conversation about how language models behave, a system design round where you design an AI feature such as a support assistant and are pushed on retrieval, evaluation, cost and failure modes, and a practical or take-home where you build or debug something with a model. Behavioural questions about a project that went wrong are common throughout.

02How do I prepare for an AI engineer interview?

Have one real project you can talk about in depth, including its evaluation set, baseline score, what you changed and what it cost. Be ready to design a RAG system and an agent on a whiteboard and defend every choice. Revise the fundamentals: tokens, context windows, embeddings, temperature, structured output, tool calling and prompt injection. And practise explaining a failure honestly, because you will be asked.

03Do AI engineer interviews include LeetCode-style coding?

Usually there is a coding round, but it's typically lighter than a pure software engineering loop: a practical problem such as parsing and chunking documents, calling an API with retries, or implementing a small eval harness. Some companies do use standard algorithm questions, so it's worth an hour or two of practice, but don't spend weeks on it at the expense of your project.

04What is the most important topic for AI engineer interviews?

Evaluation. Interviewers use it to separate people who've shipped reliable features from people who've built demos. Expect questions on how you'd build an eval set, which metrics you'd use, how you'd use a model as a grader without trusting it blindly, and how you'd catch regressions when a model version changes. If you can only prepare one area deeply, make it this one.

05What is the difference between RAG and fine-tuning, in interview terms?

RAG gives the model information at query time by retrieving relevant documents into the context; fine-tuning changes the model's weights by training it on examples. Use RAG when the knowledge changes often, needs to be cited, or is too large to train in. Use fine-tuning when you need a consistent style or format, a narrow task at high volume, or lower latency from a smaller model. Many production systems use both.

06How do I answer AI engineer system design questions?

Start with the user and the failure cost: what happens when the answer is wrong? Then walk through data ingestion, retrieval, the prompt and output schema, the evaluation set, latency and cost budgets, observability and safety. Name concrete trade-offs at each step, such as chunk size versus recall or a strong model versus a cheap one with reranking. Interviewers care more about your reasoning about trade-offs than about the exact architecture.