AI Engineer interview questions

The role is new enough that interviews vary wildly. What they share is suspicion: everyone has now met someone who called an API twice and put AI Engineer on their CV.

Interviewers are checking whether you have shipped something that real users hit, whether you can tell if it is working, and whether you understand what it costs at volume.

Every question below is one that gets asked. What we have added is the part that usually goes unsaid: what a strong answer actually demonstrates. Career1 scores interviews for a living, so this is the difference we see between an answer that lands and one that was rehearsed.

How this interview usually runs

Most companies run some version of these stages. Smaller teams often fold two of them into one conversation.

  1. 01

    Screening call

    What you have shipped that uses a model, who used it, and what your part was.

  2. 02

    Coding

    Usually ordinary software engineering. Some teams add a small exercise built on a model API, such as extracting structured data or answering questions from documents.

  3. 03

    System design

    A retrieval or assistant feature for a described product, including how you would evaluate it, what it would cost and what happens when the provider fails.

  4. 04

    Project deep dive

    One system in detail: how you measured quality, what regressed, and what you changed.

Building with models

  1. Walk me through something you shipped that used a model in production.

    What a strong answer shows: Scope and ownership. The follow-ups are always about what went wrong, which is where borrowed projects fall apart.

  2. How do you decide between prompting, retrieval and fine tuning?

    What a strong answer shows: Judgement about cost and iteration speed, rather than reaching for the most sophisticated option available.

  3. How do you choose which model to use for a feature?

    What a strong answer shows: An evaluation on your own task rather than a public leaderboard, then latency, price, context length, the provider's data terms, and how hard it would be to switch later.

  4. When would you use structured outputs or function calling instead of parsing free text?

    What a strong answer shows: That anything a program consumes should be constrained to a schema and validated, and a plan for when validation still fails: retry, repair or fall back.

  5. How would you design an agent that can take actions, such as issuing a refund?

    What a strong answer shows: Restraint: narrow tools, permission checks outside the model, confirmation before anything irreversible, step limits and a log of every tool call. Strong answers ask whether it needs to be an agent.

  6. Tell me about a model feature that did not work. What did you do?

    What a strong answer shows: Honesty and judgement. Recognising that rules or a plain search box served users better, and removing the model, is a strong answer rather than a weak one.

Retrieval

  1. Your retrieval returns plausible but irrelevant chunks. What do you do?

    What a strong answer shows: Whether you have actually debugged a RAG system: chunking, embeddings, reranking and the query itself are all fair game.

  2. How do you choose a chunking strategy?

    What a strong answer shows: That it depends on the documents and the questions: chunks that follow the source's structure, metadata carried with each chunk, and measuring retrieval separately from the final answer.

  3. When would you combine keyword search with vector search?

    What a strong answer shows: That embeddings can miss exact strings such as product codes, error messages and names, and that hybrid retrieval with a reranker is a common fix, measured on your own queries.

  4. How do you keep an index correct when source documents change or are deleted?

    What a strong answer shows: Stable document IDs, re-embedding on change, deletions that really remove chunks, and knowing that switching embedding model means re-embedding everything.

Evaluation

  1. How do you know whether a change to a prompt made things better?

    What a strong answer shows: This is the question that separates the field. Strong answers describe a fixed evaluation set built before the change, not vibes.

  2. How would you evaluate something with no single correct answer?

    What a strong answer shows: Familiarity with rubric scoring, pairwise comparison and their biases, including the fact that a model judging its own output flatters it.

  3. How do you build an evaluation set when you have no labelled data?

    What a strong answer shows: Starting small and real: sampled or realistic inputs reviewed by hand, and a set that grows with every failure users report. A few dozen cases someone read can tell you more than thousands nobody did.

  4. What did you do about hallucination in something you shipped?

    What a strong answer shows: Whether you treat it as a system problem, with grounding and citations and refusal paths, or as something to apologise for in the docs.

  5. How do you catch a regression when a provider updates a model?

    What a strong answer shows: Pinned model versions where the provider offers them, the evaluation suite run before switching, and monitoring what users feel: refusals, format errors, answer length and complaints.

Safety and security

  1. What is prompt injection, and how do you defend against it?

    What a strong answer shows: That text from users, documents or web pages can carry instructions the model may follow, that there is no complete fix, and that the real defence is limiting what the model can do and treating its output as untrusted.

  2. How do you stop a customer-facing assistant answering things it should not?

    What a strong answer shows: A tightly scoped task, classifying requests before answering, a refusal path when retrieval finds nothing relevant, and tests written to break it before users do.

  3. How do you handle confidential data sent to a model provider?

    What a strong answer shows: Knowing the provider's retention and training terms, sending only what is needed, access control on retrieval so one user cannot pull another's documents, and care with what gets logged.

Cost, latency and operations

  1. Your feature works but costs too much per request. What are your options?

    What a strong answer shows: Practical levers: smaller models for easy cases, caching, shorter context, batching, and knowing which of those hurts quality.

  2. How do you make a slow LLM feature feel fast?

    What a strong answer shows: Streaming tokens, running retrieval and other calls in parallel, a smaller model where it is good enough, prompt caching where supported, and measuring time to first token, not only total time.

  3. How do you handle a provider outage or a rate limit in production?

    What a strong answer shows: That you have thought past the happy path: fallbacks, queues, and what the user sees while it is happening.

  4. What do you log for an LLM feature in production?

    What a strong answer shows: Inputs, retrieved context, model and prompt versions, outputs, latency, token counts and user feedback, so a bad answer can be reproduced, plus a policy for the personal data inside those logs.

Now practise it out loud

Reading questions is the easy half. Career1's AI interviewer asks questions like these by voice and follows up when an answer is thin. It takes about eight minutes and costs nothing.

Practise first, in private: once a week, no video and no score, with written feedback that is never shared with companies. When you are ready, the vetting interview is a separate single attempt, no retakes: it is recorded, scored and becomes a profile companies can find. You can hide that profile at any time.

Practise this interview free

Questions people ask

What is the difference between an AI engineer and an ML engineer?

Broadly, AI engineers build products on top of existing models and are judged on the system around the model: retrieval, evaluation, latency and cost. ML engineers are more likely to train and serve models themselves. The titles overlap and the job description matters more than the label.

Do I need a machine learning degree for an AI engineer role?

Rarely. Most teams hiring for this are looking for strong software engineering plus evidence you have shipped something that used a model and can tell whether it works. Evaluation experience is scarcer than model knowledge.

What do AI engineer interviews test most often?

Evaluation. Almost anyone can wire up an API call, so interviews concentrate on how you measured quality, what you did when it regressed, and what the thing cost to run.

Do AI engineer interviews include coding?

Usually, and much of it is ordinary software engineering: an API, some data handling, tests. Strong model knowledge rarely makes up for weak general engineering in these interviews.

Should I build a portfolio project for an AI engineer role?

One small project with an evaluation set and honest notes on what failed is more persuasive than several demos. Interviewers will ask how you know it works, so have that answer ready.

Other roles

Preparing for the general questions too? Common interview questions and what they test.

Hiring instead of interviewing? See how Career1 vets applicants for you.

Career1 help

Answers in seconds, any time

Ask anything about Career1. Leave your email so we can reply if the answer needs a person.

Thinking…

Passed to a person. We will reply to .