# Data Engineer interview questions

Canonical URL: https://career1.ai/interview-questions/data-engineer · Publisher: Career1 (Answerflix Ltd)

Most data engineer interviews are not a SQL quiz. They are an hour spent finding out whether you have owned a pipeline that broke at 3am, or only built ones that worked.

Interviewers are testing three things: whether you understand what happens to data between systems, whether you plan for the day a source changes shape without telling you, and whether the thing you built could be handed to someone else.

## How this interview usually runs

1. **Screening call.** Your background, the stack you have worked in and the scale of the data. Expect to be asked what you owned rather than what the team built.
2. **SQL exercise.** Live or take-home: joins, aggregation, window functions and deduplication, with questions about what your query does with NULLs and ties.
3. **Pipeline or data model design.** How data for a described business gets ingested, modelled and kept fresh. Interviewers listen for failure handling, backfills and cost.
4. **Experience deep dive.** One pipeline you owned, end to end, including what broke. This conversation usually carries the most weight.

## Pipelines and reliability

### Walk me through a pipeline you built end to end.

What a strong answer shows: Whether you can name the source, the schedule, the failure modes and the consumer. Vague answers usually mean you inherited it.

### What happens when an upstream source changes its schema without warning?

What a strong answer shows: That you have been burned by this. Strong answers describe contracts, validation on ingest, and alerting before the dashboard goes wrong.

### Tell me about a pipeline that failed in production. What broke, and how did you find out?

What a strong answer shows: How you diagnose. People who have owned production talk about logs, lineage and the first metric that moved. Others talk in generalities.

### How do you handle a job that has to be re-run for last month's data?

What a strong answer shows: Whether your work is idempotent. Backfills separate people who designed for reprocessing from people who will corrupt a table trying.

### How do you handle data that arrives late?

What a strong answer shows: The difference between event time and processing time, a lookback window or watermark, and a plan for aggregates that were already published before the late rows turned up.

### What do you monitor on a pipeline besides whether the job succeeded?

What a strong answer shows: Freshness, row counts against a baseline, null rates and shifts in distribution. A green job that loaded zero rows is the failure strong candidates have already met.

### How do you test a data pipeline before it reaches production?

What a strong answer shows: Layers: unit tests on transformation logic, data quality checks on real output such as uniqueness and not-null constraints, and a run against a sample. Strong answers admit which of those they actually had.

## SQL

### Write a query that returns each customer's most recent order.

What a strong answer shows: Whether you reach for ROW_NUMBER() partitioned by customer and ordered by date, and what you do about two orders with the same timestamp. Ties are the follow-up.

### How would you remove duplicate rows from a table with no primary key?

What a strong answer shows: The question you ask first: what counts as a duplicate here, and which copy should win? Then a window function over those columns, keeping row number one.

### When does a LEFT JOIN quietly behave like an INNER JOIN?

What a strong answer shows: That a condition on the right-hand table in the WHERE clause throws away the NULL rows the LEFT JOIN kept, and that it belongs in the ON clause. Anyone who has debugged a report that lost rows knows this.

### What can a window function do that GROUP BY cannot?

What a strong answer shows: That a window keeps every row while calculating across related rows: running totals, rankings, or the previous value with LAG. A real use beats a definition.

## Modelling and storage

### When would you choose ELT over ETL?

What a strong answer shows: That you can argue both sides against cost and schema stability, instead of naming whichever your last employer used.

### How do you decide between a star schema and a wide denormalised table?

What a strong answer shows: That you think about who queries it and how often, not about which pattern sounds more professional.

### How would you handle slowly changing dimensions here?

What a strong answer shows: Practical familiarity. The follow-up is usually what you would do when someone needs history you did not keep.

### Why are columnar formats like Parquet faster for analytics?

What a strong answer shows: Reading only the columns a query needs, compression that works well on similar values, and column statistics that let engines skip whole row groups. Knowing the small files problem is a bonus.

### How would you partition a large events table?

What a strong answer shows: Partitioning by what queries actually filter on, usually a date, and knowing the opposite failure: thousands of tiny partitions from a high-cardinality key such as a user ID.

## Streaming and orchestration

### When would you choose streaming over batch?

What a strong answer shows: That the answer starts with how fresh the consumer really needs the data, not with the technology. Streaming costs more to run and to debug, and strong candidates say so.

### What does exactly-once processing mean in practice?

What a strong answer shows: Precision. Delivery is usually at least once, and the effect is made exactly once through idempotent writes, transactions or deduplication keys. Saying a tool simply guarantees it end to end is repeating marketing.

### How do you structure dependencies between jobs in an orchestrator like Airflow?

What a strong answer shows: Small, retryable, idempotent tasks, the difference between a schedule firing and upstream data being ready, and experience untangling a DAG somebody else built.

## Scale and cost

### A query that ran in 30 seconds now takes 20 minutes. Where do you look?

What a strong answer shows: A method rather than a guess: data volume, partitioning, the query plan, then the cluster. In that order.

### How do you keep warehouse costs from growing every quarter?

What a strong answer shows: Whether you have ever been accountable for a bill. Partition pruning, materialisation choices and killing unused tables all show up.

### A Spark job runs out of memory on one stage. What do you check?

What a strong answer shows: Data skew first: one key holding most of the rows. Then partition counts, a broadcast join on a table that is not actually small, and anything collecting results to the driver.

## Working with the people who use the data

### An analyst says the numbers in your table are wrong. What do you do?

What a strong answer shows: Reproducing before defending: find a specific record, trace it through lineage, and agree the definition. Often the data is right and two teams mean different things by the same metric.

### How do you handle personal data in a pipeline?

What a strong answer shows: Concrete practice: copying only what is needed, masking or tokenising identifiers, access controls, retention, and deletion requests that reach every downstream copy.

## Questions people ask

### What should I expect in a data engineer interview?

Usually three parts: a conversation about pipelines you have built, a SQL or modelling exercise, and a system design question about moving data between systems at some volume. The pipeline conversation carries the most weight, because it is the hardest part to prepare in advance.

### How much SQL do data engineer interviews test?

Nearly always some, and often live. Expect joins, aggregation, window functions and deduplication, and expect to be asked what your query does with NULLs and ties. Fluency matters more than clever syntax.

### How do I prepare if most of my work is under NDA?

Describe the shape of the problem without naming the employer or the data: the volume, the schedule, what broke, what you changed. Interviewers care about the reasoning, and being careful with a previous employer's details reads as a positive.

### Do I need to know Spark for a data engineer role?

Only if the job description asks for it. Far more interviews are lost on vague answers about pipelines you supposedly owned than on missing a specific framework, which most teams expect you to pick up.

### What system design question do data engineers get?

Usually a pipeline for a described business: ingest events or database changes, model them for analysts and keep them fresh. Interviewers listen for late data, backfills, failure handling and cost rather than for a particular vendor.

## Other roles

- [Interview questions and answers](https://career1.ai/interview-questions)
- [Node.js Developer interview questions](https://career1.ai/interview-questions/nodejs-developer)
- [AI Engineer interview questions](https://career1.ai/interview-questions/ai-engineer)
- [Python Developer interview questions](https://career1.ai/interview-questions/python-developer)
- [Frontend Developer interview questions](https://career1.ai/interview-questions/frontend-developer)
- [UX Designer interview questions](https://career1.ai/interview-questions/ux-designer)
- [Customer Success Manager interview questions](https://career1.ai/interview-questions/customer-success-manager)

## About Career1

Job seekers use Career1 for free: they upload a resume, practise privately with the AI interviewer, then take one vetting interview whose verified profile companies hiring through Career1 can search, unless they choose to hide it.
