Skip to content

How AI Changed the Data Science Interview

How AI Changed the Data Science Interview
  • Author Avatar
    Written by:

    Tihomir Babic

The data science job changed. Job interviews changed. Did you adapt, or are you still preparing for an interview that no longer exists?

You’re preparing for data science interviews, but still no success in finding a job? Most likely, it’s because you’re preparing for an interview that doesn’t exist anymore. 

You’re preparing to write a query from scratch, but companies are now increasingly handing out AI-generated code and asking you to find what’s wrong with it and to fix it.

In this article, we’ll walk you through the two skills interviewers actually test now, using two real interview questions. After that, you’ll know what “prepared” means today. 

AI Is Turning Data Scientists From Code Writers Into Code Reviewers

Five years ago, you would open a blank code editor and write a query line by line. 

Today, you watch gray ghost-text appear before you’ve finished typing and decide whether to accept the suggestion. 

That’s how the job changed. Data scientists no longer write every line from scratch. They work alongside AI tools that draft the first pass. 

In short, the job shifted from producing code to judging it: catching what’s wrong, tightening what’s sloppy, explaining it to whoever asks. 

Interviews Caught Up With the Job Change

Interviews today simply reflect that essential change in the job’s focus. 

How AI Changed the Data Science Interview

The third one is especially important, as AI-written code rarely throws an error. It just silently does something subtly different than what you asked for, and nothing flags that for you. That’s your job.

The New Interview Skill: Finding Bugs in AI-Generated Code

Here’s a real interview question to simulate today’s interviews.

As usual, you get a coding question – say, it’s SQL –  that requires a certain problem to be solved. In this example, the task is to find the top 2 best-selling products in each category. Not overall – per category.

Last Updated: May 2020

MediumID 10555

Management wants to identify the most popular products within each category to optimize inventory and marketing strategies. Find the top 2 products with the highest total quantity sold in each category. If products within a category have the same total quantity, order them alphabetically by product name and assign consecutive ranks (1, 2, 3, etc.).

For example, if two products in the Electronics category both sold 15 units, then iPad Pro would get rank 1 (alphabetically first) and iPhone 14 would get rank 2 (alphabetically second).

Return the category, product name, total quantity sold, and rank within category. You should expect maximum 2 products per category in your results, though some categories might only have 1 product available.

Go to the Question

The Setup

This is where today’s interviews differ from yesteryear. You’re no longer asked to write the code that finds the top 2 best-selling products in each category. Instead, the code already sits in the editor, written by AI. For example, this code. 

PostgreSQL
Go to the question on the platformTables: ecommerce_transactions

Execution is clean on the first try, with no syntax errors. 

Pay Attention to the Output

Don’t expect AI to fail at syntax; that’s not where it’ll make errors. They will be in the code logic. 

Run this against the data we have, and you get this. 

categoryproduct_nametotal_quantity_soldcategory_rank
AccessoriesTesla Model 3 Keychain61
AccessoriesTesla Model 3 Keychain32
ApparelNetflix Hoodie31
ApparelNetflix Hoodie22
ElectronicsApple AirPods Pro71
ElectronicsiPhone 1572
GamingMeta Quest 341
GamingMeta Quest 332
Smart HomeGoogle Nest Hub71
Smart HomeGoogle Nest Hub52

That’s not right. Look at Accessories. The same product – Tesla Model 3 Keychain – takes both rank 1 and rank 2. That should be your first clue that something’s off. A “top 2 products” list where both spots go to one product isn’t a top 2 products list. The output actually shows one product’s two biggest transactions. 

The Electronics category is even worse. Apple AirPods Pro shows up at rank 1 with 7 units. But that’s not AirPods’ total sales; it’s just its single largest transaction. 

Now, this is what the query should return. The output is now very different. 

categoryproduct_nametotal_quantity_soldcategory_rank
AccessoriesTesla Model 3 Keychain171
ApparelNetflix Hoodie91
ElectronicsGoogle Pixel 8361
ElectronicsiPhone 15322
GamingMeta Quest 3121
GamingMicrosoft Xbox Series X122
Smart HomeAmazon Echo221
Smart HomeGoogle Nest Hub162

Google Pixel 8 is the best-selling Electronics product, but it didn’t appear anywhere in the buggy output. It got replaced by a product whose biggest transaction happened to be bigger than any of Pixel 8’s individual transactions. The real top seller disappeared, and nothing in executing the buggy query told you that. This is why bug-finding skills are now more important than memorizing syntax. 

Where the Bug Actually Is

It’s in quantity AS total_quantity_sold. This pulls the per-transaction quantity from the table. That’s why ranking runs over individual transactions, not over products. 

What you need is to add SUM() and GROUP BY.

PostgreSQL
Go to the question on the platformTables: ecommerce_transactions

The Second AI-Era Interview Skill: Explaining Code

The other facet of today’s interviews is that you, again, get the code. But this time, there’s no trick. No bugs are hiding anywhere. The code is correct. 

Instead, you’re tasked with explaining the code out loud and clearly, to simulate how you’d defend it in front of a reviewer at your job who’s going to push back. 

An example from the platform: the question wants you to return the top 3 posts by likes for each channel; ties should skip a rank.

Last Updated: November 2024

MediumID 10538

Identify the top 3 posts with the highest like counts for each channel. Assign a rank to each post based on its like count, allowing for gaps in ranking when posts have the same number of likes. For example, if two posts tie for 1st place, the next post should be ranked 3rd, not 2nd. Exclude any posts with zero likes.

The output should display the channel name, post ID, post creation date, and the like count for each post. Because there could be ties in rankings, your output could have more than 3 rows for each channel.

Go to the Question

The official solution is this code. 

PostgreSQL
Go to the question on the platformTables: posts, channels

Explain it!

Weak vs. Strong Explanation

A weak answer reads the code back top to bottom. “I create a CTE named ranked_posts. In it, I select the post and channel IDs, the post creation date, and the number of likes. Then I use the RANK window function to rank…”. You get the idea. 

A strong answer narrates the decisions sitting inside the code. Here’s what it looks like.

“I’m filtering out zero-like posts before ranking, not after.” The order matters. Rank everything first and filter later, and a channel’s genuine top 3 could get pushed out by zero-like posts sitting in the way the whole time. 

“I’m using RANK(), not ROW_NUMBER(), on purpose.” The task calls for ties to skip a rank, e.g., if two posts are tied for first, the next one lands at third, not second. RANK() handles that natively. ROW_NUMBER() would force a not-asked-for 1-2-3 order onto a tie. 

“PARTITION BY channel_id is what makes this per-channel instead of one global top 3.” Skip this clause, and the ranking runs over the entire posts table at once – the top 3 likes company-wide, not per channel. It's an easy clause to leave out by accident, and nothing about the query breaks or errors when you do.

“I only join to the channels table at the end, after ranking and filtering already happened.” There’s no point in dragging the full raw table through a join before the result set is already narrowed down.

Takeaway: Don’t describe what a line says. Explain why it was written the way it was.

Don't Explain What the Code Says. Explain Why It Was Written That Way.

We showed you an example of what that would look like. But a walkthrough that sounds good rarely goes untested. A good interviewer will push on the moment you finish explaining a decision. This is where most explanations fall apart.

Interviewees Push Past the Rehearsed Answer

Say you gave the strong answer, like the one above, and justified using RANK(). The natural next question is: “What if I told you the business wants ties to share the same rank instead, with no gap after?”

If you simply memorized the answer (rather than understood it), you freeze because you can recite what RANK() does, but you don’t have a clue what to swap it for, or why DENSE_RANK() solves exactly that case.

You Lost the Benefit of the Doubt

When you write your own code, an interviewer can safely assume you understand it. In today’s interviews, that assumption disappears because the code in front of you was AI-generated. 

Now you have to earn that appearance of authorship. This is what separates a candidate who memorized a walkthrough from one who could maintain this code on the job. 

How AI Changed the Data Science Interview

How to Prepare for Data Science Interviews in the AI Era

The job changed, and it changed the shape of the interviews. Now, you must change how you prepare for the interview. Practicing your own code in isolation won’t get you there. 

Here’s what’s worth doing instead. 

How to Prepare for Data Science Interviews in the AI Era

1. Read Other People’s Code on Purpose

Like you read books written by other people, read queries that you didn’t write. Don’t just review your own queries. Bugs hide better in unfamiliar code, because you don’t have the mental shortcut of remembering what you meant when you wrote it. 

How to do it? Pull queries from old projects, teammates, forums, anywhere. Practice asking “does this do what it claims?” before you even run it.

2. Practice Narrating Code Out Loud, Even When It’s Correct

Don’t let the interview be the first time you have to explain a PARTITION BY to another person.

Most people only talk through code when something’s broken. However, narrating correct code is harder, because there’s no bug pulling your attention toward what to say. You have to decide, unprompted, which decisions in the query are even worth defending.  

How to do it? Take a query that you know works and explain every clause using the phrase “instead of”. For example, not “this partitions by channel”, but “I partitioned by channel instead of leaving it as one global ranking, because otherwise the top posts from one channel would crowd out every other channel entirely. Do this out loud. If you can stand listening back to your voice, record yourself.

3. Learn the Bug Patterns That Repeat

There’s typically a small, recurring family of mistakes that show up across AI-generated queries. (We already explained the first one.)

How to Prepare for Data Science Interviews in the AI Era

These aren’t random. They all stem from the same AI’s blind spot: the code optimizes for looking right, not for matching intent exactly. 

How to do it? Keep a catalog of every AI-generated bug you personally catch, no matter how small. Don’t trust a query before you run it against that list. Did you confirm this join key is really unique? Did you check the boundary on every date filter? Did you verify the filter appears at the right step?

AI Didn’t Make SQL Less Important. It Changed What “Knowing SQL” Means

No, you may not be excused for not knowing SQL just because “AI writes SQL now.” SQL knowledge still matters, but a different kind.

Two Different Things Both Called “Knowing SQL”

Interviews have moved from insisting on syntax fluency to checking semantic judgment. 

How to Prepare for Data Science Interviews in the AI Era

Syntax fluency is remembering that RANK() takes an OVER clause, that a CTE starts with WITH, that you group before you filter with HAVING. This is the part AI is genuinely good at. It rarely misspells a keyword or forgets a comma.

Semantic judgment is knowing what a query actually does to the data, independent of whether it runs. This is the part every example in this article has tested – ranking on the wrong grain, the ordering of filter versus rank, the reason RANK() beats ROW_NUMBER() for a specific tie-breaking rule. None of that shows up as a syntax error. All of it shows up as a wrong answer.

A Small Example

Take a filter like WHERE status != 'canceled'. 

Syntactically, nothing is wrong with it. It'll run, and it looks like it does exactly what it says. 

Semantically, it silently drops every row where status is NULL, because SQL's three-valued logic treats NULL != 'canceled' as unknown, not true. (Read more about NULL semantics specifically.) The row disappears from your result set without a single error, warning, or wrong-syntax squiggle anywhere.

Practice the Skills Modern Data Science Interviews Actually Test

StrataScratch’s coding questions page now has two modes or filters.

Solve-the-Question is the classic setup: a problem, a blank editor, you write the query.

Debug-AI-Code is for practicing what we discussed in the article: the code editor already has AI-written code sitting there, and it doesn’t return the correct output. Your task is to find out why and fix the code. There are about 50 questions available in that mode currently.

Conclusion

The SQL interview didn’t get easier or harder because of AI. It just changed to reflect what the job actually looks like: AI writes the code, you decide if you trust it. (Whether that’s a good thing is up for debate.)

Today, everything comes down to two skills: catching the bug that runs clean but returns the wrong answer, and explaining a decision. 

In short, it’s about semantics, not syntax.

FAQ

1. How has AI changed data science interviews?

AI changed the actual job; interviews just reflect that. Interviews shifted from testing whether you can write a query from scratch to testing whether you can judge one someone (AI) already wrote. 

That means two things: finding bugs in code that runs without errors and explaining the reasoning behind a query.

2. Will AI replace SQL coding questions in data science interviews?

No. If anything, SQL questions are becoming more central. What’s changing is the format. 

The interview prompt is not “write the code”, but “here’s a query, find what’s wrong with it” or “explain why this was written this way”.

3. Can you use AI during a data science interview?

It depends on the company you’re interviewing for and the specific interview. Check with your recruiter rather than assuming either way. 

However, even in AI-assisted rounds, you’re still expected to catch its mistakes and explain its output.

4. What SQL skills are becoming more important because of AI?

Semantic judgment is now more important than syntax fluency. AI already has it covered. Because of that, knowing that a function exists matters less than knowing, for example, how NULL behaves in a comparison, or what a join does to your row count, or why an operation’s ordering changes a query’s result.

5. What are common mistakes in AI-generated SQL?

Four patterns appear frequently:

  1. Ranking or aggregating on the wrong grain, so a window function runs over raw rows instead of totals per group
  2. A missing PARTITION BY that lets a ranking run over the whole table instead of resetting per group
  3. A join that quietly duplicates rows because a key isn’t as unique as assumed
  4. A date filter that’s off by one at a boundary
  5. A filter applied after an aggregation when it needed to happen before

6. How do you debug AI-generated SQL in an interview?

Run it and look at the actual output. If it’s wrong, check the recurring failure points first: whether the data is aggregated to the right grain before a ranking or window function runs over it, window functions without a PARTITION BY, join keys that aren’t unique, filter ordering, and boundary conditions on dates.

7. How should you explain SQL code during an interview?

Justify decisions instead of describing syntax. Don’t say that “this filters for active rows”; the interviewer can see it themselves. Instead, say something like, “I filtered before ranking instead of after because ranking first could push out real results with zero-value rows”; it shows you understood the trade-off. 

Also, expect follow-up questions that test how well you understand what you just said.

8. How should I prepare for AI-era data science interviews?

Read code you didn’t write, not just your own. 

Practice narrating correct code out loud, not only broken code. 

Keep a running list of the AI bug patterns you personally catch, so you start recognizing them on sight instead of relearning them each time.

9. Is practicing SQL still worth it when AI can generate SQL queries?

Yes, even more than before. AI generating a query doesn’t mean the query is logically correct, i.e., that it returns what’s intended. Interviews are now built around that gap between AI writing syntactically perfect and semantically far-from-perfect code.

Share