Writing

What I look for when hiring an ML engineer

In short: I look for someone who checks the data before the model, gets suspicious when a result looks too good, has shipped at least one thing end to end, and can explain it to someone who isn't an…

published
read time
5 min
words
956
lang
en
filed under
Career

In short: I look for someone who checks the data before the model, gets suspicious when a result looks too good, has shipped at least one thing end to end, and can explain it to someone who isn't an engineer. Knowing the newest architecture is the easiest part to teach.

Both sides of the table

I've been on both sides of this table. I've interviewed for machine learning roles in a hospital network, at a large hardware company and at a medical AI company, each with a very different idea of what "ML engineer" means. And while leading machine learning work at a hospital network, I've been on the side that has to decide who joins. The two experiences taught me different things, and they mostly agree on what matters.

What looks good, and what tells me something

Looks good in an interview

  • A long list of frameworks
  • The latest model names
  • A top leaderboard score
  • A polished notebook
  • Confident answers to everything

Tells me something

  • One project explained three levels deep
  • A story about a result that turned out wrong
  • Something running behind an endpoint
  • Questions about the data before the model
  • "I don't know, here's how I'd find out"

The left column isn't bad. It just doesn't separate people. Anyone can list PyTorch, scikit-learn and Hugging Face. What I can't fake in a conversation is having been burned by real data and learned from it.

Five things I look for

They look at the data first

I like to describe a messy situation and ask what they'd do in the first hour. Say, voice recordings from three clinics, labelled by different people, and a request to predict a diagnosis. A strong candidate asks where the recordings came from, whether the same patient appears more than once, how the labels were made, and what the class balance looks like. A weaker one starts picking a model. In health data especially, the model is rarely where projects die.

They distrust good numbers

I ask about a time a result looked great and wasn't. Everyone who has done this work for real has a story: a patient in both train and test, a scaler fit on the whole dataset, a feature that quietly encoded the answer. If someone has never had that experience, they either haven't worked on messy data or haven't checked. Both are fine for a junior role. Neither is fine for a senior one.

They've shipped something

Not necessarily in production at scale. A small API in Docker, a demo someone else used, a pipeline that ran every night without them. Taking a model from a notebook to something another person can call teaches you about versioning, latency, input validation and logging. Those are most of the job. I'll take a modest model that's deployed over a brilliant one that lives in a notebook.

They can explain it to a non-engineer

In a hospital you explain your model to clinicians. At a hardware company you explain it to product people. Either way, if you can't say what the model does, when it fails and what the numbers mean in plain words, it won't get used. I sometimes ask a candidate to explain precision and recall as if I were a nurse deciding whether to trust an alert. It's a short question that tells me a lot.

They know where their knowledge ends

"I haven't used that, but I'd start by reading how it handles X" is a great answer. Bluffing is the one thing that ends an interview for me, because in this field a confident wrong answer becomes a confident wrong model.

What I care less about

  • Memorised theory. I don't ask people to derive backpropagation on a whiteboard. I want to know they'd notice if training was broken.
  • The exact stack. GCP or AWS, Flask or FastAPI, MLflow or Weights and Biases. Good engineers switch in a week.
  • Leaderboard results. Useful as a sign of persistence. Not a sign they can handle a dataset nobody cleaned for them.
  • Domain knowledge, at first. I've moved between health, hardware and voice myself. Curiosity about the domain matters more than arriving with it.

How I'd run the conversation

My preferred interview is one project, explored deeply. The candidate picks something they built. I keep asking why. Why that split? Why that metric? What happened when it broke? What would you change now? Three or four levels down, you find out whether someone made the decisions or just followed a tutorial. It's also kinder than a trivia quiz, because people talk best about their own work.

I'd add one short practical piece: a small dataset with a deliberate problem in it, like duplicated patients or a leaky column, and a question about what they notice. Not a take-home that eats a weekend. Thirty minutes, together, talking out loud.

From the other side of the table

When I've been the one interviewing for a job, the conversations that went well were the ones where I talked about what went wrong and what I changed. The ones that went badly were the ones where I tried to sound impressive.

If you're the candidatePick one project and prepare to go deep on it: the data, the split, the metric, the failure, the fix, and how it ran once it left your laptop. Write those six answers down before the interview. That one project will carry more weight than ten listed on your CV.

If you're the one hiring, try replacing one round of theory questions with a deep dive into a single project the candidate chooses. Keep asking why until they say "I don't know". How they handle that moment is most of what you need to learn.

related

Keep reading