Three years ago, hiring an ML Engineer meant finding someone who could train a model.
Today, most ML Engineers being hired into SaaS teams will never train anything from scratch.
They will pick a foundation model, wrap it in retrieval, build an evaluation harness, and then spend eighteen months making the whole thing reliable enough to put in front of a paying customer. The title has not changed.
The job underneath it has, and most hiring processes have not caught up.
The job moved sideways, not up
The instinct when LLMs arrived was to assume the bar had risen. Surely you now need people who understand transformer architectures, attention mechanisms, the lot.
For a small number of teams, yes. For almost everyone building product on top of models they did not train, the bar did not rise. It moved. The scarce skill is no longer designing the model. It is everything that sits around it: prompt and context management, retrieval quality, latency and cost budgets, fallback behaviour when the model returns nonsense, and a way of knowing whether last week’s change made things better or worse.
Which means the strongest candidate for your ML role may well be a very good backend engineer who has spent a year shipping model-backed features, rather than a researcher with three papers and no production experience.
Evaluation is the skill that separates people
If there is one thing worth optimising your whole process around, it is this. Ask a candidate how they know their system is working.
Weak answers describe vibes. They demoed it, it looked good, the team was happy. Slightly better answers mention a small test set someone put together once. Strong answers get specific quickly. They talk about building a golden dataset from real user queries, about labelling disagreements between reviewers, about the difference between offline evals and what happens in production, about tracking regression when you swap model versions, about LLM-as-judge and its failure modes.
Evaluation is unglamorous and it is where most AI features quietly die. Candidates who have lived through that will tell you about it in detail. Candidates who have not will change the subject to model selection.
Retrieval is an engineering problem wearing an AI costume
Most teams discover this the hard way. The model was never the reason the answers were bad. The chunks were bad, the embeddings were stale, the metadata filtering was wrong, and nothing in the pipeline told anyone.
So screen for the boring parts. Does the candidate think about document chunking strategy as a real design decision? Have they dealt with keeping an index in sync with a source of truth that keeps changing? Do they know when a hybrid keyword and vector approach beats pure semantic search? Can they explain why re-ranking helped, or whether they ever measured it?
These are data engineering instincts more than machine learning ones, which is exactly the point.
What to stop screening for
A few habits from the old process have outlived their usefulness for most product teams.
Whiteboard derivations of backpropagation tell you almost nothing about whether someone can ship a reliable retrieval-augmented feature. Kaggle placings signal comfort with clean, static, well-specified datasets, which is close to the opposite of the job. Insisting on a PhD narrows a market that is already tight, for a role where the day-to-day work is production engineering.
Framework name-checking is the other trap. Whether someone has used a particular orchestration library matters far less than whether they can explain what it does and why they would or would not reach for it. Those libraries change every few months. Judgement does not.
What to screen for instead
Four things carry most of the signal.
Production instinct. Has anything they built been used by people who did not work at their company? What broke, and what did they do about it?
Cost and latency awareness. Anyone who has run a model-backed feature at any scale has a story about a bill or a p95 that got out of hand. If they have never thought about token spend or caching, they have not run one.
Scepticism. The best people in this space are cheerfully cynical about model output. They assume it will be wrong sometimes and they design for that. Uncritical enthusiasm is a risk, not energy.
Product sense. Knowing which problems are worth pointing a model at, and which ones are better solved with a rule, a form, or a bit of SQL. The engineers who reach for the model every time create more work than they finish.
The interview: Give them something broken
The most useful exercise we see teams run is not a build task. It is a debugging task.
Give the candidate a small system that mostly works. A retrieval pipeline over a few hundred documents, an obviously flawed prompt, no eval set, and a handful of real questions where the answers are subtly wrong. Then ask them to improve it and explain their reasoning as they go.
What you learn in an hour of that is worth three rounds of conversation. Do they measure before they change anything? Do they look at the retrieved context, or do they start rewriting the prompt? Can they tell the difference between a retrieval failure and a generation failure? Do they know when to stop?
It also respects candidate time, which matters more than ever in a market where good people are choosing between offers rather than chasing them.
Where the market actually sits
Compensation for this profile has decoupled from traditional data science bands and sits closer to senior backend or platform engineering, with a premium at companies where the AI feature is the product rather than an addition to it.
The people you want are rarely on the market for long, and they are unusually motivated by what they will get to work on. Vague AI ambition does not land. A specific problem, real usage data, and honesty about what is not working yet will get you replies from people who ignore everything else in their inbox.
One last thing about the title
If the work is applied, shipping features on top of foundation models, some teams now call it AI Engineer and find their applicant pool improves noticeably. If the work genuinely involves training and fine-tuning, keep ML Engineer and be specific about it in the advert.
Either way, write the job description around what the person will actually do in their first six months. The candidates worth hiring will recognise the honesty, and the ones who would have been wrong for it will select themselves out before you spend four rounds finding out.
Choosing the right recruitment agency is about fit, expertise, and trust.
By asking the right questions and digging into their processes, you can find a partner who not only fills roles – but helps you build a Product & Engineering team that fuels long-term SaaS growth.
Invest in a Product & Engineering Recruitment agency and accelerate your path to success.
Reach out to a member of the team here, or see more about how we can support your growth here.
