Hiring for AI in 2026 has one dominant failure mode, and it is not a shortage of candidates. It is that every profile now claims AI experience, so the usual screen — read the résumé, trust the keywords, interview the top ten — selects for people who write well about AI rather than people who have shipped it. This guide is the process we would use, in order, with the parts that actually change outcomes marked as such.
1. Decide which of the five AI roles you are hiring
Most bad AI hires are category errors made before the job spec was written. The five roles below look adjacent on paper and are different jobs in practice.
- [AI engineer](/hire/ai-engineers) — builds product features on existing models. Retrieval, tool calling, evals, latency, cost. The most common first AI hire and usually the right one.
- [LLM engineer](/hire/llm-engineers) — the narrower version of the same job, for teams whose whole product is model-shaped.
- [Machine learning engineer](/hire/machine-learning-engineers) — trains, serves and operates models you own. The hire you need when you have proprietary data, not when you have an API key.
- [AI researcher](/hire/ai-researchers) — creates capability that does not exist yet. Expensive, slow to pay off, and correct only when your differentiation genuinely depends on it.
- [AI product manager](/hire/ai-product-managers) — owns what good means for a non-deterministic feature, which is a harder question than it sounds.
The test: write down the first three things this person will ship. If they are features, you want an AI engineer. If they are models, you want an ML engineer. If they are papers or novel capability, you want a researcher. If you cannot write the three things, you are not ready to hire yet — that is a strategy problem wearing a recruiting costume.
2. Write the spec around evidence, not around years
Years of experience is a broken filter here. The tools that define the job are two years old; asking for five years of LLM production experience selects for people willing to misrepresent, and excludes the strongest builders in the field.
Replace the experience line with an evidence line. Instead of "3+ years building AI systems", write "has shipped at least one LLM-backed feature to real users and can describe its evaluation". That sentence is checkable, and it filters harder than any number of years.
The best predictor we have found is not seniority or pedigree. It is whether someone can describe a failure of their own system in specific terms — and what they measured to find it.
3. Source from evidence rather than from job titles
Job titles lag reality by about eighteen months in this field. Plenty of people doing excellent AI engineering are titled backend engineer, data scientist, or founder. Searching titles alone will miss them.
This is the specific problem AI MAXXERS exists to solve: public professional signals are read and scored 0–100 across nine AI capabilities, with the evidence behind each score attached, so you can search for people who demonstrably do the work rather than people who describe it. Search by skill — RAG, agents, evals, fine-tuning — as well as by role, because the skill is what you are actually buying.
- Filter to the capability you need, not the title.
- Set a minimum score to skip the long tail, then read the evidence lines rather than the numbers.
- Shortlist ten. Contact five with something specific about their work in the first sentence.
4. Run an interview that a bluffer cannot pass
Take-home projects are a poor filter for AI roles: the whole job is now assisted by AI, so a take-home mostly measures who had a free weekend. Replace it with a conversation about something they already built. Three questions, in this order:
- Walk me through one AI system you shipped, end to end. Listen for whether the architecture has reasons attached to it.
- What broke, and how did you find out? People who have operated a system in production answer instantly and in specifics — retrieval recall, a bad chunking strategy, a silent provider change. People who have not, generalise.
- How did you know your change made it better? This is the evaluation question, and it is the single highest-signal question in an AI interview. The absence of an answer is itself the answer.
If you want a practical exercise, do it live and let them use whatever tools they use daily, including AI. Watching someone debug a broken retrieval pipeline for forty minutes with their real toolkit tells you more than a week of unobserved take-home work.
5. Expect to be outbid, and compete on something else
You will not win a bidding war against a funded AI lab, and you do not have to. What consistently moves strong candidates: ownership of a whole surface, a short decision loop, access to real users, and a team that already has evaluation discipline so their work will not be judged on vibes. Say those things concretely in the first message.
For the numbers themselves — bands by role, seniority and market, and what actually moves them — see what AI talent costs.
6. Close in days, not weeks
The most common reason a good AI candidate says no is that someone else finished first. Compress: one screen, one technical conversation, one meeting with whoever they will work with, decision within a week. Every extra round costs you more candidates than it filters.
A hiring loop you can run this week
- Write the three things this person ships in their first quarter.
- Pick the role category that matches those three things.
- Search the database by capability, shortlist ten from the evidence.
- Send five specific messages referencing their actual work.
- Run the three-question conversation. Decide inside a week.
The whole loop is roughly ten hours of your time. Most of the failure in AI hiring happens in the first step, which takes twenty minutes and is usually skipped.