Hiring an AI Developer? 10 Questions to Ask Before You Start

Published: August 9, 2026 — Most AI projects don't fail on the demo. They fail in production, in the security review, or three months in when the freelancer goes quiet. These 10 questions — drawn from how working AI consultancies evaluate themselves — surface that risk while you can still walk away.

🔍 Quick Takeaways

Why the Demo Isn't the Test

Anyone can demo an AI agent in 2026. Far fewer have shipped one that survived real users, real data, and a real compliance review. The gap shows up in a specific place: the last 20% of the work — securing the system, hardening it for real traffic, handling edge cases, passing a security review — which is where most AI projects die.

Ask these questions before you sign, not after the project stalls. Strong developers answer most of them before you ask, because owning the last 20% and protecting your data is simply how they work.

The 10 Questions

1. Who owns the last 20% — security, architecture, and production hardening?

Ask directly: who secures the system, handles real traffic, and fixes edge cases? Green flag: the developer treats production hardening as their job, and talks about evals, guardrails, monitoring, and failure modes unprompted. Red flag: "we deliver the model and you productionize it" — you're buying a prototype, not a product.

2. How do you handle our data — and will the model train on it?

This question decides your security review. Green flag: specifics, not reassurance — where data goes, who can access it, whether it leaves your environment, no training on your data by default, audit trails, and familiarity with the compliance you need (SOC 2, HIPAA, GDPR, financial rules). Red flag: vagueness. If your data is sensitive, this alone is disqualifying — see Data Sovereignty: Why Your AI Should Run Where Your Data Lives.

3. What does "done" mean, and how will we measure it?

"Done" is where scope disputes live. A system that works in a demo can be wrong 15% of the time in production. Green flag: acceptance criteria up front — accuracy targets on your data, latency, uptime, evaluation methods. Red flag: "we'll see" or an undefined success bar. If "done" is undefined, every miss becomes your problem.

4. How do you price — and what happens when scope changes?

AI projects evolve as you learn what the model can and cannot do. Green flag: transparent pricing (fixed-scope sprints or clear rates), a defined change process, no long-term lock-in — and ideally a small paid pilot. Red flag: big commitment required before any trial. For budget reality, see How Much Does a Custom Local AI System Cost in 2026?

5. Who actually writes the code — and how senior are they?

Many firms sell you seniors in the pitch and staff the build with juniors. Green flag: named, senior engineers who own your project end to end — ask for your day-to-day contact by name and whether they've shipped production AI. Red flag: a rotating pool of contractors and meetings that stay with the salesperson.

6. What happens when the model hallucinates or takes a wrong action?

Every AI system fails sometimes; the question is whether the developer designed for it. Green flag: grounding in your data (RAG), constrained tools, human-in-the-loop for high-stakes actions, and monitoring that catches regressions before users do. Red flag: reliability treated as an afterthought. Related: Why RAG Hallucinates and How to Fix It.

7. Can you show real production work, not just a demo reel?

Green flag: real case studies with outcomes, references you can call, and specifics about what broke and how they fixed it. Red flag: a demo reel and no production references. Experience is the one thing a slick pitch cannot fake.

8. How do you evaluate the system before you call it done?

Ask what test set they'd build from your documents and how they score retrieval versus generation. Green flag: a concrete evaluation plan — test questions mined from your real usage, faithfulness and relevance scored separately, thresholds set. Red flag: "we'll test it in the demo." This is how you'd verify it yourself: How to Evaluate a RAG System Before You Trust It.

9. What's the plan for your data to stay private — architecturally?

If your use case is confidential (legal, medical, financial), ask specifically: on-premise or local deployment, where models run, what leaves the building. Green flag: the developer proposes running models inside your boundary by default and can explain the trade-offs. Red flag: the only option on the table is sending data to a third-party API. Deep dive: How to Deploy an LLM Inside Your Own Security Boundary.

10. What happens after launch — who maintains it, and what does that cost?

Models drift, libraries update, security patches land. Green flag: a clear maintenance plan and honest cost (industry norm is 20–30% of build cost per year for enterprise systems). Red flag: disappearing after delivery. A handoff without maintenance is a system that quietly decays.

The Cheat Sheet: Green vs. Red Flags

Question Green flag Red flag
The last 20%Owns production hardening"You productionize it"
Your dataSpecifics + no training by defaultVague reassurance
Definition of "done"Measurable acceptance criteriaUndefined, "we'll see"
PricingTransparent + trial or pilot offeredBig commitment, no trial
Who codesNamed senior engineersRotating juniors
Failure handlingGuardrails + evals by designReliability as afterthought
ProofReal case studies + referencesDemo reel only

Frequently Asked Questions (FAQ)

What should I ask before hiring an AI developer?

Ask about the last 20%: who owns production hardening, security, and edge cases? Then data handling (where does your data go, is it used for training?), a measurable definition of done, transparent pricing with a pilot, who actually writes the code, what happens when the model hallucinates, and proof of real production work.

What are the biggest red flags when hiring for AI?

Vague data-handling answers, no measurable definition of done, reliability treated as an afterthought, a rotating pool of junior contractors, and a demo reel with no real production references. Any one of these is a reason to keep looking.

Should an AI developer offer a trial or pilot?

Yes — a risk-free trial or small paid pilot is one of the strongest signals of confidence. It lets you verify quality, communication, and fit before a larger commitment, and a partner sure of their work will offer it.

How do I know if an AI partner is actually senior?

Ask who your day-to-day engineer is by name, whether they have shipped production AI before, and to speak with them directly before signing. A partner staffing seniors will introduce them; one hiding juniors will keep you talking to a salesperson.

What does 'done' mean for an AI project?

A shared, measurable definition of success agreed before work starts: accuracy targets on your data, latency, uptime, and evaluation methods. For AI especially, a system that works in a demo can be wrong 15% of the time in production — 'done' must be measured, not felt.

🔍 Want to work with someone who answers these questions before you ask?

I build private, local AI systems — RAG with cited answers, on-premise LLMs, and multilingual assistants — through Haal Lab, with transparent scoping and no lock-in. Contact me for a scoping conversation, no obligation.