How to Choose an Agentic AI Development Partner: An Evaluation Framework
Mohammed Usman is the founder and CEO of Masarrati with 15+ years in product engineering. He has led the development of 10+ production AI, blockchain, and cybersecurity platforms for enterprise clients across UAE, MENA, and Europe.
TL;DR
Demonstrations no longer separate suppliers, so score on problem framing, evaluation and observability, integration reality, governance, and continuity — then eliminate on ownership and transfer before comparing anything else. Ask what a supplier refused to automate, who owns the prompts and evaluation sets, and whether you could run the system without them in six months. Test with a bounded paid discovery on real data, and do not hire at all when the capability is core, the process is ambiguous, or a mature product already fits.
Updated August 1, 2026
Every firm demonstrates well now. The tooling is good enough that a competent team can show a convincing agent in a fortnight, which makes the demonstration close to worthless as a selection signal. What separates suppliers is not what they show you in a meeting. It is what they do in the months when the data turns out to be inconsistent, the workflow has undocumented exceptions, and the agent has to fail safely in front of real customers.
This is a framework for assessing that. It assumes you are buying a delivery system rather than a prototype, and it ends with the questions we think you should put to any supplier, including us.
Decide what you are actually buying
Three different purchases get discussed under one heading, and suppliers are rarely equally good at all three.
Feasibility. Can this workflow be handled by an agent at all, and what would it cost to run? The output is evidence and a recommendation, sometimes a negative one.
Delivery. Build it, integrate it, get it into production and keep it running.
Capability transfer. Build it alongside our team so that we can carry it afterwards.
Say which one you are buying in the first conversation. A supplier optimised for delivery will quote a build when what you needed was a short feasibility answer, and you will not discover the mismatch until the invoice arrives.
Prepare your side of the table first
Before you shortlist, write down four things: the specific workflow with its exceptions; who the internal decision-maker is and how often they are available; what data you can realistically provide and how quickly; and what a correct outcome looks like, expressed as cases rather than adjectives.
Suppliers who ask for these in the first meeting, and who push back when the answers are vague, are showing you how they will behave later. Suppliers who move straight to architecture are showing you something too.
Five dimensions worth scoring
Problem framing. Will they tell you not to build something? Judgement about when an agent is the wrong instrument is the rarest capability in this market and the most expensive to lack.
Evaluation and observability. Whether they can prove what the system does, before and after it reaches production.
Integration reality. Whether they have worked with enterprise systems, permission models, partial failures and imperfect data, rather than clean APIs and curated samples.
Governance and security. Data handling, autonomy boundaries, auditability, and whether these are designed in or retro-fitted when your security team starts asking.
Operating model and continuity. Who works on it, what is handed over, and whether you could run it without them.
The questions, and what a good answer sounds like
Ask these in a working session rather than in writing. You are assessing how they think, and the useful signal is in the follow-up.
Tell me about a project where you advised a client not to build an agent. A good answer is specific: the workflow, why a deterministic pipeline, a retrieval interface or a rules engine fitted better, and what the client did instead. A weak answer claims agents suit almost anything, or offers an example so trivial that declining it cost the supplier nothing.
How do you define correct for this workflow, and when do you build the evaluation set? A good answer puts evaluation before implementation, builds the set from your real and messy records including adversarial and edge cases, treats the set as yours, and runs it as a regression on every change. A weak answer describes manual testing or leans on the general capability of the underlying model.
Show me the observability you ship. After the fact, what can I see about a single decision? A good answer covers step-level traces, tool calls with inputs and outputs, versioned prompts and configuration, cost and latency per task, and the ability to replay a case. A weak answer is application logs and a dashboard.
What happens when the agent is not confident? A good answer describes explicit thresholds, a defined human queue with an owner, what the system does while it waits, idempotent actions, an undo path for anything already done, and a documented stop control. A weak answer is that it retries or asks the user.
Which parts of this workflow would you refuse to automate in the first release? A good answer names them without prompting and gives the criteria under which scope would later expand. A weak answer accepts your full scope enthusiastically.
Where does our data go, what is retained, and what is used for training? A good answer is precise about residency, retention periods, redaction, tenancy separation and the terms of any model provider involved, and is willing to put all of it in the contract. A weak answer reassures without specifics.
What is your model strategy if a provider changes pricing, deprecates a model or is unavailable? A good answer describes an abstraction layer, per-task model selection, evaluation-driven substitution, fallbacks and, where required, a self-hosted route. A weak answer is loyalty to one provider.
Who owns the code, prompts, evaluation sets and infrastructure at the end? A good answer is that you do, from day one, in your repositories and your cloud accounts. A weak answer involves hosting on the supplier's platform with no described exit.
How will you estimate and control run cost? A good answer instruments cost per task, tiers models by task difficulty, caches and batches where it can, sets budget caps and alerts, and is candid that a firm figure is impossible before real volumes are known. Treat early precision here as a warning rather than as reassurance.
Who exactly works on this, and what happens if they leave? A good answer names people, describes how work is documented so that it survives them, and accepts a continuity clause. A weak answer is a capability statement about the firm.
What did you do when the source data turned out to be poor? A good answer talks about reconciliation, retries, partial failure, back-pressure, permission inheritance and the unglamorous cleanup that preceded the interesting work. A weak answer has never met bad data.
Could we run this without you in six months, and what would that take? A good answer is yes, with a plan: runbooks, training sessions, a defined transfer period and paid exit assistance if wanted. A weak answer treats the question as a sign of mistrust.
What went wrong on your last engagement, and what did you change afterwards? A good answer is structural and admits fault. A weak answer blames a client. Ask as well for a reference from a project that was cancelled or reduced in scope. How a supplier behaves when work shrinks is more informative than how they behave when it grows.
Scoring and elimination
Score each dimension, weight evaluation, observability and continuity most heavily, and apply elimination rules before comparing totals. Eliminate any supplier that cannot answer the ownership question clearly, that will not commit to a transfer plan, that declines a bounded paid discovery, or whose references are all demonstrations rather than systems still running.
A supplier scoring moderately across every dimension is usually safer than one scoring at the top on capability and poorly on continuity. The second kind produces systems you can neither maintain nor leave.
Test cheaply before you commit
Run a bounded, paid discovery before any large engagement. Give it a fixed scope, your real data rather than a sample chosen to flatter, defined acceptance criteria, and a deliverable set that includes an evaluation harness, an architecture proposal, a run-cost model and an honest recommendation, including the option not to proceed.
Then watch how they behave when the data is worse than they expected, because it will be. Do they renegotiate scope openly, or absorb the problem quietly and surface it later as a delay? That behaviour is the thing you are actually buying.
When you should not hire a partner at all
There are situations where the right decision is to keep the work inside, or not to start yet.
If the workflow is core to your product, will change continuously, and you can hire and retain the engineers, build the team. Renting a permanent capability is a slow way to lose it.
If the real obstacle is that the process is ambiguous, ownership is contested or the underlying data is unreliable, no supplier can resolve that for you. Bringing one in at that stage converts an internal problem into a contractual one, at a worse price.
If nobody internally can be freed to make weekly decisions and supply data, delay. Engagements without an available owner do not fail loudly. They drift, and the drift is charged to you.
And if a mature product already does most of the job, buy the product. Configuring something that exists is almost always the cheaper route to the same outcome, and a supplier worth hiring will say so before you ask.
Red flags
Certainty in the first meeting, before anyone has seen the data. A proposal with no line item for evaluation or observability. Pricing by number of agents rather than by outcome or scope. Reluctance to name the team. Ambiguity about ownership. Demonstrations only on the supplier's own data. And a roadmap that increases autonomy on a fixed schedule rather than in response to evidence.
Put these questions to us as well
Masarrati is one of the firms you might be evaluating, and this framework is not written to flatter us. Ask us the ownership question, the transfer question and the what-went-wrong question, and compare our answers with anyone else you are speaking to. If ours are thinner, that is useful information and you should act on it.
Deciding
Choose on evidence of production systems rather than demonstrations. On evaluation and observability practice rather than familiarity with model names. On willingness to reduce your scope rather than accept it whole. And on whether you could run the result without them.
Define what you are buying, prepare your own side of the table, test with a bounded paid discovery, and eliminate on ownership and continuity before comparing anything else.
If you want to see how we structure discovery, delivery and transfer, our engagement models and SaaS to agentic AI pages are the place to start, and unfamiliar terminology is defined in the glossary.