Three proposals sat side by side in a manufacturer's meeting room. All three promised efficiency through AI, and all three cost about the same. The managing director finally asked one question: once this is live, which number moves, and by how much? Two providers scrolled back to the demo. The third said: I haven't measured your baseline yet, so I don't know — let's measure it first.
I invented that meeting for this article; I did not invent the difference between the three proposals. Most providers that have appeared in Türkiye over the past two years under the label of an AI agency sell the same two products: a chatbot and a content automation. Both are real work. Neither one changes a factory's quoting process, a retailer's stock flow or a service network's work-order routing.
The line starts here: an AI consultant is not in the business of installing a model but of building the process and the measure that model attaches to. The twelve questions below exist to bring that line to the surface in the first meeting. Under each one sits what to listen for in the answer — because the real information isn't in the question, it's in where the provider hesitates.
What is the difference between an AI agency and an AI consultant?
An AI agency usually delivers a tool: a chatbot, a content pipeline, an interface over an off-the-shelf model. An AI consultant commits to an outcome — naming in advance which metric in which process will improve and by how much, measuring it before the work starts, then building the system around that measure. They aren't rivals but different jobs; the trouble is that both get sold with the same deck. The title hasn't settled either: some say AI consultant, others AI transformation advisor — both describe the same work.
The difference is not academic; it lands on the invoice. Standing up a chatbot takes a few days today, and the price should say so. Translating an entire quoting process into an engineer's language — wiring the catalogue, the CRM and the response flow into one system — is months of engineering work, and it returns a number you can measure. The frame in this article came out of the projects we run on the AI advisory side.
The twelve questions, at a glance:
- Are you selling a model or an outcome?
- Which case and which number back your efficiency claim?
- Where in our business would you refuse to use AI?
- Which process would you pick for the first pilot, and why that one?
- Who does the data preparation, and how long does it take?
- Which baseline do you measure before the project starts?
- How many weeks until the pilot is live?
- How do you connect to the systems we already run?
- What happens when the model gets an answer wrong?
- Where is our data processed, and who carries the compliance duty?
- Which model provider are we locked into?
- When the project ends, who owns the system and who on our team has learned what?
1. Are you selling a model or an outcome?
The point of the question is to hear how the other side describes itself. A provider selling a model opens with technology: which model it uses, which interface it builds, how many integrations it has shipped. A provider selling an outcome opens with your numbers — how many hours a process consumes, what those hours cost, and past which threshold the investment pays for itself.
That is why the first five minutes tell you something. Opening with technology isn't wrong, it's incomplete: which model gets used is an implementation detail and may change within six months, whereas which metric moves is the subject of the contract and shouldn't.
Listen for this: is the provider asking about the numbers in your business, or presenting its own toolkit?
2. Which case and which number back your efficiency claim?
A good answer carries three components: a piece of work, a metric and a time frame. "We delivered serious efficiency gains for our clients" is a courtesy, not evidence; don't move to the next agenda item before you hear which work, which number and over what period.
Here is one from our own side, to show how the claim should be built. At Meccanotecnica Umbra Türkiye we connected the product catalogue to an AI technical advisor that works out the right equipment for an engineer describing their plant, and to a quote portal: quote requests rose tenfold and the time between request and response fell by ninety percent. The sentence names the work (quoting), the metrics (request volume, response time) and both direction and size.
Listen for this: does the number come with a baseline? "Up forty percent" that never says forty percent of what is decoration rather than measurement.
3. Where in our business would you refuse to use AI?
A provider who says "everywhere" has probably not yet used it seriously anywhere. An experienced consultant will name at least two areas: decisions where the cost of error keeps a human in the loop, and processes whose history is too thin to feed a model at all.
Drawing that line is concrete work. Price approval, workplace safety calls and one-off strategic choices are not the model's job. Repetitive work with a writable rule and a record behind it — preparing quotes, matching orders, searching technical documents — is where the model pays first and pays most.
Listen for this: can they say what they won't do? A provider who can't draw the boundary can't measure the result either.
4. Which process would you pick for the first pilot, and why that one?
A good answer rests on three criteria: how often the process repeats, what an error costs today, and what state the data is in. A consultant can't propose a process without asking all three; if they do, the proposal comes from their shelf, not from your business.
The logic of that choice runs against instinct. The most profitable first area is usually not the most visible one but the most repeated: a ten-minute task done thirty times a day is a bigger lever than a two-day task done once a month. We set out the full sequence of a transformation in what AI transformation is.
Listen for this: is the proposed process a bottleneck you described, or an example already sitting in the deck?
5. Who does the data preparation, and how long does it take?
In most projects this is the largest part of the work, and in most proposals it is the missing line. Data preparation — collecting sources, cleaning them, labelling them and opening access rights — takes longer than standing up the model, so who does it belongs in writing before the contract.
"You supply the data, we'll handle the rest" quietly leaves half the cost with you. A sound answer names the split: which source comes from you, which transformation the provider performs, how many weeks each step takes and where the schedule slips if one runs late.
Listen for this: is the time allotted to data preparation shorter than the time allotted to building the model? If it is, either they can explain why or they have never done this work.
6. Which baseline do you measure before the project starts?
Nobody can prove that an unmeasured process improved. A serious provider records today's value before starting: how many minutes, how many steps, how many errors, how many requests, how many people.
You cannot take that measurement afterwards. The only reason we can say response time fell by ninety percent at Meccanotecnica Umbra is that we measured the time before it fell. A baseline of zero is a fine start too: at SIM Printing Suppliers, visibility across AI engines read zero at first measurement and reached 40,000 once the programme ran — a result that starts from zero needs no interpretation.
Listen for this: is the baseline a line item in the proposal, or a verbal intention?
7. How many weeks until the pilot is live?
An enterprise pilot's time to live is measured in weeks, not months — and that number is a direct read on the provider's engineering depth. "First output in six months" describes a programme rather than a pilot; "we'll have it up in two days" describes a demo rather than a pilot.
The range shifts with the type of work; the logic holds. A pilot is the smallest working system that runs in one process, with one team, on real data. Its purpose is to measure rather than to impress, and its output is a decision rather than a presentation — scale it or stop it.
Listen for this: is there a defined exit criterion? A pilot without one doesn't end; only its budget does.
8. How do you connect to the systems we already run?
AI's value sits not in the model but in what it attaches to. A system that can't read from and write back to your ERP, CRM or production software hands your staff a second screen and a third workload — and falls out of use within six months.
The integration layer is usually where the real engineering happens. The platform we built for our German client MKComputer syncs more than 200,000 products in five minutes and routes orders with zero manual steps. At that scale the question isn't which model; it's at what speed, at what error rate and under which failure scenario the data moves.
Listen for this: is it settled who writes the integration? "Your IT team opens the API" is an assumption, not a plan.
9. What happens when the model gets an answer wrong?
Every model gets things wrong; a serious provider says so before you ask and shows a design that caps the cost of being wrong. A good answer has three layers: the source each answer is grounded in, a rule that hands the task to a human below a confidence threshold, and a log that makes every answer traceable after the fact.
For a manufacturer selling technically this is not negotiable: recommending the wrong equipment costs more than one lost customer. A model inventing answers where it has no data — hallucination, in the industry's term — can be bounded by design; once each answer is grounded in product data, the model stops inventing products of its own.
Listen for this: when errors come up, do they get defensive or do they describe the design?
10. Where is our data processed, and who carries the compliance duty?
The answer should fit in one sentence: on whose servers and in which country the data is processed, how long it is retained, and whether it feeds model training. Where personal data is involved you are the data controller and the provider is the processor — and that distinction belongs in the contract in writing.
The practical consequence is blunt: if the contract omits processor status, retention period and the list of sub-processors, the whole duty stays with you. Enterprise buyers read that clause before the price, because price is negotiable and liability is not.
Listen for this: can the provider name its own sub-processors? If not, they don't know where the data chain ends.
11. Which model provider are we locked into?
The model layer is the fastest-moving and fastest-cheapening part of this field; today's best model can be second-best in six months. A well-built system keeps the model provider as a replaceable component — the business logic, the data flow and the interface stay put while the model changes.
The commercial translation is plain. If your system has to be rewritten when the provider doubles its price or retires a model, that system is your supplier's asset rather than yours. Dependency isn't something you avoid entirely; it's something you measure and bound.
Listen for this: asked how long it takes to swap the model, can they name a duration?
12. When the project ends, who owns the system and who on our team has learned what?
Ownership is defined across four items: the source code, the rule and prompt sets that carry the system's operating logic, the data it produces, and title to the accounts. If all four don't end up with you, you didn't buy the system, you leased it — without knowing when the rent goes up.
The second half of the answer is handover, and most proposals never mention it. In a healthy advisory relationship at least two people on your team end up able to run the system day to day, and that comes from weeks worked side by side rather than a training deck. The consultant's aim is not to make themselves indispensable but to leave you running without them.
Listen for this: does the proposal carry a line for handover and training? A handover nobody discussed never happens.
In closing: the one thing all twelve questions measure
All twelve questions measure the same distinction: is the party across the table installing a tool, or taking responsibility for an outcome? The tool-installer finishes on delivery day; the one accountable for the outcome starts on delivery day. On a price list the two sit next to each other; in what you get back for the invoice they land far apart.
The number of AI companies in Türkiye has grown fast, and that is good news — five years ago there weren't enough providers to ask these questions of. The bad news is that decks multiplied at the same rate. The way to tell them apart in the first meeting isn't to ask for a better presentation but to ask a harder question.
There is a test you can run today, and it takes ten minutes. Open the most recent proposal on your desk and look for one thing: does it name a baseline to be measured before the project starts? If it doesn't, that document commits to an installation rather than an outcome — and the gap between those two documents decides who turns out to be right at the end of the project.
We wrote the marketing-side equivalent of this discipline in eight questions to ask an agency; different questions, same logic. Ask us these twelve too: the answers to the second and the eighth already sit inside this article, with their numbers attached.