Selin, transformation manager at a cable manufacturer, walked out of a board meeting carrying a single sentence: let's do something with AI and see the result within three months. On her desk sat two vendor decks, a forty-page roadmap draft from a strategy consultant, and a question nobody had answered — what happens first thing on Monday morning?
I invented Selin for this article; I did not invent the question on her desk. We hear it from a different company every month, and the answer is neither buying a platform nor commissioning a six-month strategy exercise. AI transformation starts on one process, with a success measure written in advance and a ninety-day calendar. This article opens that calendar up block by block: which process to choose, how to take a baseline, when to stop a pilot, which lines make up the budget, and how the call gets made on day ninety.
I set out what AI transformation is, along with its four-stage roadmap, in what AI transformation is; how to test a consultant in the first meeting is collected in 12 questions to ask an AI consultant. This piece fills the gap between the two: what gets done after the decision is taken, before the contract is signed, and while the pilot is running.
Where does AI transformation actually start?
AI transformation starts with a ninety-day pilot on a single process — the one that repeats most and keeps a record. The starting point is not a technology, a department or a strategy document; it is a piece of work with a clear owner, a value you can measure today and output someone can check the same day. A strategy document is not wrong, only out of order: a roadmap written before the first pilot has produced any data is a calendar built on assumptions.
Few companies in Türkiye have made that start. According to the Artificial Intelligence Statistics bulletin that TÜİK, the Turkish Statistical Institute, published for the first time on 1 October 2025, 7.5% of enterprises with 10 or more employees use any AI technology at all, up from 2.7% in 2021. Eurostat's figure for the same size band, released on 11 December 2025, puts the EU average at 20.0%. The gap is not a technology gap but a starting gap, and closing it does not require anyone to build a large programme first.
The practical question — how do you actually get an AI project started — is answered by three decisions: which process, who owns it, which number has to move. Talking to vendors before those three are on paper means negotiating the price before the scope. The framework below fits those three decisions into the first fifteen days and everything else into the remaining seventy-five.
Which blocks make up a 90-day AI pilot project?
The ninety days split into four blocks: selection and ownership (days 0-15), baseline and data preparation (days 16-30), build and field test (days 31-72), then evaluation and decision (days 73-90). Each block produces the input for the next, and none opens before the previous one closes. The calendar is how we scope the work rather than an industry norm; the order of the blocks, though, is not up for negotiation.
In the AI transformation piece I scoped a pilot at eight to twelve weeks. The ninety-day frame lays that range onto a calendar: the roughly eleven weeks from day sixteen to day ninety are the pilot itself, while the first fifteen days are the selection work that makes sure the pilot opens on the right process.
Days 0-15 — Selection and ownership
The first two weeks produce a single page: the chosen process, a named process owner, an executive sponsor and the one number expected to move. The candidate list never runs past two or three processes; each one is scored against the criteria below and the highest scorer wins. If even the candidates are unclear, this block stretches to three weeks with our Digital Transformation Audit package, which ranks three to five pilot candidates from data before the calendar picks up again.
Days 16-30 — Baseline and data preparation
Two jobs run in parallel in the second block. The first is the baseline: the chosen process's current value, recorded over a two-week window on real work. The second is data preparation: where the process's record lives, which of its fields are empty and who has to grant access. The block closes when the process owner and the sponsor sign the success and stopping criteria; nobody moves on to the build with unsigned criteria.
Days 31-72 — Build and field test
The third block matches the six weeks of our AI Pilot package. In the first four weeks the use case is validated, the data inventory and quality check are completed, the model is chosen, and the prototype and interface are built and wired into the existing system; in the final two weeks real users work with it on real work. The field test is not simulated, because what gets measured is not laboratory accuracy but the unit time of the process. Scope is frozen in this block: a new idea, a second use case or a neighbouring department goes onto the list for the next pilot.
Days 73-90 — Evaluation and decision
In the final block the baseline is repeated with the same method, the difference is calculated, and the pilot's running cost — model usage fees, infrastructure, human review — is turned into an annual figure. Day ninety is a meeting with a single agenda item: scale, extend, fix and rerun, or stop. The pilot report reaches the sponsor's desk at least a week before that meeting; the decision comes from a report read in advance, not one opened for the first time in the room.
Which process should the pilot run on?
The process to pick for a pilot is one that runs hundreds of times a month, leaves a digital record, fails in ways someone spots the same day, has a clear owner, and can show its improvement within ninety days. The first three criteria are the three filters from the AI transformation piece; pilot selection adds two more, because not every piece of work that suits AI makes a good first pilot.
Scoring the five criteria one by one turns the argument over the candidate list into a calculation:
- Repetition: how many times a month does the process run? Work done rarely will not produce a meaningful sample during the field test.
- Record: does the process's history sit in a digital trace — an email, an ERP row, a form, a call log? Without one, the second block turns into data collection and the pilot into a record-keeping project.
- Visible error: does someone catch a wrong output the same day? An error nobody catches cannot be measured in the field test, and in production it multiplies quietly.
- Ownership: is there a named process owner who will stand behind the result? A pilot owned by IT finds no backing in the unit that actually runs the process.
- Ninety-day fit: if the model gets it wrong, can the result be undone, and can it be wired into the existing system within weeks? Decisions carrying workplace safety, product conformity or legal liability are never the subject of a first pilot.
The logic of the choice often runs against instinct. The most visible process — the bottleneck the managing director mentions in every meeting — is usually not the best first pilot, because most visible processes involve approval, authority or negotiation, and those are management's job rather than AI's. A good first pilot is often dull: preparing quotes, entering orders, searching technical documents, classifying incoming requests. The advantage of dull work is that it repeats often, its errors show, and its result is beyond argument.
Meccanotecnica Umbra is what that logic looks like in the field. The bottleneck was quoting: an engineer buying equipment could not work out what suited their plant without expert support, and requests arrived by phone and email to be handled by hand. The process repeated often, left a trace in the email traffic, and could be measured on every request. The work was a full 22-week project rather than a pilot, yet its result came from the accuracy of the choice — once the AI technical advisor and quote portal went live, quote requests rose tenfold and the time between request and response fell by ninety percent.
How do you take the baseline before the pilot?
Taking a baseline means recording the chosen process's current value over a two-week window, on real work, with a method that will be repeated exactly after the pilot. Three numbers are enough: time per unit, the error or correction rate, and the number of items waiting in the queue. An estimate, a recollection or a figure from a departmental presentation does not count as a baseline; the only thing that counts is the record kept across the window.
The method has three rules. First, the measuring is done not by the person doing the work but by a second person the process owner assigns; people measuring their own work measure optimistically. Second, the window falls in an ordinary period — the run-up to a public holiday, year-end closing or a trade-fair week does not count as a baseline. Third, the method is written down: on day ninety the same definition is measured over the same length of time, by the same person where possible. Change the method and the difference you see belongs to the method, not the pilot.
The baseline has a second job: sometimes it makes the pilot unnecessary. If the measurement shows that most of the unit time goes into an approval queue, or into moving data between two systems by hand, the right move is not AI but a fix on the digital transformation or process design side. That is not a failed pilot; it is a cheap lesson learned in two weeks.
How do you write success and stopping criteria?
Success and stopping criteria are written before the build starts, using the baseline figure, as two thresholds: a continue threshold and a stop threshold. The continue threshold is the smallest improvement the pilot has to reach in order to scale; the stop threshold is the value below which the old way of running the process is cheaper. The space between them is the fix-and-rerun zone, and a pilot landing there is allowed a single extra round.
The thresholds come from a calculation, not a guess. The continue threshold is the improvement that covers the post-pilot annual running cost — model usage fees, infrastructure, integration maintenance, human review; any gain below that, however impressive it looks, does not pay the invoice. The detail of that calculation is a return-on-investment topic in its own right; what matters at the pilot stage is that the threshold is a cost equivalent, not a percentage someone hoped for.
Alongside speed, the criteria carry quality and usage. If unit time falls while the error rate climbs, the pilot has not succeeded; what sped up is the mistake. If the field-test users are not really using the system, the prototype's accuracy says nothing about the scaling decision. That is why the criteria sheet is a single page with four lines:
- Speed: below which threshold, relative to the baseline, must time per unit fall?
- Quality: above which threshold must the error or correction rate never rise?
- Usage: how much of the field-test users' work has to run through the system?
- Stop: which figure, if missed, closes the pilot on day ninety?
How ready does the data need to be for a pilot?
The data does not have to be perfect for a pilot; it is enough to have a record that represents the chosen process's recent history, can be accessed, and has a known owner who keeps it up to date. Cleaning the whole company's data is not the pilot's business; a pilot only concerns itself with the record of its own process. If there is no record at all, the right step is not a pilot but a digitisation step that creates that record first.
Data is where pilots fall over most often, and that is not only our observation. In a statement dated 26 February 2025, Gartner predicted that through 2026 organisations will abandon 60% of AI projects unsupported by AI-ready data; according to the survey in the same statement, 63% of 248 data management leaders said their organisation either lacks the right data management practices for AI or is unsure whether it has them. That is why the framework's second block is given over to data preparation.
Data preparation asks four questions: where the record sits, how far back it goes, which of its fields are empty and who grants access to it. The fourth usually takes longest. The technical work can finish within days while approval for ERP or CRM access takes weeks, so the access request goes out on the last day of the first block rather than the first day of the second.
Processes involving personal data add one more question: in which country, and on whose servers, will the data be processed? Under Turkish data protection law (KVKK) the company itself is the data controller, and no customer or employee data enters the field test until that split is written into the pilot contract. For the calendar, the point is that this approval also belongs inside the second block; a data approval sought after the field test has started stops the pilot, not just the schedule.
Who belongs on the pilot team, and who should own it?
A pilot team has five roles: the executive sponsor, the process owner, the technical counterpart, the field users and the external team. The owner is always the process owner — not IT and not an outside consultant. That is the person who will defend the result, recommend the day-ninety decision, and run the new process if the pilot scales; if nobody is named for that role, the pilot does not start.
- Executive sponsor: owns the budget and the day-ninety decision. Attends not the weekly meetings but three decision points: process selection on day fifteen, signing the criteria on day thirty, and the decision on day ninety.
- Process owner: today's manager of the chosen process. Assigns the baseline measurement, runs the field test and recommends the decision.
- Technical counterpart: opens data and system access, and is the in-house person responsible for the connection to ERP, CRM or production software.
- Field users: the operators, sales engineers or customer representatives who use the prototype on real work for two weeks. Their feedback is part of the measurement, not a courtesy.
- External team: takes on the build, the model choice and the pilot report. It does not make the decision; it produces the number the decision rests on.
The team question carries extra weight in Türkiye. In the same TÜİK bulletin, among enterprises that are considering AI but not yet using it, the most common reason given is a lack of relevant expertise in the business, at 74.2%; it is followed by costs being too high, at 67.4%, and legal uncertainty over who is liable if AI causes harm, at 62.4%. A well-built pilot answers all three barriers on a small scale: it brings expertise in from outside and leaves it with your team, ties the cost to a fixed scope, and puts responsibility in writing from the first day.
Handover is an output of the pilot itself. By the end of the field test your team needs to be able to run the system without the external team, recognise when it goes wrong and step in. That comes not from a training deck but from six weeks at the same table; if the knowledge walks out of the door with the external team when the pilot ends, the company has bought a dependency rather than a system.
Which lines make up the budget for an AI pilot project?
The budget for an AI pilot project has six lines: the external team and build, internal team time, data preparation, integration, model usage and infrastructure fees, and the post-pilot running cost. The first line is usually fixed and visible in the proposal; the other five vary and most proposals never mention them. The honest answer to what AI costs is the sum of all six.
Let me use our own pricing as the example, because it shows how the lines ought to separate. Our AI Pilot package runs six weeks at a fixed €15,000: it covers use case selection, the data inventory and quality check, model selection, prototype and interface development, and a two-week field test with real users, and the source code is handed over in full ownership. Model usage fees, cloud infrastructure and tool licences sit outside the price because they depend on consumption; they are not fixed, and they appear in the production roadmap as estimates. Where the candidate process is unclear, the Digital Transformation Audit that comes first runs three weeks at €5,500.
The largest line missing from any proposal is internal team time. The process owner running the baseline, the technical counterpart opening access, and field users working with a new system for two weeks are all real costs — as is the process running slower than usual in the first days of the field test. A company that leaves this line out of the budget thinks the pilot is cheaper than it is, and gets its surprise on the second one.
The last line, post-pilot running cost, is the denominator of the scaling decision. Model usage fees, infrastructure, integration maintenance and human review are permanent costs, and they arrive at the day-ninety meeting as an annual figure. A company that gets the pilot approved on the build price alone ends up making the decision again when the first usage invoice lands.
What does a 90-day pilot mean for an SME?
For an SME, a 90-day pilot means trying AI not as an investment programme but as a capacity test on one process. The SME's advantage is decision speed and a short chain of command; its disadvantage is a process owner who is usually running three jobs at once. That is why the most important decision in an SME pilot is not the technology but how much of the process owner's time is protected for ninety days.
TÜİK's figures show the SME picture plainly: AI use stands at 6.6% among enterprises with 10-49 employees and 9.6% among those with 50-249. The rate is low, but the reading is not a handicap: an SME starting today sets off from a point where the great majority of its same-sized competitors have not yet begun.
An SME pilot follows three rules. First, off-the-shelf models and existing software: training a custom model is rarely needed at this scale, and the gain usually comes from wiring a ready-made model correctly into the accounting, order or e-commerce system. Second, automation first: whatever part of the process can be solved by writing a rule is solved on the business automation side, and the model only steps in where no rule can be written. Third, a short record: an SME's history is shorter than a large company's, so the pilot is built on work that reads, classifies or drafts text rather than forecasting from a long past.
Typical SME candidates are dull, frequent tasks: keying orders that arrive by email into the system, drafting quotes, classifying customer questions and routing them to the right person, matching supplier invoices to accounting codes. In every one of them unit time can be measured, errors show the same day, and the process owner is already clear — often the founder. The big advantage of a founder-owned pilot is that the day-ninety decision can be made at the same table, on the same day.
What does a pilot change for manufacturers and exporters?
For a manufacturer or exporter, the first pilot often changes the sales and technical support line before it touches the production line. AI in manufacturing brings quality control and maintenance planning to mind first; both can be the subject of a pilot, but only where sensor data, labelled images or a fault history is already on record. Without that record, half of the ninety days goes on data collection and the pilot no longer fits its own calendar.
In AI for manufacturing, the pilots that fit most comfortably into ninety days are desk processes: preparing quotes, searching technical documents and catalogues, confirming orders, handling customer correspondence across languages. At a manufacturer selling technical products, the engineer who prepares quotes is often the only copy of knowledge written down nowhere. To the extent the pilot captures that knowledge as rules and examples, it buys continuity as well as speed; when the senior engineer retires, the quoting process does not leave with them.
For an exporter, language is an extra lever. At Meccanotecnica Umbra the search and content architecture was built in four languages — Turkish, English, Arabic and Russian — and reached adjacent markets as well. Answering a foreign buyer's technical question in their own language on the same day is one of the shortest routes into a new market without growing the sales team, and it is a process a pilot can measure within ninety days: requests per language, response time, and the share of requests that become quotes.
An exporter's second question is compliance. The EU AI Act entered into force on 1 August 2024 and applies in stages; its provisions on prohibited practices and AI literacy have applied since 2 February 2025. The Act also covers providers and deployers established outside the EU where the output of their AI system is used in the Union. If you are building a system that produces quotes or technical recommendations for a European buyer, writing down what the system does and where human approval sits on the pilot's scope sheet is cheaper than a compliance exercise after the fact. The legal assessment is your legal adviser's job; the pilot's job is to prepare the information that assessment needs from the start.
Why is a pilot run differently in a large company?
In a large company the problem is not starting a pilot but finishing one. According to TÜİK, 24.1% of enterprises with 250 or more employees use AI technology — so in roughly a quarter of large companies AI is already running somewhere. The risk here is a growing number of pilots that run unaware of one another, none of them measured and none of them ever closed.
The problem is not unique to Türkiye. In a statement dated 29 July 2024, Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value. All four map onto a block of the ninety-day framework, and all four can be put in writing before the pilot starts.
An enterprise AI pilot differs in three ways. First, the approval calendar: an information security review, the procurement process and data access approval can take longer than the technical build, so in a large company the first block starts all three approvals while the candidates are still being scored. Second, portfolio discipline: no second pilot opens in the same unit before the first one's result has been read, and every pilot has its own process owner and its own stop threshold. Third, the rollout path: before the pilot starts, it is written down which units it will spread to, and in what order, if it succeeds; a pilot without a rollout plan stays one team's tool even when it works.
A large company's advantage is data: most of its processes have lived in ERP, CRM and call-centre records for years. Its disadvantage is that the owner of that data is rarely the owner of the process. Bringing the data-owning unit onto the team from the first block protects the second block's schedule. Tying pilots to a single measurement frame and basing the continue, fix or stop decision on measurement is the method behind our AI consulting work.
How is the scaling decision made when the pilot ends?
The scaling decision is made on day ninety, against the criteria signed on day thirty, as one of four options: scale, extend, fix and rerun, or stop. No new criteria are written in the decision meeting. Changing the criteria at the moment of decision is the sign that the pilot has stopped being a measuring instrument and become a tool of persuasion.
- Scale: the continue threshold was passed, the annual running cost is covered, and the process owner is ready to run the new process. The system moves into production and the same process spreads across the rest of the company.
- Extend: the result is above the continue threshold and the same set-up can be carried over to a neighbouring process. Because the build cost has been paid once, the second process comes in cheaper than the first.
- Fix and rerun: the result sits between the two thresholds and it is clear what fell short. A single extra round opens with a rewritten scope; a pilot that lands in the same zone twice is stopped.
- Stop: the result fell below the stop threshold. What was learned is written down, the source code stays with the company, and the next candidate is taken from the second line of the list.
All four options are legitimate outcomes of a well-run pilot. The failed pilot is not the one that gets stopped but the one that ends undecided: neither scaled nor closed, an experiment that runs on quietly through budget cycle after budget cycle. The only reason the day-ninety decision is mandatory is to make that outcome impossible.
Three documents come to the decision meeting: the pilot report setting the baseline beside the post-pilot measurement, the annual running cost, and a roadmap showing the technical steps, estimated budget and timeline for moving to production if the pilot scales. If one of the three is missing the decision is not shelved; the missing document is completed and the meeting slips by a week at most.
What are the most common mistakes in a 90-day pilot?
The most common mistake is opening the pilot to an entire department, the second is skipping the baseline, and the third is changing scope during the field test. All three share a root: treating the pilot as a showcase rather than a learning experiment. Here are seven mistakes, with the day on the calendar where each one shows up:
- Department scope (days 0-15): "let's transform the sales department with AI" is a programme, not a pilot. A pilot opens on one process.
- Starting unmeasured (days 16-30): a pilot built without a recorded baseline cannot prove its result, however good it is.
- Unsigned criteria (day 30): if the success and stop thresholds were not written by day thirty, day ninety produces an argument instead of a decision.
- Scope creep (days 31-72): every new request added during the field test resets the measurement; new ideas go onto the list for the next pilot.
- The ownerless pilot (every day): a pilot carried on IT's back gets no defence from the unit that runs the process.
- The hidden cost (days 73-90): if internal team time and running cost were never budgeted, the scaling decision is reversed at the first invoice.
- Leaping from one success to a programme (after day 90): opening five new pilots at once after the first success scatters the discipline the first one taught. The second pilot opens using the first one's criteria sheet as its template.
The first step to take on Monday morning
Back to Selin's question: what happens first thing on Monday morning? A page gets opened and three columns get written — two or three candidate processes, the current owner of each, and the one number that has to move in each. The page fills in half an hour; if it does not, the problem is not AI but processes nobody owns, and the first job is to establish that ownership.
If the page is full, the first fifteen of the ninety days have begun. Score the candidates against the five criteria, set the baseline window for the top scorer two weeks out, and send the access requests today. AI transformation starts not with a big decision but with a small, measured one; you can run the ninety days with your own team or with outside support. If you choose the second route, you will find the method in our consulting scope and the evidence in our case records.