Two proposals sat on Deniz's desk. The first guaranteed a thirty percent lift in conversion rate within three months, at half the price of the other. The second gave no lift figure at all; instead it listed, line by line, which measurements would be repaired in the first six weeks. Deniz picked the first.
By the end of the fourth month what remained was a fourteen-slide deck, three "winning" tests and a conversion rate that had barely moved. None of the tests had run longer than two weeks. None came with a calculation of how many visitors they were based on. And the thirty percent in the contract was never tied to a metric, a period or a baseline. I invented Deniz for this article; that is the only invented thing in it.
Choosing a CRO agency is not choosing a supplier; it is choosing a measurement partner. A wrong choice costs more than the budget — it corrupts the next year of decisions, because a badly measured test leaves false information behind dressed as fact. Below are five evaluation criteria, eight questions for the first meeting and three promises that should end the conversation. How the work itself runs is set out on our conversion optimisation service page, and the definition of the term sits in what CRO is.
What does a CRO agency sell, and what doesn't it?
A CRO agency does not sell traffic; it sells the method for pulling more sales or enquiries out of the traffic you already have. If that distinction isn't settled before signing, both sides spend the engagement waiting for the wrong work.
A conversion optimisation agency delivers three things: a numbered account of where visitors drop off, a prioritised set of hypotheses for fixing those points, and a tested result for each hypothesis. What it does not sell is media buying, a new source of visitors, or product-market fit.
The boundary matters, because CRO consultancy applied to the wrong problem produces the most expensive mistake of all. On a site with a few hundred visitors a month, an A/B test measures noise. On a site whose pages take seconds to load, the problem is infrastructure rather than persuasion. A good agency says this in the first meeting and points you elsewhere; a weak one signs the contract.
How does the agency treat your measurement stack?
This is the first criterion: does the agency validate your existing measurement before it runs a single test? A badly defined conversion event invalidates every test built on top of it — and nobody notices, because the report still arrives full.
Ask this in the meeting: "What will you do on the measurement side in the first two weeks?" A good answer is concrete — conversion events redefined, double counting removed, channel attribution checked, mobile and desktop flows verified separately. A weak answer is one sentence: "We'll use your current setup."
Repairing measurement first has a visible cost and a delayed return — but it does return. In the SOYLU AVM case, pixels and conversion tracking were rebuilt from scratch before any campaign went live; what made the $1.5M recorded in the first 6 days and the 150% rise in total traffic readable at all was exactly that order. Measurement is not a later line item; it is the first one.
Hypothesis discipline, or a long list of ideas?
The second criterion is whether the agency turns ideas into hypotheses. A hypothesis has three parts: what we change, why we change it, and how much movement we expect in which metric.
"Let's make the button orange" is an idea. "The add-to-cart button falls below the fold on mobile, so its click rate is low; pinning it should lift mobile add-to-cart rate by at least 10% relative" is a hypothesis. The difference isn't style: the second can be proven wrong, the first cannot. A sentence that can be wrong is measurable; one that cannot is only arguable.
Ask the agency for five anonymised lines from a current client's test backlog. If each line carries an expected impact and an implementation effort, the prioritisation discipline exists. If all you see are change titles, what you have is a to-do list rather than a method. And if losing tests are still on the list, you are at the right table.
Is the agency honest about test duration?
The third criterion is the one most often skipped: does the agency calculate how long a test needs to run before it starts? An agency that skips the sample size calculation decides when to stop a test by looking at the result — which stops it from being a measurement at all.
Put numbers on it. On a page with 20,000 visitors a month converting at 2%, separating a 10% relative lift from noise takes tens of thousands of sessions per variant — on most sites, weeks rather than days. An agency that can run that calculation in the meeting has a method. One that says "we'll have results in a few days" is selling an impression, not a statistic.
The second marker of honesty is the early-stopping policy. A good agency writes down the decision threshold and the duration before the test starts, then holds to them. If a test is closed on day three because it looks good, the winning variant is usually noise — and that noise then ships to the live site permanently.
How transparent is the reporting?
The fourth criterion reduces to one question: are the losing tests in the report? A document that carries only winners is a presentation, not a report.
A transparent CRO report carries four things: each test's hypothesis, the duration and sample collected, the outcome (won, lost, no difference), and the decision that outcome changed. Tests that come back "no difference" are not failures but inventory; knowing which idea didn't work protects next quarter's budget.
Who presents the report is a signal too. If the monthly meeting is run by an account manager rather than the analyst, a layer is filtering the information. Ask this in the same meeting: when the engagement ends, who keeps the testing setup, the conversion definitions and the archive of past tests? If the answer is "we do", what you bought is dependency rather than optimisation.
Which pricing model is right for you?
The fifth criterion is not the price but the shape of the price. Four models are in circulation, and each rewards a different agency behaviour.
- Monthly retainer: a fixed fee and a continuous testing cycle. Sound for high-traffic sites that change often — but the number of tests per month and the average test duration belong in the contract from day one.
- Project-based: fixed scope, fixed duration, fixed price. The right model for a first audit and a measurement repair; on its own it does not carry continuous optimisation.
- Performance-based: part of the fee is tied to the lift. It sounds fair and has exactly one condition — the definition of the lift, the source of measurement and the baseline period must be written into the contract. Unwritten, the model works in the agency's favour.
- Tool licence plus setup: the agency is essentially a reseller for a testing tool. The licence may well be necessary, but on its own what you have bought is software, not method.
One question settles the choice: which agency behaviour does this pricing reward? A model that rewards test volume produces many short tests; a model that rewards learning produces fewer long ones. The second looks slower and moves faster.
Which eight questions do you ask in the first meeting?
The eight questions are a listening device, not a filter. The real information isn't in the answer given — it is in where the agency hesitates.
- What exactly will you do on the measurement side in the first two weeks?
- How do you calculate the sample size a test needs?
- How do you decide when to stop a test, and is that decision written down in advance?
- How many of your tests lost last quarter, and what did you learn from which one?
- What have you done at a similar scale in my sector, and can I see the numbers?
- What do losing tests look like in the report?
- When the engagement ends, what is left behind and who holds the setup?
- Is there a situation in which you should turn this work down?
The eighth question yields the most. "We work with any site" is usually true and means nothing on its own. An agency that can name the case in which it would turn you away knows its own scope; one that cannot will learn that scope on your budget.
These cover the CRO-specific layer of the choice. For the general layer — how data is used, whether channels cohere, how the team reacts in a crisis — the eight questions to ask an agency before you sign applies the same discipline to a wider frame.
Which three promises should end the meeting?
Three promises make the rest of the meeting unnecessary the moment you hear them. All three hide the same habit: talking without measuring.
The first is the guaranteed lift. "We guarantee a 30% conversion increase in three months" cannot honestly be said: the outcome depends on the size of the existing problems, on traffic and on category. A range is honesty; a guarantee is a sale — and guaranteeing proposals usually leave the definition of the lift blank.
The second is the agency that never calculates a sample size. If "how long will you run the test?" gets "until the result is clear" instead of a duration, you are buying an impression, not a measurement. The third is the licence-only model: heatmaps, session recordings and testing tools are instruments of the work, never the work. A proposal that ends at tool installation is missing hypothesis generation, prioritisation and decision — the things the fee is for.
There is a fourth marker, more dangerous because it is quieter. An agency that prices before it has seen a plan will later either shrink the scope or grow it. Both lead to the same place: the invoice, rather than the work, sitting at the centre of the relationship.
How do you ask for a case with numbers?
Asking for a case study is easy; asking for the right one is not. The wall of logos is a client list, not evidence.
Ask for three things together: the starting value, the ending value and the time between them. "We increased the conversion rate" is a sentence; "the conversion rate went from 1.4% to 2.1% in nine weeks, on this page" is a record. A percentage without context is the easiest number in the world to produce.
Here is one from our own side, held to the same standard: at GYMWOLVES the target was to double sales in 3 months. The data flow was repaired, the funnel rebuilt and the campaign fed with social proof shot with athletes; by the end of the third month sales were up 12×, session duration 3× and engagement 8×. The useful information there isn't the 12× — it is the order: measurement first, then the funnel, then the campaign.
Then ask for one more thing: a piece of work that failed. An agency that can walk you through a project which missed its expected result is the one that will keep telling you the truth for the next eight months.
Agency, a CRO specialist, or an in-house team?
The decision depends on where the loss sits. If the loss is on one page or at one step, a CRO specialist is enough; if it is spread across measurement, interface, content and technical infrastructure, one person cannot hold those layers at the same time.
Building in-house makes sense under three conditions: monthly traffic sustains continuous testing, the product team can ship a winning variant within two weeks, and the company does not treat a losing test as a failure. The third is the hardest, and it is usually a cultural decision rather than a budgetary one.
For most mid-sized brands the hybrid model is the right answer: over six months the agency builds the measurement, runs the first wave of tests and hands the routine to the in-house team; the outside role then narrows to hypothesis generation and review. Here the agency's success is measured by how unnecessary it makes itself — so put the handover session in the contract.
Conclusion: the one test to run before you choose
Choosing a CRO agency is an audit of method, not a comparison of decks. An agency that repairs measurement first, turns ideas into hypotheses, calculates test duration in advance and shows its losses in the report is a better investment than one offering guarantees, in every case. The first kind leaves you with the reasoning as well as the result.
Here is the concrete test you can run today: pick a single page on your own site, note its current conversion rate and monthly visitors, then ask every agency you meet the same question — "How many visitors and how many days would it take to detect a 10% relative lift on this page?" An agency that can give you a range in the meeting knows the calculation. Think twice before signing with one that says it will look into it and get back to you.
Ask us the same questions. The sequence of our method is written out on our conversion optimisation service page; put the answers side by side and compare them.