Two GEO proposals sat on Selin's desk. The first guaranteed that within three months her company would be "the first brand named in your category" on ChatGPT, with an AI visibility score reported every month; it cost half as much as the other. The second guaranteed nothing. Instead it listed, line by line, which 10 questions would be asked each month, to which three engines, under which rules. Selin picked the first.
By the end of the fourth month she had an installed llms.txt file, a score chart that crept up a little every month and three screenshots from ChatGPT. All three showed her brand being named, because all three prompts contained the brand's name. When she removed the name and asked the same question herself, the answer still listed her competitors. I invented Selin for this article; she is the only invented thing in it.
Choosing a GEO agency is not choosing a content supplier; it is choosing a measurement partner. A wrong choice costs more than the budget: a badly built measurement makes a programme that doesn't work look successful and one that does work look like a failure, and the next year's decisions get built on that error. Below are six evaluation criteria, ten questions for the first meeting, three promises that should end the conversation, and how the price is put together. The concept itself is explained in the guide to standing out in AI search; how we work as a GEO agency is written out on the service page.
What does a GEO agency sell, and what doesn't it?
A GEO agency does not sell you rankings; it sells the work that raises the chance of your brand being cited when ChatGPT, Gemini, Perplexity and Google AI Overviews answer a question. That chance is built in four layers: technical groundwork that lets an engine read your site, content that still stands when a passage is lifted out, a brand described in the same sentence everywhere, and a measurement that tracks all of it every month.
What it doesn't sell should be just as clear. A good GEO agency does not claim to control how engines pick sources, because no engine discloses that choice. It does not sell ad space; advertising on ChatGPT is a separate job and, as I set out in the piece on ChatGPT ads, it sits on top of organic visibility rather than replacing it. Nor does it sell a ranking report: a report on Google rankings can be useful, but it is SEO's report.
The boundary matters because GEO is built on top of SEO without replacing it. A generative engine cannot read a page that isn't indexed, loads slowly or hides its content behind JavaScript. A good agency tells you in the first meeting whether your site carries that foundation and, if it doesn't, proposes fixing that first. A weak agency presents a content calendar without looking at the foundation.
What does the agency measure success by?
This is the first criterion, because the value of the other five depends on it: what does the agency measure GEO's outcome by? The right answer is not rankings but mentions — whether your brand appears in the answer to specific questions, in which sentence, and which of your pages is cited.
Let me use our own method as the example, because you should hear the same level of detail from the other side of the table. Every month the same 10 prompts go to ChatGPT, Gemini and Perplexity: four category questions, three constraint questions, three comparison questions. The prompts are written once and never change; none contains the brand name, the questions are asked in a clean session and the answers are recorded by hand. Thirty queries a month, three fields for each: did we appear, in which sentence, which page was cited. The reasoning behind the method is laid out in the measurement section of the GEO guide.
The quickest way to tell whether a measurement is honest is to ask the agency for its own number. Ours is this: in the baseline round on 30 August 2026 we appeared in none of the 30 queries; in the 1 September 2026 round Perplexity cited our article on choosing a CRO agency for a question on that subject, and the number became 1/30. A small number, but we know where it came from and which page was cited for which question. A report that says "your visibility is up forty percent" but cannot tell you in the answer to which question is an impression, not a measurement.
Where does the agency start with your technical groundwork?
The second criterion is the groundwork: before talking about content, does the agency check whether your site is open to AI engines at all? The groundwork reads on five signals — the permission given to AI crawlers in robots.txt, the llms.txt file, structured data, language signals and the share of headings written as questions.
You can measure those five yourself before the meeting. The GEO Visibility Checker is free and returns a score out of a hundred within seconds; we explained how the points split across the five signals in the article on Türkiye's first GEO audit tool. Note the score and the two lowest items, then ask the agency in the meeting: what will you fix on my site in the first week? If the answer matches the findings already in your hands, the agency has genuinely looked at your site.
Ask one more concrete question: which crawlers should robots.txt allow explicitly? A good answer names them — GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended and so on — and knows the trap as well: a crawler named in its own block never reads the wildcard block, so unless the restriction list is repeated in every block, admin paths open up too. That is why our own site lists ten AI crawlers by name and repeats the same restrictions in each block; the detail is in the llms.txt guide. A weak answer is one sentence: "The bots can already reach your site."
Who writes the content, and on what basis?
The third criterion is content structure. Generative engines don't waste time on long introductions; they take the first clear answer under a heading. So the question to ask an agency isn't "how many articles will you produce?" but "how will you choose the headings?"
A good answer names its sources: the Search Console query report, the first minutes of sales calls and support logs. A heading comes from a sentence a customer actually wrote; the paragraph under it still makes sense when lifted out of context and carries its figure and its constraint inside the sentence. A weak answer is a volume: a hundred AI-generated pages a month. If none of those hundred pages holds a citable paragraph, volume only means more pages to crawl.
One test: ask the agency for a page it has written before, cut out a random paragraph and read it out of context. Is its subject clear, does its claim stand in a single sentence, is its figure inside it? If you cannot say yes to all three, that kind of paragraph won't be cited on your site either.
Does it fix how your brand is described beyond your own site?
The fourth criterion rarely appears in proposals at all: entity consistency. When a model describes a brand it doesn't read only your site; it also reads your LinkedIn page, your Google Business Profile, sector directories and whatever else has been written about you. If the site says one thing, LinkedIn another and a directory a third, the model can't decide which to write, so it writes the competitor it is sure about.
Ask the agency: who will write our one-sentence brand description, on which profiles will it be updated, and how will consistency be checked? This is editorial rather than technical work, and it is usually work nobody inside the company owns. A description without an owner scatters again within six months. A good agency draws up the list of profiles, makes the first round of corrections and leaves the record with you.
What should you see in the report?
The fifth criterion is the report, and it comes down to one question: are the zero months in the report too? A document that carries only screenshots of the answers you appeared in is a shop window, not a report.
A transparent GEO report carries four things: how many queries on the fixed question list you appeared in, which sentence named you in each of those answers, which of your pages was cited, and which competitors appeared in the same answers. One further distinction matters: is the brand name in the text of the answer, or only in the source list? The first citation in our own measurement was of the second kind — our page was cited but our name didn't appear in the answer text. Those are different gains, and the report should record them separately.
The record of which page was cited is the least-read and most valuable line in the report. Whichever page was cited for whichever question, the next piece of content is written in that page's structure. An agency that can't read its report back from that line repeats the same mistake month after month.
What evidence should you ask for?
The sixth criterion is evidence. A wall of logos is a client list, not evidence. What you should ask for comes together: the starting value, the ending value and the time between them — and in GEO a fourth, the questions and engines the measurement was based on.
Here are two from our side, held to the same standard. At SIM Printing Suppliers the site was rebuilt as a five-language Next.js application and the content programme was written with SEO and GEO in mind together; in six months organic traffic grew 15× and visibility in AI engines went from zero to 40,000. At İstanbul Ortez Protez the content was written with a question-and-answer structure and self-contained passages that AI engines can cite; reaching the organic top 3 for priority keywords took fifteen months. The real information in the second is not the figure but the duration: where trust is expensive, results come in months, and an agency that says so up front is telling you the truth.
Then ask for one more thing: a piece of work that moved slowly or missed its expected result. GEO is a young field; an agency whose history shows no deviation at all has either done very little work or reads its reports selectively.
Which ten questions do you ask in the first meeting?
The point of the questions is not to put the agency through an exam but to hear its method. Pay as much attention to which question gets a vague answer as to the answers themselves.
- Which questions, which engines and what frequency will you measure success by?
- Will our brand name appear in the prompts?
- Which technical issues will you fix on our site in the first week?
- Which sources will you draw the headings and questions from?
- On which external profiles will you correct our brand description?
- How will zero months, and answers where we appear only in the source list, look in the report?
- Can you show the start and end values of similar-scale work, together with the measurement method?
- Which items make up the price, and which are excluded?
- When the work ends, who keeps the question list, the measurement records and the access?
- Is there a situation in which you should turn this work down?
The second question yields the most in the least time. A yes means most of the success to be reported will come from the question itself: a model hands back a brand when the brand's name is in the question. The tenth shows whether the agency knows its own scope. An agency that can tell a company with an unindexed site, or no content at all, that it needs a different job first knows where its boundaries are.
That is where the GEO-specific layer of the questions ends. For the general layer of an agency relationship — how data is used, whether channels cohere, how the team reacts in a crisis — the 8 questions to ask an agency before you sign applies the same discipline to a wider frame. If what you need is a consultant to put AI into your business processes, the right list is the 12 questions to ask when choosing an AI consultant.
Which three promises mean the meeting is over?
Three promises make the rest of the meeting unnecessary the moment you hear them. All three hide the same thing: selling unmeasured work as if it had been measured.
The first is a ranking or mention guarantee. "In three months you'll be the first brand ChatGPT names" cannot honestly be said, because no engine discloses how it picks sources and the same question can produce two different answers on the same day. A guaranteeing proposal usually doesn't state which question, which engine or which session conditions it means either — so what is being guaranteed stays undefined.
The second is the proposal with no measurement method. If there is no fixed question list, engine list and record format, or if measurement comes down to a single "AI visibility score" produced by a tool, what you hold is an indicator, not a result. The score itself is not the problem; our own tool produces one. But that score measures the groundwork, not the answer. An agency that reports a groundwork score as the result is selling the thermometer as the cure.
The third is the llms.txt-only model. llms.txt is a proposed file that gives language models a plain-text map of your site; it is cheap to set up, does no harm, and no engine has declared it a condition for citing a source. A proposal that installs the file and calls the job done is selling the footnote. The same goes for only adding schema or only editing robots.txt: all three are part of the groundwork, and none of them on its own is GEO.
There is a fourth marker, more dangerous because it is quieter: a price given without looking at your site. A price written before the state of the groundwork, the number of pages and the number of languages are known tells you the scope will later be either narrowed or expanded.
How is a GEO agency's price put together?
GEO has no price list everyone agrees on, and I won't give a figure in this article; any figure I gave would belong to an average site, not yours. Instead I'm setting out the items that make up the price and what makes each one grow, because the way to compare two proposals is to look at the items rather than the total.
- Diagnosis: the five-signal audit and the baseline measurement. It grows with the number of key pages and engines to measure.
- Technical groundwork: crawler permissions, structured data, llms.txt and rendering fixes. It grows with the site's stack; on an off-the-shelf platform it may be a few settings, on a custom-built site it may be development work.
- Content: pages to rewrite and pages to create. It grows with the number of pages and languages; each language should be written to its own buyers' questions, not translated.
- Entity consistency: correcting external profiles. It grows with the number of profiles and is often the cheapest item.
- Measurement: the monthly round and report. It grows with the number of months, and after handover it can move to your team.
The question to ask when choosing a model is the same in any agency work: which agency behaviour does this pricing reward? A per-page price rewards producing many pages; an open-ended monthly fee rewards stretching the work. A healthy structure is often a mix of the two: a project fee for the diagnosis and groundwork, a monthly fee with its duration written in up front for content and measurement, and a handover at the end. Whichever model you choose, get the exclusions in writing: ad budget, replatforming, building a site from scratch and translation.
On our side this service has no fixed package; the price is written after the diagnosis, against its findings. We have listed plainly on the service page which items fall within the scope of our GEO consulting and which stay outside it.
What changes for an SME, an exporter and a large company?
The criteria stay the same; their weights change. At each of the three scales, what you should expect from an agency concentrates on a different item.
SMEs
A small business's advantage is a narrow category. Competing with the big players on a category question such as "which firms do this in Türkiye" is hard; but on constraint questions carrying a location, a budget or a technical requirement — "is there a firm in Istanbul that does both this and that" — the chance of being named can arise earlier. For an SME the right start is closing the groundwork gaps the tool shows, pulling the Google Business Profile and directory descriptions into one sentence, and writing pages that answer the five to ten questions customers ask most. Part of this the in-house team can do; what you want from an agency is measurement discipline and prioritisation, not a big content calendar.
Exporters
For an exporter the issue is language. A foreign buyer's supplier search increasingly starts as a conversation, and that conversation happens in the buyer's own language. The first question for an agency is whether it will measure separately in each target market's language: a round run with Turkish prompts won't tell you whether you appear in the answer given to a buyer asking in German. The second is hreflang and language signals; the third is whether the content will be translated or written to that language's own questions. At Meccanotecnica Umbra the SEO and GEO architecture was built in four languages (TR, EN, AR, RU) at once; at SIM Printing Suppliers the five-language structure opened the door to demand from outside Türkiye.
Large companies
A large company's problem is not being invisible but being visible wrongly. The authority and the content already exist; the risk is a model writing an old product name, a closed facility or a subsidiary's description as if it were the brand's own sentence. What you want from an agency here is governance: who owns the brand description, how legal sign-off works, how the descriptions of several subsidiaries are kept apart. The second risk is defence. As I set out in the piece on ChatGPT ads, a challenger can rent the space beneath an answer that names the big brand; holding the organic source is therefore a line item for a large company too.
Agency, consultant or in-house team?
The decision depends on where the work is stuck. If the problem is only the groundwork — a closed crawler rule, missing schema, content hidden behind JavaScript — a good developer closes it in a few weeks and no agency is needed. If the problem is the content and nobody measures monthly, you need a team that brings method and discipline from outside.
An in-house team makes sense under three conditions: someone knows the customers' questions and can write, a technical person has publishing access to the site, and someone will give roughly an hour a month to the measurement round. The third is the most often skipped; GEO work that isn't measured quietly stops after a few months.
For most companies the right answer is the hybrid: the agency builds the groundwork and the first wave of content, teaches the measurement round to your team, and then stays only for periodic review. In this model the agency's success is measured by how unnecessary it makes itself, so put the handover in the contract. The question list, the measurement records and all access should stay with you.
How does INDOLES run this work?
INDOLES runs this work in four steps and holds itself to the criteria in this article: diagnosis and baseline, technical groundwork, content structure, then measurement and handover. In the diagnosis the site is scanned on five signals and your category's 10 questions go to three engines without the brand name. No content is written before the groundwork is set. Content is written against a question map drawn from customers' own sentences, and the brand description is aligned to one sentence across external profiles. The measurement repeats with the same 30 queries every month, and at the end the round is handed over to your team.
We built the same structure on our own site first, and we report the result as it is. In Google Search Console for 22 August – 19 September 2026, our average position was 1.2 for the query "yerli geo aracı" (local GEO tool) and 2.5 for "türkçe geo aracı var mı" (is there a Turkish GEO tool). In the monthly round we are at the start of the road: we are cited in 1 of 30 queries. We put the two numbers side by side because this is exactly what you should hear from an agency: where we lead, where we don't, and what we measure it with.
The full scope, the deliverables and what stays out of scope are written out on our GEO service page.
Conclusion: the test to run before comparing proposals
Choosing a GEO agency is an audit of measurement, not a comparison of decks. An agency that measures with fixed questions, keeps the brand name out of prompts, looks at the groundwork before the content, shows its zero months in the report and itemises its price is a better investment than one offering guarantees, because it leaves you the reasoning as well as the result.
Here is the concrete test you can run today: write three questions about your category the way a customer would ask them, put them to ChatGPT, Gemini and Perplexity without your brand name, and note who appears in the answers. Then ask every agency you meet to explain how it would measure those same three questions, under the same conditions, six months from now. An agency that can describe the method knows how to measure; think twice before signing with one that says it will report a score.
Ask us the same questions. Part of our answer — the measurement method, our own numbers, what stays out of scope — is already written in this article and on the service page; put the rest side by side and compare.