How to Choose a GEO Agency: Criteria, Questions, Red Flags

A year ago hardly anyone used the acronym GEO. Today almost every SEO agency has it on the menu. The problem is that two completely different things are sold under the same name: real work on making your brand citable in AI models, or classic SEO in new packaging.
This is not a ranking. You will not find a list of “the best GEO agencies” here, because no agency is a credible judge in its own case, and we are a party to it as well. What you will find are criteria you can use to evaluate any vendor yourself, seven questions to ask before you sign a contract, and a description of how we checked who actually shows up in AI answers today.
What a GEO agency actually does
Before you evaluate an offer, it helps to know what the work consists of. The tasks of a generative engine optimization agency come down to four areas.
Baseline measurement. A set of buying questions from your industry, asked of several models, to check whether your brand comes up in the answers at all and who gets named instead of you. Without it there is nothing to improve and nothing to compare the result against six months later.
Rebuilding content for citation. Models extract facts, definitions, and numbers from text. Content written for the click, with a long intro and the point at the end, cites poorly. Content in which each section answers one question and contains something concrete cites well. We cover the source selection mechanism in more detail in our article on how AI models choose sources.
Working on how the company is described beyond its own website. Models build a picture of a brand from many sources at once. Consistent business data, presence in industry directories, media mentions, and publications that others reference weigh more here than another page on your own site.
Citation monitoring over time. A repeatable measurement that shows whether the changes you made did anything.
Any good SEO agency can handle the first three. The fourth is what separates GEO from SEO with a new label.
What exactly happens in a GEO project
It helps to know what the workflow looks like, because that lets you judge whether the proposal in front of you is complete. Generative engine optimization usually splits into four stages.
Stage one: competitor analysis in model answers. The point is not who ranks high in Google, but who the models give as a source for questions in your industry. Those two lists can differ a lot, because traditional search engines and generative engines pick material differently. The result of this analysis sets the real baseline.
Stage two: strategy and choosing key questions. Instead of a keyword list, you get a list of questions you want to be the answer to. That is a real shift in thinking: in classic SEO you target a phrase, in GEO you target the whole situation in which a customer asks for a recommendation. A good GEO strategy also sets the order, because narrow, industry-specific questions are won faster than general ones.
Stage three: optimizing content for fact extraction. The model reads text in fragments and picks the ones it can quote without losing meaning. So what counts is self-contained sections, a definition stated up front, specific numbers with a source, and avoiding sentences that only make sense in the context of the whole article. The same content written for clicks and written for citation looks different.
Stage four: building authority beyond your own site. Models weigh how a company is described in the results they can access: industry directories, publications, and roundups prepared by someone other than you. This is the slowest part of the work, but also the most durable.
How we checked who is visible in AI answers
So that this article is not just another opinion, we describe the method we use ourselves and that we used while preparing this piece. You can repeat it for your own industry.
We take fifteen buying questions, the kind a customer looking for a vendor actually asks, for example “which SEO agency is worth recommending for a service business” or “what is GEO and who offers it.” We ask them of four models: ChatGPT, Perplexity, Gemini, and Claude, with web search turned on. For every answer we record two things: whether the brand was mentioned and whether its website was given as a source. That is sixty answers per run.
The result is a hard number, not an impression, and it can be compared month to month. Our measurements lead to one general observation: nobody is far ahead in this field yet, and anyone who claims otherwise should show the numbers from their own measurement.
The second measurement is about who gets cited instead of you. For questions about GEO, the models pointed to a handful of individual domains as sources. For questions about Google Ads and Facebook Ads, there were more than twenty cited domains. The difference matters: in GEO, competition for citations is today many times smaller than in classic advertising services. That is a window that will be much narrower a year from now.
Who is visible today: a measurement from September 1, 2026
Instead of a star ranking that could not be honestly justified, here is the result of a measurement with the method above, with a date. This is a snapshot from one of our European home markets, so treat it as a worked example of the method, not as a list of who to hire in the US. The pattern is what matters, and it is easy to repeat for your own market.
In Google’s top ten for the phrase “GEO agency” that day, three of the results were not agency websites at all but rankings of GEO agencies. That says a lot about the content format this query rewards.
For questions about GEO asked of AI models, the sources cited were largely the authors of those roundups, not the firms with the strongest SEO profiles. The model answers with what it read in the roundups, and that is the practical lesson from the whole measurement.
Both lists change over time, so treat them as a snapshot from a specific day and a starting point for your own measurement, not as a verdict on who to choose.
How to do this measurement yourself before you hire anyone
You do not have to wait for a proposal to learn your baseline. You can do this measurement in an hour and walk into the conversation with an agency knowing what you are talking about. It is also the cheapest form of verification: if a vendor later shows you a result significantly better than yours, it is worth asking where the difference comes from.
Step 1. Write ten key questions. Not phrases, but full sentences, the way a customer would ask. Three types: a question about a vendor (“which X agency is worth recommending for a service business”), a question about price (“how much does X cost in 2026”), and a question about the problem you solve. Save them in a file, because a month from now they have to be exactly the same, otherwise you cannot compare results.
Step 2. Ask them of four models with search turned on. ChatGPT, Perplexity, Gemini, and Claude. Search has to be on, because without it you are testing the model’s memory, not its access to current sources. Ask each question in a new conversation so the context of earlier answers does not affect the next ones.
Step 3. Record two things for each answer. Whether your brand was mentioned and whether your website appeared among the sources. These are two different situations: a model can recommend a company without linking to it, and the other way around. With ten questions and four models you have forty answers and two numbers that are your baseline.
Step 4. List the domains cited instead of you. This is the most practical output of the whole exercise. You get a list of companies and websites the models currently treat as authorities in your industry. Notice how many of them are not agencies but directories, rankings, and industry portals, because that hints at where it is worth securing a presence.
Step 5. Repeat in a month, with the same method. A single measurement says little, because models can be unstable and the same answer can look different two days later. Only a series of results shows a trend and lets you judge whether the optimization work changes anything.
If after this exercise your brand does not show up even once, that is no reason to panic. In most industries that is exactly what the starting point looks like today, which is why the barrier to entry is still low.
Seven questions to ask an agency before you sign
This is the practical part. The answers to these questions separate a vendor from a salesperson.
1. How will you measure the baseline and what will you show it with? Expected answer: a specific number of prompts, specific models, a report before the work starts. An evasive answer like “we will assess your AI visibility” means there is no method.
2. Which models do you monitor and how often? ChatGPT and Perplexity are the minimum, because they behave differently: Perplexity almost always cites sources, ChatGPT less often. ChatGPT alone is too narrow. A frequency below once a month does not let you separate the effect of the work from model fluctuation.
3. What will you do beyond my website? If the whole plan fits on your site, it is SEO. Citability is largely built off-site.
4. How will you separate the GEO effect from the SEO effect? A control question. The honest answer is: partly you cannot, because they are connected vessels, which is why we measure citations and traffic separately. An answer that promises full separation is false.
5. Can you show traffic coming from AI tools in my analytics? This can be done and it can be checked. We describe it in our article on how to find AI traffic in Google Analytics. An agency that cannot set this up will not be able to show you the result either.
6. What will you do if citations do not grow after three months? You are checking whether a plan B exists and whether the contract accounts for it.
7. Do you use GEO for yourselves and what are your results? The simplest test. Ask a model about a GEO agency and see whether your candidate shows up. If they cannot get themselves cited, it is worth asking why. “We are busy with clients” is acceptable once, but not in a service whose whole point is visibility.
Red flags
Guaranteed positions or citations. Nobody controls what a model generates. A guarantee in this field is either ignorance or deliberate misdirection.
Results in a few weeks. Models have to recrawl and reprocess the content. That takes time and cannot be sped up with a wire transfer.
A proposal based only on the number of articles. “Ten articles a month” is a production metric, not a result metric. In GEO what counts is citability, and that does not grow linearly with word count.
No measurement of any kind in the proposal. If the offer says nothing about how you will check the result, you are buying hope.
Selling GEO as a replacement for SEO. GEO is built on top of SEO, it does not replace it. We break this down in our article on the differences between SEO and GEO.
What to realistically expect
An honest set of expectations looks like this. During the first month the measurement and the content rebuild take place, so nothing visible happens. Between the second and third month the first mentions appear, usually for narrow, industry-specific questions, not general ones. General questions such as “the best marketing agency in the US” are the hardest, and the effect comes last there.
Traffic from AI tools is still small in absolute numbers and will stay that way for a long time. Its value is different: a person who comes to you from a model’s recommendation arrives after the decision, not in the middle of comparing offers. That is usually a much shorter path to a sales conversation than a click on a search result.
Table: what signals competence and what signals a sales pitch
If you have several proposals in front of you, run them through this table. The middle column describes an answer that indicates real work, the right one is a warning sign.
| Criterion | Signals competence | Red flag |
|---|---|---|
| Baseline measurement | A specific number of prompts and models, a report before work starts | ”We will analyze your AI visibility” with no details |
| Models monitored | At minimum ChatGPT and Perplexity, plus Gemini or Claude | ChatGPT only, or no answer |
| Measurement frequency | At least monthly, the same method every time | A one-time measurement or “at the end of the engagement” |
| Scope of work | Content plus presence beyond your own site | Only publications on your website |
| Reporting | Citations and mentions separate from organic traffic | One combined “visibility” chart |
| Analytics | They can show traffic from AI tools in GA4 | The topic is skipped or brushed off |
| Promises | Time ranges and a scenario for when it does not work | Guaranteed citations or positions |
| Their own result | They share their own numbers, including weak ones | Generalities about “numerous successes” |
One note on using this table. The point is not for the vendor to tick all eight boxes, because the field is young and hardly anyone has a full set today. The point is for any gap to be conscious and explained, not hidden.
What a good GEO report looks like
This is a good control question at the proposal stage: ask for a sample report, even an anonymized one. It should contain four things.
The list of questions you measure visibility on. Not “industry queries,” but specific sentences that can be repeated. Without that, the next measurement will not be comparable with the previous one, and the whole method loses its point.
The result in numbers, broken down by model. How many mentions, how many citations, across how many answers. The breakdown by model matters, because visibility in Perplexity and in ChatGPT follows different rules and mixing them blurs the picture.
The list of domains cited instead of you. This is the most practical part of the whole report. It shows who the model currently considers an authority in your industry and, along the way, hints at where it is worth securing a presence. This list is often completely different from Google’s top 10, and that is exactly why it is valuable.
A comparison with the previous measurement. A single number says nothing. Only the difference over time shows whether the work is paying off.
If the report contains only the number of published articles and an organic traffic chart, you are getting an SEO report with the word GEO added to the header.
When GEO does not make sense
Something agencies are reluctant to talk about. GEO is not for everyone.
If your website is not indexed in Google or takes fifteen seconds to load, GEO will be a waste of money, because models reach for content search engines already know. Foundation first, then the layer on top.
If you sell a low-value impulse product where the customer does no research and the decision takes seconds, GEO makes little sense. It delivers the most value where a customer looks for a vendor, compares, and asks for a recommendation, which means services and B2B sales.
If you have budget for one channel and nothing is promoting you today, paid advertising will deliver faster. GEO is a medium-term investment, not a way to get leads this week.
How we do it at adsfox
We have been running performance marketing since 2018, have worked with more than 350 companies in over 20 industries, and are a Meta Business Partner and Google Partner. We treat GEO as an extension of SEO, not a separate discipline, and we measure it with the method described above, the same one for clients and for ourselves, every month, on the same list of questions. We show the baseline and the method before the engagement starts, so there is always something to compare the result against.
If you want to check how your brand shows up in model answers today, start with our AI SEO agency page or get in touch, and we will run a baseline measurement for your industry.


