Skip to content

AVE Studio · AI Visibility Engineering

When not to trust a GEO agency: the 7 red flags

The mark of a bad GEO partner is not bad intent — it is the unmeasurable promise. If the offer guarantees a position, proves itself with a single screenshot, or promises AI visibility by a calendar date, that is not measurement, it is hope retail. Each of the seven red flags below comes with a control question — one that any agency, ourselves included, must be able to answer on the spot.

Request the AssessmentAssessment · a fixed fee, credited to the work

Updated ·

We did not invent this list

On 11 August 2026 we asked three AI engines, with repetition (k=5): who should you NOT trust with AI visibility work? Across 90 answers the engines never once named a company to avoid — every answer was a list of criteria. ChatGPT, verbatim: “caution is advised with anyone who guarantees immediate results”. On 12 August we repeated it at twice the scale: across 178 extracted mentions there was again ZERO counter-recommendation, and on the six negative questions ChatGPT did not search once in thirty runs — it worked purely from what it had learned earlier. The seven flags below build on the criteria classes the machines themselves use, extended with our own 18-point standard.

The 7 red flags

Seven signs, any one of which is reason enough to get a second proposal. Each comes with the question to ask back — and what an honest provider would answer.

1. Guarantees a position or an appearance

AI answers reshuffle from run to run — two answers to the same question almost never share an order. Whoever promises a guarantee is selling something they do not control.

What actually happens: A guarantee is dangerous not because it promises a lot, but because it replaces measurement. If the contract says “appearance in ChatGPT answers”, then the proof of delivery becomes a single favourable run — and there is always a favourable run if you ask often enough. Your money buys a lucky sample rather than work. The honest phrasing is not a promise but a direction: which rate are we trying to raise, from what baseline, and when do we measure again.

Control question: “From how many repetitions, and with what margin of error, do you measure the result?”

2. A single screenshot is the proof

One run is not a measurement: the same question can return a different answer five minutes later.

What actually happens: A screenshot is dangerous evidence because it cannot be reversed: you cannot tell which attempt it was. The same question can surface different sources five minutes later, and in our own measurement 12 to 40 new domains entered the sources behind each question within a single day. A picture shows one instant of that. Ask for this instead: in how many runs out of how many did the brand appear, and when did they run.

Control question: “How many times did you ask this exact question, and in what share of the answers did the brand appear?”

3. Promises results by a date

Model refresh cycles are neither public nor controllable — anyone quoting a calendar date is guessing.

What actually happens: A promised date conflates two different things: what the provider does, and what the model does. You can commit to a deadline on the first — page shipped, measurement run, report delivered. Not on the second: model refresh cycles are not published. In our own measurement Perplexity picked up a new study within days, while ChatGPT's memory layer is slower by an order of magnitude. Anyone quoting one date for both does not understand one of them.

Control question: “What can you measure from your own work, and what depends on the model provider?”

4. Keeps the question set secret

If you cannot see the question set, the result cannot be interpreted — a pretty number can be measured onto anything.

What actually happens: The question set is the most important part of a measurement, because it decides what is being measured at all. If you cannot see it, you cannot check the most convenient cheat: querying your own brand name. Your buyer does not search that way. Ask for the questions verbatim and see whether you recognise your buyer's own sentences in them — if you do not, the report measures another company's problem.

Control question: “Do I get a per-question breakdown of what the AI answers to MY buyers' questions?”

5. Sells an “AI position” or average ranking

Rankings change run to run; the only honest number is a mention rate with a margin of error.

What actually happens: “You are third in ChatGPT” sounds good because it is familiar: we brought it from Google. But there you have a stable result list, whereas here every run builds new text. A position is therefore a property of that run, not of you. The only number that carries into next month is the rate: in how many runs out of how many you appeared — and how uncertain that is.

Control question: “Why am I getting a ranking instead of a rate with a confidence interval?”

6. Never says “we don't know”

Machine measurement has three states: measured, not measured, inconclusive. Whoever has a number for everything is inventing one somewhere.

What actually happens: The absence of “we don't know” is telling because a great deal genuinely cannot be measured. In our own work, for instance, a fifty-URL indexing probe returned nothing for all fifty because of a provider quota — and what went into the report was that we do not know, not that the pages are unindexed. The difference between those two is that one of them is worth spending money on and the other is not.

Control question: “Show me a report that contains the line: this we could not measure.”

7. Uses your own brand name as proof

A nice answer to “what do you know about company X?” is not a recommendation — your buyers do not ask that way.

What actually happens: Querying the brand name is the most widespread mistake because it always produces a pretty result. If your company name is in the question, the model has to talk about you — that is the echo of your question, not a recommendation. Your buyer, meanwhile, types the problem, not your name. In our measurement the same company shows zero on the unbranded category question and looks excellent on the branded one: one company, two entirely different realities.

Control question: “Do you measure unbranded buyer questions, or do you query my company name?”

What they say, and what they should be saying

The same six situations, phrased two ways. The left column is not a lie — it simply cannot be checked. Every sentence in the right column can be questioned back, and the answer either exists or it does not.

What the bad offer saysWhat a provider who measures says
“You will get into ChatGPT's recommendations.”“Today you appear in three of twenty-five buyer questions. That is the number we want to raise, and we re-measure monthly.”
“Look, here's a screenshot — it mentions you.”“Four mentions in ten runs: 40% [12–74% confidence interval], on 12 August 2026.”
“It'll be done in three months.”“Our work takes six weeks. When the model picks it up is not ours to decide — which is why we measure instead of promising.”
“We measure with our own proven question set.”“Here are the twenty-five questions, verbatim. Read them: is this what your buyer asks?”
“You've moved up to third place.”“Position shifts run to run, so we report a rate. What did change is the source set: two new domains appeared.”
“We have data on everything.”“We measured these three; these two we could not — and here is why.”

What to ask for before you sign

Six documents. None of them is a trade secret, and anyone who genuinely measures can produce all six in five minutes. If any of them draws the answer “that's part of our methodology”, that is itself an answer.

  1. 1. The question set, verbatim

    Not a list of topics but the exact sentences as they go into the engine. This decides whether they are measuring your market or a generic one.

  2. 2. The repetition count and the measurement date

    How many runs per question and per engine, and when. Without repetition the number is an anecdote; without a date you cannot tell when it expires.

  3. 3. A margin of error next to every rate

    Four out of ten runs is not “40%” but 40% with a fairly wide band. Dropping the band makes the measurement look more precise than it is.

  4. 4. The list of things not measured

    The report should contain a section on what could not be measured and why. Its absence is the fastest filter on this whole list.

  5. 5. The raw answers, not just the summary

    Ask for the engine's actual text and cited sources, at least as a sample. A summary can show anything; a raw answer cannot.

  6. 6. The terms of the re-measurement

    When, with the same questions, at the same repetition count. If the re-run uses a different set, the two numbers are not comparable — and that is precisely what tends to happen.

How to use this in the conversation

The seven flags are worth something only if you say them out loud. In practice what decides is not the content of the answer but whether it exists at all — someone who measures has these numbers to hand; someone who does not starts talking about their methodology.

Start with the question set, not the price

Your first question should be whether you get the questions verbatim. This decides whether your market is being measured, and it is the point where fewest people can dodge convincingly. If you do get them, read them as if you were your own buyer: do you recognise yourself?

Push back on a single number

Pick one claim from the offer and ask how many runs it comes from. You need no statistics for this: “we measured that over five repetitions, on 12 August” is a different kind of answer from “in our experience”. The difference is audible.

Watch what they say when they don't know

The most revealing moment is when you ask something that genuinely has no data behind it — such as when ChatGPT's memory will pick up a change. Someone who measures will say that this cannot be known, and tell you what they can measure instead. Someone who does not will give you a deadline.

The positive pair of this list — what to demand before you hire anyone: How to choose a GEO provider — what to demand

What an honest measurement costs

We publish the price because when choosing a provider, knowing the order of magnitude is itself a defence. With us the entry point is the Citation Tracker, from €240 net; one-off instruments run from €240 to €490 net, and continuous measurement from €200 net per month. This is our list, not a market average — but if someone offers a “full AI visibility audit” for a fraction of it, it is worth asking how many runs it consists of.

This yardstick can be turned on us too

We fall short of our own standard on three counts today: our methodology is self-published rather than independently audited; we have few publishable, independent client results; and we measure at k=10 repetitions, not k=100. A fourth came out of the most recent round: when comparing the length of competitor pages we took a language model's ESTIMATE as a measured number and published an error of 1.5–2×, corrected in our public findings register. We write this down because it is the counterpart of the seventh flag above: a standard its author is exempt from is decoration.

Frequently asked questions about choosing a GEO agency

Eight questions about choosing a provider: what to ask for alongside the proposal, what to measure yourself, and when to walk away.

How do I know an AI visibility report is a real measurement?

Three signs: a repetition count (k) and a margin of error next to every rate; a per-question breakdown; and at least one “we could not measure this” line. If all three are missing, it is not a measurement.

Why won't the AI name whom I should avoid?

We measured it: out of 90 negatively framed questions, 0 answers named a culprit. Engines give criteria and positive counter-examples — which is why this page gives criteria too.

Can an appearance in an AI answer be guaranteed?

No. The answer is rebuilt on every run; what can be measured and influenced is the source set the engine works from. If someone claims otherwise, ask them flag one's control question.

What are k and a confidence interval?

k is the number of repetitions: we ask the same question k times. The confidence interval tells you how much to trust the resulting rate. At k=1 there is no interval — that is an anecdote.

What does an honest AI visibility measurement cost?

With us the entry point is from €240 net, one-off instruments from €240 to €490 net, and continuous measurement from €200 net per month. The number matters less than what sits behind it: how many questions, how many repetitions, how many engines. The same sum can be expensive for one run and cheap for three hundred.

What if the agency won't share the question set?

Then the report cannot be checked, and you have no answer to the question that matters most: did they measure your buyer's sentences? A question set is not a methodological secret — the methodology is what you do with the results. If the questions are confidential, the number cannot be authenticated.

Can't I just ask ChatGPT about my company myself?

Fine for orientation, not for measurement, for two reasons. One: if your company name is in the question, the model has to talk about you. Two: a single run tells you nothing, because the answer reshuffles run to run. If you do try it, use your buyer's words, leave your name out, and ask five times.

How do I know you don't do exactly the same?

You don't — which is why the seven check questions are on this page, and why we wrote down where we fall short of our own standard. The practical test is simple: ask us for the same six documents we suggest asking anyone else for. If any of them draws an evasive answer, the same conclusion applies to us.

If you have already signed, and now recognise the list

This is not necessarily a problem, and it does not mean you were cheated — most providers use the metrics the market has accepted, in good faith. Three steps, in this order. First, ask for the six documents above, without a deadline and in a neutral tone: HOW the answer arrives tells you more than what it contains. Second, ask for a re-measurement with exactly the question set the first one used — if it was not kept, there is nothing to compare against, and that is worth knowing. Third, if the numbers cannot be reproduced, do not attack the contract; restore the measurement instead: agree on a shared question list and repetition count for the next cycle. The goal is not to be proven right, but to have something to measure against next month.

Turn the yardstick on us

Ask us the seven control questions too — we answer them standing. And if what you need is not an audit but a change: AVE Studio ships the fixes in code that make the engines answer differently about you. Write or call:

hello@avestudio.pro+36 30 900 1356