The AVE Studio GEO methodology: 18 named procedures for measuring AI visibility, checkable line by line
AVE Studio measures per engine whether answer engines name a brand and whether they cite it, publishes every number with a confidence range, and executes the fixes in code. This page shows what that means in practice, line by line: 18 named procedures, each one checkable on our own public surface.
Published ·

Measuring AI visibility is not yet standardised. The same word — visibility, share of voice in LLMs, citation — means different things at different vendors, and most reports hand you a single number where a range would be the correct answer.
AVE Studio therefore does not publish claims about its methodology. It publishes named procedures. A procedure has a name, has an error, and can be held to account. Every section below introduces one.
What AVE Studio's methodology does, in one sentence
AVE Studio measures whether answer engines can find a brand, understand it correctly and name it, then fixes the gaps at code level and re-measures on the same question set.
The measurement runs on two axes. Generative Engine Optimization, or GEO, covers retrieval: what an engine finds about the brand at the moment the question is asked. Knowledge Engine Optimization, or KEO, covers parametric memory: what the model knows about the brand with no web access at all. Together they make up machine equity.
AVE Studio measures six engines: ChatGPT, Google Gemini, Anthropic Claude, Google AI Overviews, Google AI Mode and Perplexity. One of the six, Google AI Mode, deliberately carries zero weight wherever results are aggregated. AVE Studio measures it and reports it, but has no measured basis for how much weight it should carry against the others — and without evidence there is no weight. Zero weight is not a downgrade of the engine; it is the honest consequence of missing evidence.
The AVE Studio panel of 31 fixed questions run on 2 August 2026 covered four engines: ChatGPT, Google Gemini, Anthropic Claude and Google AI Overviews.
The measurement toolkit consists of eleven named procedures: the Wilson score interval, the Newcombe difference interval, the naming-rate gate, the competitor-gap ranking, exclusion of branded prompts, unaided and aided recall, per-language question-set design, FCrDNS bot verification, a single central robots interpretation, deterministic knowledge-graph deviation, and the refuted-phrase gate.
The AVE 18-point methodology standard, and why AVE Studio publishes it
The AVE 18-point methodology standard is a set of 18 yes-or-no questions that let anyone check an AI visibility provider's methodology on that provider's own public surface.
AVE Studio publishes the standard because the buyer of an AI visibility audit currently has nothing specific to ask. Each of the 18 questions is one where the answer is either present on a public page or it is not.
The questions of the AVE 18-point methodology standard:
- Does it report results per engine, or give one blended score?
- Does it separate being named from being cited?
- Does it state how many times a prompt is run per engine?
- Does it compute and publish a margin of error, with a named method?
- Does it test whether the difference between two measurements is significant?
- Does the user see a range, or a single number?
- Does it penalise statistically unstable questions when picking targets?
- Does it exclude brand-name prompts from the visibility metric?
- Does it measure whether the engine names any company at all for a question?
- Do target questions come from the market, or from the client's own domain?
- Does it measure what the model knows with no web access?
- Does it measure per market and per language, or run translated prompts?
- Does it verify that an engine's crawler can actually reach the page?
- Does it audit the brand's entity record and identifier consistency?
- Does it measure whether the AI states the brand correctly?
- Does it deliver a report, or code-level execution as well?
- Does it re-measure with the same method after fixes and after model changes?
- Does it publish its methodology, and an explicit list of what it will not promise?
The AVE Studio row is the first row of the matrix at the foot of this page, with the place each answer can be checked.
AVE Studio measures per engine and separates naming from citation
AVE Studio reports results per engine and never blends them into a single score, because naming and citation behave in opposite directions across engines.
The Semrush and Kevin Indig Ghost Citations study of June 2026 analysed 3,981 domain appearances across 115 prompts, 14 countries and four engines. One major engine named the brand in 83.7 percent of appearances but supplied a source link in only 21.4 percent; another did the reverse, with 87 percent citation against 20.7 percent mention. A blended score misrepresents both.
A ghost citation is what happens when an answer engine cites your page as a source but never says your brand name in the text of the answer. In the same study, 61.7 percent of brand appearances were of this kind, and only 13.2 percent produced a citation and a mention together. Brand visibility in AI answers and citation tracking are therefore two separate metrics in every AVE Studio report, each with its own confidence range.
The third hard rule enters at the same point. A prompt that asks for the brand by name is never a visibility metric at AVE Studio. Asking for your own name guarantees a self-mention, inflates the number, and says nothing about whether a buyer would find the company. AVE Studio handles branded prompts in a separate brand fact-check panel.
AVE Studio runs ten repetitions and states what may be read out of them
AVE Studio runs a prompt ten times per engine, and its unit of measurement is not the individual question but the portfolio: 25 to 40 questions, five engines, ten repetitions, giving 1,250 to 2,000 binary observations.
Generative search is probabilistic. In the study by Atil and colleagues presented at the Eval4NLP workshop in 2025, five language models across eight tasks and ten runs produced accuracy differences of up to 15 percent between runs, and a gap of up to 70 percent between the best and worst achievable performance — under settings configured to be deterministic.
The five engines in that count are ChatGPT, Google Gemini, Anthropic Claude, Perplexity and Google AI Mode. Google AI Overviews is measured but left out of the repetition count, because it runs through SerpApi, where repetition carries a different cost and a different character.
The AVE Studio client report labels the prompt-level table as indicative rather than a measurement, and states the per-question uncertainty outright. Across ten runs, computed with the Wilson interval, uncertainty at the extreme cells — zero out of ten, ten out of ten — is roughly plus or minus 14 percentage points, and plus or minus 26 at the middle cells. At portfolio level, with 2,000 observations and a naming rate of 20 percent, the same figure is plus or minus 2 percentage points.
Ten runs are therefore still not enough to decide anything about a single question. AVE Studio does not allow per-question decisions to be drawn from them: the unit of measurement remains the portfolio.
Two further procedures govern targeting. The naming-rate gate measures whether an engine names any company at all for a question: if the answer is generic, the question is unwinnable from the start and drops out of the target set. In AVE Studio measurement, which-company questions carry an average naming rate of 100 percent, while generic discovery questions carry 22 percent. The competitor-gap ranking then orders targets by naming rate multiplied by the absence of the client's own presence: the place to aim is where the engine lists companies but not the client. High-variance prompts take an automatic deduction, because statistically unreliable fights are not worth recommending as targets.
AVE Studio uses the Wilson interval, and that choice decides what happens at zero
AVE Studio computes a Wilson score interval at 95 percent for every proportion it publishes, and names that estimator openly.
The choice of estimator is not a detail. The naive normal-approximation interval, applied to the observed proportion, returns zero percent with zero uncertainty when nothing is observed — which is statistically false. Zero mentions in twenty runs does not mean the true rate is zero; it means the true rate sits under a low but non-zero upper bound. The conservative variant of the same estimator, which maximises variance at a proportion of fifty percent, makes the opposite error: it returns an identically wide band for every cell, regardless of what was measured.
The Wilson interval does neither. It is asymmetric, it adapts to the observed proportion, and it does not collapse at the edges of the scale.
That decision matters exactly where most AI visibility tracking takes place. On unbranded category questions a Central European mid-market company typically sits near zero, and the near-zero region is where estimators diverge most. One finding in the statistical framework published by Ronald Sielinski in March 2026 and revised in June points the same way: once confidence ranges are drawn around citation share estimates, much of the apparent difference between competing domains overlaps. The paper's conclusion, verbatim, is that "single-run visibility metrics provide a misleadingly precise picture".
AVE Studio uses a Newcombe difference interval to say whether anything changed
AVE Studio computes a Newcombe hybrid-score difference interval for the gap between two measurements, and flags movement only when the confidence ranges do not overlap.
Testing the significance of change is the step that separates measurement from dashboard-watching. The difference between two readings may be real movement, and it may be noise. A report that does not say which leaves the decision to the eye.
The work submitted by Julius Schulte, Malte Bleeker and Philipp Kaufmann on 8 April 2026 reaches the same conclusion for generative search: answers vary across runs, prompts and time, so one-off observation is unreliable, and visibility has to be characterised as a distribution rather than a single-point outcome.
The AVE Studio significance flag is a deterministic check running in code, with a regression test behind it. It is not an editorial judgement, and it is not eyeballing.
AVE Studio also measures closed-book: the KEO axis
AVE Studio runs closed-book measurement, meaning it measures what the model knows about a brand with no web access, and reports the result on two separate metrics: unaided recall and aided recall.
Knowledge Engine Optimization, or KEO, is the layer that retrieval-based GEO cannot see. Unaided recall measures which companies the model names for the category, with no mention of the brand name. Aided recall measures what the model says once it is given the brand name. The gap between the two shows whether the brand exists in the model's parametric memory at category level, or only when asked about directly. Model Mindshare is the metric that shows how much room a brand occupies in a model's category answers.
The work by Kandpal and colleagues published at ICML in 2023 showed that language models struggle to learn long-tail knowledge: the ability to answer a factual question tracks how many relevant documents the model saw during training. A Central European mid-market company sits in the long tail by definition, which is why the closed-book axis is not a secondary concern for it but the binding constraint.
In the AVE Studio survey of fifteen Hungarian companies conducted on 30 June 2026, closed-book recall measured with a single trial was accurate for three companies, partial for eight and absent for four. The sample is a national convenience sample measured with one model and one trial, and is therefore indicative rather than representative — a limitation that travels with the downloadable data file.
AVE Studio designs question sets per language rather than translating them
AVE Studio designs a separate question set for every market and never runs translated prompts, because a translated prompt is not the same prompt.
The study by Dayeon Ki and colleagues, accepted as an ICML 2026 Spotlight, found across eight languages and six open-weight models that models preferentially cite English sources when the query is in English, and that the bias is stronger for lower-resource languages. An April 2026 paper showed the same effect in the reranking layer: systems systematically favour English and the query's own language. Both are effects measured in controlled experimental settings rather than the live behaviour of commercial engines, and no Hungarian-specific benchmark exists.
That is what can be said, and no more: a measurable language-preference signal exists.
The AVE Studio panel of 2 August 2026 therefore ran in five languages — Hungarian, English, German, French and Polish — with a separately designed question set for each, 31 fixed questions and four engines. A country selector and per-language question design are not the same thing: the first sets where the query runs from, the second decides whether the question being asked is one the market actually uses.
AVE Studio starts with the access layer
AVE Studio begins every audit by verifying the access layer, because if an engine's crawler cannot load the page, the quality of the content is irrelevant.
In the AVE Studio survey of fifteen Hungarian companies conducted on 30 June 2026, five of the sites loaded for a simple crawler. Ten returned HTTP 403. The limitation is the same as for the survey's other figures: a national convenience sample, indicative rather than representative.
AVE Studio uses a single central robots interpretation across the whole system, guarded in code so that contradictory implementations cannot drift apart. It verifies crawlers using forward-confirmed reverse DNS: it does not trust the user agent, it checks whether the engine's crawler was really there. A user agent on its own is a claim; forward-confirmed reverse DNS is the verification of that claim.
The rest of the layer is the dual user-agent robots matrix, llms.txt, the JSON-LD graph and Atom feeds. Together these describe what engines are allowed to reach, and what machine readability they find once they reach it.
AVE Studio runs an entity audit and a factuality audit
AVE Studio measures whether a brand exists as an identifiable, consistent record in the machines' registries, and whether the AI states correctly what it says about the brand.
Entity SEO means the brand exists as an entity in the knowledge graph, not merely as text on a page. Structured data, the schema.org vocabulary and JSON-LD schema markup carry it, and identifier consistency across pages is the precondition. In the AVE Studio survey of 30 June 2026, eight of fifteen Hungarian companies had a credible entity record, and average entity strength on a zero-to-two scale was 2.0 for large companies, 0.8 for mid-sized and 0.4 for small ones. Knowledge graph optimization starts here.
The factuality audit goes one step further. AVE Studio builds a deterministic knowledge graph from the client's official data and compares it with what the LLM asserts about the brand; structural deviation flags the hallucination. The procedure follows the methodological pattern set out by Gupta and colleagues at the AAAI Symposium Series in 2025, which runs a deterministic and an LLM-generated knowledge graph in parallel and measures divergence through instantiation ratios. The AVE Studio implementation adapts that pattern to brand data; it is not the paper's original procedure.
The AVE Studio panel of 2 August 2026 documented that engines sometimes name companies that do not exist. The factuality audit gives that phenomenon a falsifiable measure.
AVE Studio executes fixes in code and re-measures on the same set
AVE Studio carries out fixes in code rather than in a PDF recommendation: at the level of schema, server-side rendering, the access layer, entity registration and content layers.
Execution includes the layer built for machines. MCP, the Model Context Protocol, belongs here, together with the agentic commerce standards: ACP, released by Stripe and OpenAI on 29 September 2025 under an Apache 2.0 licence, and UCP, announced by Google and Shopify on 11 January 2026, served through a well-known path manifest.
AVE Studio re-measures on the same question set with the same method after fixes and after model changes. An improvement measured on a different question set is not an improvement; it is a different measurement. AVE Studio reports re-measurement results in a temporal frame and asserts no causal link between the intervention and the movement.
What AVE Studio will not promise, and why that is a gate rather than an intention
AVE Studio does not promise guaranteed positions, guaranteed appearances or specific percentage visibility gains, and enforces that prohibition through a gate running in code.
A refuted-phrase gate lives in the repository and scans every AVE deliverable, including AVE Studio's own marketing. Anything that has once failed adversarial verification cannot be shipped again. On 7 June 2026, AVE Studio put 25 of its own claims through a three-vote refutation test: 19 were confirmed, six were rejected, and the six were withdrawn from all communication. Published manuscripts are content-locked: the shipped text is generated bit-for-bit from the ratified manuscript, and continuous integration fails if anyone edits it by hand.
The best-known number in generative engine optimization falls under the same gate. In the KDD 2024 work by Aggarwal and colleagues, the best-performing methods improved on the baseline by 41 percent on the position-adjusted word count metric and by 28 percent on the subjective impression metric — two numbers belonging to two metrics, not to two interventions. Adding source citations produced a 115.1 percent visibility increase for pages ranked fifth in the results, while the visibility of top-ranked pages fell by an average of 30.3 percent. All of it was measured in a 2023 experimental setup using GPT-3.5-turbo with Google retrieval, not on a 2026 commercial engine, and the paper's own abstract presents its forty percent figure as a ceiling rather than an average. This is precisely why the AVE Studio prohibition list rules out using that number as a promise.
Variance sets the limit of what can be promised at all. In the January 2026 work by Camuffo and colleagues, minor design choices in LLM-based text annotation shifted outcomes by 12 to 85 percentage points. At that order of instability, promising an outcome would be a sales decision rather than a methodological one.
GEO, AEO, LLMO, KEO, AI SEO: which is which?
AVE Studio does the work that the market searches for under four or five different names, so it is worth separating the terms.
Generative Engine Optimization, or GEO, means preparing content to appear inside the answers that answer engines generate. Answer Engine Optimization, or AEO, is the same work framed from the buyer's question. LLMO, AI SEO and LLM SEO name the same practice in different industry dialects, and AI search optimization is the broadest umbrella. Anyone comparing GEO agency offers is shopping in this set. KEO measures something different: the model's trained, parametric memory, with no web access.
The practical question buyers ask usually sounds like this: how do I get into ChatGPT's answers. The answer differs by engine. ChatGPT visibility, Google AI Overviews optimization, Google AI Mode, Gemini, Perplexity optimization, Claude and Copilot draw on different sources, and ChatGPT SEO does not transfer automatically to the rest. Prompt tracking and citation tracking have to run per engine for the same reason.
Google's own disclosures set the scale. AI Mode passed one billion monthly users, announced at the company's developer conference in May 2026. AI Overviews reached two billion monthly users, reported by Alphabet in its second-quarter 2025 earnings communication. Availability across more than 200 countries and territories and more than 40 languages was announced by Google in May 2025 with the international expansion. Three numbers, three dates, three sources.
Zero-click search is the shared consequence. The Pew Research Center analysis published on 22 July 2025 examined the browsing data of 900 US adults across 68,879 searches from March 2025. Where an AI summary appeared, users clicked a traditional result on 8 percent of visits, against 15 percent where none appeared, and clicked a source inside the summary itself on 1 percent. Browsing sessions ended after 26 percent of pages carrying an AI summary, against 16 percent of standard results pages.
On GEO vs SEO the best empirical answer is the Ahrefs correlation study of 75,000 brands, extended in December 2025 to ChatGPT and Google AI Mode. YouTube mentions showed a Spearman correlation of 0.737 with AI visibility, branded web mentions 0.664, backlinks 0.218, page count 0.194, and Domain Rating between 0.266 and 0.326. The Ahrefs researchers state themselves that this is correlation and not causation: classical search authority is a weak predictor of AI citation, but the measurement does not say what the cause is.
Where the field stands, and how AVE Studio measured it
AVE Studio examined the public surfaces of sixteen AI visibility providers on 6 August 2026 against the AVE 18-point methodology standard, using only each provider's own documentation, methodology page, help centre and published research.
The review uses three evidence levels. The first means the content of a methodology or documentation page was read verbatim. The second means content from the provider's own domain was seen, but the full documentation surface was not walked. The third means no claim surfaced from a primary surface in this round. Third-level providers were dropped from both the numerator and the denominator of every ratio below; there were seven of them, named in the matrix at the foot of this page. This page asserts nothing about them. What a provider does not publish supports one statement only: that it does not publish it.
Nine providers remained at the first or second evidence level. As of 6 August 2026:
- Four state how many times they run a prompt.
- Two publish a margin of error.
- None names the estimator it uses on its methodology page.
- None publishes a test for the significance of change.
- One publishes that it measures the model's trained, web-free knowledge.
- Three publish per-market runs; none publishes question sets redesigned per language.
- None publishes an explicit list of what it will not promise.
- None was found that carries the full chain — measurement, target selection, gap analysis, code-level fixes, re-measurement — end to end and publishes it.
AVE Studio regards one development in 2026 as the most important of these: the category has begun publishing its own uncertainty. Two years earlier, none of these numbers was public anywhere.
How to refute this page
AVE Studio publishes the standard line by line so that the path to refuting it is short.
Open any provider's methodology page and look for a counter-example to the aggregates above. Specifically, these are worth looking for:
- A provider that names the estimator used to compute its margin of error.
- A provider that publishes a test for when the difference between two measurements is significant.
- A provider that measures with question sets redesigned per language rather than translated prompts.
- A provider that publishes an explicit list of what it will not promise.
The AVE Studio row sits in the same matrix, answering the same questions, and is refutable the same way. A standard is worth something only if the party issuing it submits to it too.
This page was written with the same rubric AVE Studio applies to client pages: named sources, numbers, dates, samples and attributed references. If an answer engine lifts one paragraph out of it, that paragraph has to stand on its own.
The author is the founder of AVE Studio. AVE Studio is an AI visibility engineering studio based in Budapest: it measures and builds whether answer engines find, correctly understand and name a brand. Every AVE Studio figure on this page is verifiable in the linked research releases, together with the limitations attached to it.
The 18 criteria — AVE Studio is the first row
Each row is one yes-or-no question, with the AVE Studio answer and where it can be checked. The right-hand column records the field as of 6 August 2026.
| # | Criterion | The question | AVE Studio | Where it can be checked | Field |
|---|---|---|---|---|---|
| 1 | Per-engine breakdown | Does it report per engine, or give one blended score? | Six engines measured (ChatGPT, Gemini, Claude, AI Overviews, AI Mode, Perplexity), reported per engine; wherever results are aggregated, Google AI Mode carries zero weight, because there is no measured basis for a weight | /methodology, /geo | of the 6 most fully documented, 5 report per engine; 1 publishes a blended score |
| 2 | Mention vs citation | Does it separate being named from being cited? | Two separate metrics, each with its own confidence range | /methodology, /research/model-mindshare-cee-2026-q3 | 4 of the 6 separate them in published material |
| 3a | Sample size disclosed | Does it state how many times a prompt is run? | k = 10 per prompt per engine, disclosed; portfolio of 25-40 questions × 5 engines × 10 repetitions = 1,250-2,000 observations | client report, /methodology | 4 of the 9 reviewed disclose it |
| 3b | Margin of error published | Does the number carry a margin of error, with a named method? | Wilson score 95% CI; per-question uncertainty stated: ±14 percentage points at the extreme cells, ±26 at the middle cells; at portfolio level, with 2,000 observations and a 20% naming rate, ±2 percentage points | client report | 2 of the 9 publish a margin of error; 0 name the estimator on their methodology page |
| 3c | Change significance tested | Does a named test decide whether a difference is significant? | Newcombe hybrid-score difference CI, with the "ranges do not overlap" criterion | client report | 0 of the 9 publish a change-significance test |
| 4 | Uncertainty shown to the user | Does the user see a range, or a single number? | A range; the prompt-level table is labelled "indicative, not a measurement" | client report | 1 of the 6 shows a range, 5 show a point score |
| 5 | Variance-aware targeting | Does it penalise statistically unstable questions? | Automatic deduction for high-variance prompts | /methodology | 0 of the 6 publish it |
| 6 | Branded prompts excluded | Does it exclude brand-name prompts from the visibility metric? | Hard rule; branded prompts handled in a separate brand fact-check panel | /methodology | 0 of the 9 publish the exclusion; 1 publishes that it runs both branded and unbranded queries |
| 7 | Naming-rate gate | Does it measure whether the engine names any company at all? | Yes; measured example: which-company questions 100%, generic discovery questions 22% | /methodology | 0 of the 6 publish it |
| 8 | Market-seeded targeting | Do target questions come from the market or from the client's domain? | From the market; the domain is a scope filter, not a topic source | /methodology | 4 of the 6 publish a market-based prompt source; for two of them this is a licensed conversation panel |
| 9 | Closed-book measurement | Does it measure the model's trained, web-free knowledge? | Yes; unaided recall and aided recall as separate metrics | /methodology, KEO subpage, /research/hungarian-ai-search-readiness | 1 of the 9 publishes that it measures trained knowledge; 0 publish the two-way recall split |
| 10 | Per-language measurement | Does it measure per language, or translate? | 5 languages, each with a separately designed question set | /research/model-mindshare-cee-2026-q3 | 3 of the 9 publish per-market runs; 0 publish question sets redesigned per language |
| 11 | Access layer verification | Does it verify that an engine's crawler can get in? | Single central robots interpretation, FCrDNS bot verification, dual user-agent robots matrix, llms.txt | audit module, /research/hungarian-ai-search-readiness | 2 of the 6 publish access checks; 0 publish crawler verification as opposed to user-agent logging |
| 12 | Entity / knowledge-graph layer | Does it audit the entity record and identifier consistency? | Yes; entity strength on a 0-2 scale, identifier consistency across pages | KG Doctor, /research/hungarian-ai-search-readiness | 1 of the 6 publishes entity-consistency work |
| 13 | Factuality audit | Does it measure whether the AI states the brand correctly? | Deterministic knowledge graph vs. LLM-asserted knowledge; structural deviation flag | factuality module | 2 of the 9 publish partial accuracy monitoring (pricing, product descriptions); 0 publish deterministic knowledge-graph comparison |
| 14 | Execution | Report only, or code-level execution as well? | In code: schema, server-side rendering, access layer, entity registration, content layers, MCP/ACP/UCP | /geo, /agentic-commerce | 3 of the 6 publish an execution layer |
| 15 | Re-measurement | Does it re-measure with the same method? | Same question set after fixes and after model changes; temporal frame, not causal | /methodology | 1 of the 9 publishes an explicit re-measurement loop; for the rest it is implicit in time-series running |
| 16 | Published methodology + no-promise list | Is the methodology public? Is there a "what we will not promise" list? Is there published primary research? | Published methodology, "what we will not promise" section, refuted-phrase gate running in code, two primary studies with downloadable data, own visibility baseline published | /methodology, /research/* | 5 of the 9 publish methodology or primary research; 0 publish an explicit "what we will not promise" list; 0 publish their own measured visibility on unbranded category questions |
The field reviewed, provider by provider
Sixteen providers were reviewed: nine at evidence level E1 or E2 (these form the denominator of every ratio), seven at E3 (excluded from both numerator and denominator). The AVE Studio row sits above them as the reference point.
Evidence levels. E1 = the content of the methodology or documentation page was read verbatim. E2 = content from the provider's own domain was seen, without walking the full surface. E3 = no claim surfaced from a primary surface in this round; counts towards neither numerator nor denominator.
| Provider | Level | 3a | 3b | 3c | 6 | 9 | 10 | 13 | 15 | 16 |
|---|---|---|---|---|---|---|---|---|---|---|
| AVE Studio | own | publishes (k=10 per engine) | publishes (Wilson 95%, ±14 / ±26 per question, ±2 at portfolio level) | publishes (Newcombe) | publishes | publishes | publishes | publishes | publishes | publishes |
| Evertune | E1 | publishes (100× per model) | publishes (±1 / ±3 / ±10 points by sample level); does not name the estimator | does not publish | does not publish | publishes (baseline knowledge isolated via direct API vs. live consumer app) | does not publish | does not publish | partly | publishes |
| Gumshoe | E1 | publishes (~800 observations per run) | publishes (±5 pp @95%, ±2.5 pp on repeat); formula on the blog, estimator not named | does not publish | does not publish | does not publish | does not publish | does not publish | partly | publishes |
| Peec AI | E1 | publishes (1 run per prompt per model per day) | does not publish | does not publish | does not publish | does not publish | publishes (per-prompt country, 80+ markets) | does not publish | partly | publishes |
| Otterly.ai | E1 | partly ("daily", k not stated) | does not publish | does not publish | does not publish | does not publish | publishes (per-prompt country, 50+) | does not publish | partly | publishes |
| Scrunch | E1 | partly (daily / 3-day refresh) | does not publish | does not publish | does not publish | does not publish | does not publish | partly (pricing, product descriptions) | partly | partly |
| Yext Scout | E1 | publishes (per-location question set, monthly) | does not publish | does not publish | partly (runs both branded and unbranded queries; states no exclusion) | does not publish | partly (location level) | partly | publishes | publishes |
| Profound | E1 | does not publish | does not publish | does not publish | does not publish | does not publish | partly (region filter) | does not publish | partly | partly |
| Semrush AI Toolkit | E2 | does not publish | does not publish | does not publish | does not publish | does not publish | publishes (14 countries in the Ghost Citations study) | does not publish | partly | publishes |
| Ahrefs Brand Radar | E2 | does not publish | does not publish | does not publish | does not publish | does not publish | partly | does not publish | partly | publishes |
| Goodie AI | E3 | — | — | — | — | — | — | — | — | — |
| Brandlight | E3 | — | — | — | — | — | — | — | partly | — |
| Athena HQ | E3 | — | — | — | — | — | — | — | — | — |
| Daydream | E3 | — | — | — | — | — | — | — | — | — |
| Conductor | E3 | — | — | — | — | — | — | — | — | — |
| BrightEdge | E3 | — | — | — | — | — | — | — | — | — |
| seoClarity | E3 | — | — | — | — | — | — | — | — | — |
Note on the extended criteria (1, 2, 4, 5, 7, 8, 11, 12, 14): these were coded on the top six (Profound, Evertune, Peec AI, Semrush, Ahrefs, Yext), plus wherever the same document happened to answer them. Scrunch publishes at E1 level on criterion 11 (bot traffic tracking, robots allowlist verification, a code-light page version served to agents) — the strongest published access layer in the field, and the matrix records it as such.
No silent cuts. What was not reviewed: the documentation surfaces of the seven E3 providers; the Evertune help centre and whitepapers; the Profound documentation subdomain; the open-source GEO/AEO tooling landscape.
Frequently asked questions
- What makes an AI visibility methodology advanced?
- An AI visibility methodology is advanced when it states the size of its own error. The AVE 18-point methodology standard breaks that into 18 yes-or-no questions: per-engine breakdown, disclosed sample size, a margin of error computed with a named estimator, change significance, closed-book measurement, access layer, factuality, code-level execution and re-measurement.
- What is the difference between GEO, AEO, LLMO and KEO?
- GEO, AEO and LLMO name the same work: preparing content to appear inside answer engine responses. Knowledge Engine Optimization measures something different — the model's trained, web-free parametric memory. AVE Studio measures the two separately, and calls the combination of both machine equity.
- How many times does an AI prompt have to be measured to be reliable?
- Once is never enough. AVE Studio runs ten repetitions per prompt per engine and states the resulting per-question uncertainty: ±14 percentage points at the extreme cells and ±26 at the middle ones. The unit of measurement is therefore the portfolio: 25-40 questions, five engines, ten repetitions, 1,250-2,000 observations, where the range narrows to ±2 percentage points.
- Why can appearing in ChatGPT not be guaranteed?
- Generated answers are probabilistic. In a 2025 study, five models across eight tasks and ten runs produced accuracy differences of up to 15 percent under settings configured to be deterministic. AVE Studio therefore promises no position, no appearance and no percentage gain, and enforces that through a gate running in code.
- What is a ghost citation, and why does it matter?
- A ghost citation is when an answer engine cites your page as a source but never says your brand name in the answer. In a June 2026 study of 3,981 domain appearances, 61.7 percent were of this kind. AVE Studio therefore reports naming and citation as two separate metrics.
- How can what a model knows about a brand without the web be measured?
- Through closed-book measurement. AVE Studio queries the model with web retrieval disabled and reports two separate metrics: unaided recall, where the model names companies for the category with no mention of the brand name, and aided recall, where it is given the brand name. The gap between them shows the strength of parametric memory.
- How can a GEO provider's methodology claims be checked?
- Open its methodology page and look for answers to the AVE 18-point methodology standard. Four questions filter fastest: does it state the sample size, does it name the estimator behind its margin of error, does it test change significance, and does it publish an explicit list of what it will not promise.
- What does AVE Studio measure that an AI SEO agency typically does not?
- AVE Studio measures parametric memory alongside retrieval, publishes its margin of error with a named estimator, decides change significance with a named test, designs question sets per language, and executes fixes in code. All 18 points are checkable line by line in the matrix.
How to cite this page
AVE Studio (2026). *The AVE Studio GEO methodology: 18 named procedures for measuring AI visibility, checkable line by line.* AVE Studio, Budapest. Published: 2026-08-06. Available at: https://avestudio.pro/en/blog/ai-visibility-methodology
Sources
Every source with a resolvable URL and the date it was accessed. Vendor documentation changes; the article's field aggregates record the state as of the date below.
Academic
- GEO: Generative Engine Optimization — Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande — KDD 2024, 2024Accessed:
- Non-Determinism of "Deterministic" LLM Settings — Atil et al. — Eval4NLP 2025, 2025Accessed:
- Quantifying Uncertainty in AI Visibility — Ronald Sielinski (IQRush), 2026Accessed:
- Don't Measure Once: Measuring Visibility in AI Search (GEO) — Schulte, Bleeker, Kaufmann, 2026Accessed:
- Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG — Ki, Carpuat, McNamee, Khashabi, Yang, Lawrie, Duh — ICML 2026 Spotlight, 2025Accessed:
- All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG — arXiv preprint, 2026Accessed:
- Large Language Models Struggle to Learn Long-Tail Knowledge — Kandpal, Deng, Roberts, Wallace, Raffel — ICML 2023, PMLR 202:15696–15707, 2023Accessed:
- Continuous Monitoring of Large-Scale Generative AI via Deterministic Knowledge Graph Structures — Gupta, Haque, Ali, Kamal, Alam, Rahman — AAAI Symposium Series 2025, 2025Accessed:
- Variance-Aware LLM Annotation for Strategy Research — Camuffo, Gambardella, Kazemi, Malachowski, Pandey — Bocconi, 2026Accessed:
Industry study
- Ghost Citations Study — Semrush × Kevin Indig, 2026Accessed:
- The ghost citation problem — Kevin Indig — Growth Memo, 2026Accessed:
- An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied) — Louise Linehan, Xibeijia Guan — Ahrefs, 2025Accessed:
- Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews (75k Brands Studied) — Louise Linehan — Ahrefs, 2025-12Accessed:
- Google users are less likely to click on links when an AI summary appears in the results — Athena Chapekis, Anna Lieb — Pew Research Center, 2025-07-22Accessed:
Primary company statement
- 100 things we announced at Google I/O 2026 — Google, 2026-05Accessed:
- Alphabet Q2 2025 earnings call: CEO's remarks — Sundar Pichai — Alphabet, 2025-07-23Accessed:
- AI Overviews international expansion (May 2025 update) — Google, 2025-05-20Accessed:
- How Evertune Measures AI Visibility — Evertune, 2026Accessed:
- Repeated sampling enables GEO action by topic and prompt — Will Robinson — Evertune, 2026Accessed:
- How Gumshoe Works — Gumshoe, 2026Accessed:
- How Much Data Do You Need to Measure AI Visibility with Confidence? — Nick Clark — Gumshoe, 2026Accessed:
- Peec AI documentation — setting up your prompts — Peec AI, 2026Accessed:
- Otterly.ai help — search prompt monitoring — Otterly.ai, 2026Accessed:
- Guide to Boosting Brand Presence in AI Search — Scrunch, 2026Accessed:
- Introducing the Visibility Score in Yext Scout — Yext, 2025Accessed:
- What Makes Yext Scout the Most Comprehensive Brand Visibility Tool? — Yext, 2026Accessed:
- Prompt Volumes — Profound, 2026Accessed:
Own research and statements
- Hungarian AI-search readiness — AVE Studio, 2026Accessed:
- Model Mindshare Report — CEE 2026 Q3 — AVE Studio, 2026Accessed:
- Methodology — AVE Studio, 2026Accessed:
- Vocabulary — AVE Studio, 2026Accessed: