Skip to content

← Research · Measurement note

Published · · Measured by · · k=5

Which AI searches, and which answers from memory?

A field note from AVE Studio · 180 measured runs across three engines · Budapest · 12 August 2026

What we measured

AVE Studio put twelve buyer questions to three AI engines on 12 August 2026 — six questions a company actually types when looking for a provider, and the inverse of each one ("who should I avoid"). Every question ran five times on every engine (k=5), for 180 runs in total.

We did not study the wording of the answers. We recorded which sources each answer stood on, because the wording reshuffles from run to run while the source set stays comparatively stable.

Finding 1: the three engines are not one thing

EngineRuns that used any sourceQuestions with any sourceDistinct domains cited
ChatGPT7 / 60 (12%)2 / 1241
Gemini43 / 60 (72%)11 / 12284
Perplexity60 / 60 (100%)12 / 12218

In AVE Studio's measurement, ChatGPT drew on live sources in 12% of runs while Perplexity did so in every single one. That is not a small difference in degree; it is a difference in kind. Work that makes a site more reachable and more current acts directly on Perplexity, partly on Gemini, and — for these questions — barely on ChatGPT, which mostly answered from what it already knew.

Finding 2: on the negative question, ChatGPT never searched at all

Split by question polarity, the gap widens. Across the six "who should I avoid" questions, ChatGPT used a source in 0 of 30 runs in AVE Studio's measurement, while Perplexity did so in 30 of 30.

The second half of the same finding: across 178 extracted mention records, the engines named a company to avoid exactly zero times. They answered with criteria instead — and where they named anyone, it was as a positive counter-example. The "who should I avoid" answer is assembled from a checklist, and it reasons with the framework of whoever published that checklist.

The illustration, selected by rule

Genre rule: the quoted output is picked by a stated, reproducible rule, never by hand. The rule here: of the six negative questions, take the one whose text sorts first alphabetically in Hungarian; for each engine take the earliest run of the 12 August 2026 round, ordered by run timestamp.

That resolves to "Kihez ne forduljak, ha a cégemet nem említi az AI, amikor a vevőim kérdezik?" — "Who should I NOT turn to when the AI doesn't mention my company to my buyers?"

  • ChatGPT, 14:42:52Z — 0 sources. Opens: "If you want AI systems to know and mention your company, here are some steps worth taking…"
  • Gemini, 13:10:43Z — 0 sources. Opens by listing categories of people to avoid.
  • Perplexity, 12:58:56Z — 20 sources, among them bakokrisztian.hu, bamboo-h.com, marketingprofesszorok.hu and minner.hu. Opens: "If you're asking who to turn to when your company doesn't appear in AI answers…"

Finding 3: two of the three engines answered the negative question as a positive one

Look again at what ChatGPT and Perplexity actually did above. Neither answered the question that was asked. Both silently converted "who should I NOT turn to" into "who should I turn to" and answered that instead. Only Gemini stayed with the negative framing.

This is one illustrative case, not a rate — we have not yet measured how often the conversion happens. We are publishing it because it points at something the aggregate numbers hide: the absence of counter-recommendations may be partly a refusal to answer negatively at all, rather than only an absence of negative material to draw on.

Finding 4: searching often and drawing widely are not the same property

The last column of the table is worth a second look. Perplexity searched most often (60/60), yet Gemini cited more distinct domains — 284 against Perplexity's 218 — despite searching in only 72% of runs. ChatGPT stopped at 41.

So "how often it searches" and "how widely it draws" are two separate properties. In AVE Studio's measurement Gemini reached for the web less often, but when it did, it brought sources from more varied places. The practical consequence: a narrow, well-optimised set of pages earns repeat placement in Perplexity sooner, while Gemini responds more to presence across many different sites.

What we did not measure

We publish this section because our own standard demands it — and because without it the numbers above would claim more than we know.

  • We did not measure how often the conversion happens. Finding 3 comes from one selected case; what share of negative questions get flipped to positive is an open question.
  • We did not measure why ChatGPT searched in that 12%. It had sources on two questions and none on ten; what distinguishes those two does not follow from this round.
  • We did not measure indexing. The probe for it returned nothing on all fifty URLs because of a provider quota — so this note says "we don't know", not "not indexed".
  • We did not assess answer quality. Only which sources the answers stood on.

How to reproduce this

The method is deliberately simple, so anyone can run it. Pick six questions your buyer actually types, and form the inverse of each ("who should I avoid…"). Run them all on the same day, across several engines, at least five times per question per engine. Record not the wording of the answer but the cited sources — and separately, whether there was any citation at all. The three numbers that fall out (runs with a source, questions with a source, distinct domains) are already comparable against next month. AVE Studio re-measures these same six pairs and publishes the result even when nothing has changed.

What this changes in practice

If your question is "why doesn't ChatGPT mention us", crawlability and freshness are not the lever — in these runs ChatGPT was not reading the live web at all. What reaches that layer is what other people have written about you, in enough places and for long enough. If your question is "how do we appear in Perplexity", the classic work applies immediately and shows up in days.

Both statements are about this measurement, on these twelve questions, on this day.

Limits

One round, three engines, twelve questions, k=5, on a Hungarian-and-English question set chosen by AVE Studio for its own market. It is not a national statistic and it does not transfer unexamined to another category. Engine behaviour also changes without notice — which is why the date is in the title of this note rather than in a footnote. The value here is the pattern and the method, not the decimal.

Citing this note

Free to quote and reproduce with attribution and a link to this page. The measurement is AVE Studio's; if you use the numbers, please carry the date (12 August 2026), the repetition count (k=5) and the engine list with them — a rate without those three is not a finding.