K3 · Stage 3 — Anchor: the corpus
Corpus Campaign· KEOpresence in the sources AI models learn from
The verified, canonical, readable, citable record AI models learn from — so each new model is trained to know and recall your business correctly.
Read the definitionGreat content is worthless if it never appears in the sources AI models can learn from. We plan your brand's presence on exactly those channels — with verifiable evidence and no impossible promises.
The Machine equityMachine equity is the business asset that arises from AI systems being able to recognize, understand, recall, recommend — and, on a user's behalf, transact with — a brand. Composed of three layers: rented reach (retrieval visibility), owned memory (trained model recall), and machine buyability (agent-side transactability). Measured honestly only as a portfolio of ranged, per-layer metrics — never as a single blended score.
Read the definition loop
Machine equity is the business asset that arises from AI systems being able to recognize, understand, recall, recommend — and, on a user's behalf, transact with — a brand. Composed of three layers: rented reach (retrieval visibility), owned memory (trained model recall), and machine buyability (agent-side transactability). Measured honestly only as a portfolio of ranged, per-layer metrics — never as a single blended score.
Read the definition- 1 · MEASURE — where you stand and which questions matter (Citation Tracker, Selector)
- 2 · FIX — why the AI skips you, fixed in code (Gap Analysis, Gap Closure)
- 3 · ANCHOR — an unambiguous entity, present where models learn (Entity Anchor, Corpus Campaign)
- 4 · PROVE ↺ — closed-book baseline and the delta at every new model generation (Model Memory Audit, Release Audit)
KG Doctor and the Claim Coherence Engine — the studio's own instruments — serve the entire loop.
AVE Studio doesn't run short-term campaigns: we build our partners' machine equity, across model generations.
AI models learn their knowledge of the world from vast text collections called corpora. Every model has a knowledge cutoff: only information available before it can enter training; anything published later can be learned at the earliest by a following model generation. If your brand isn't present in the right sources, the next generation won't necessarily know about it either.
01Placements in media with documented AI licensing agreements
We plan targeted placements in outlets with documented AI licensing deals (AP, FT, Guardian, Le Monde, News Corp, Axel Springer, Schibsted, Reddit — every claim backed by a source link), and on Hungarian surfaces that AI crawlers visit with high frequency.
02Timing aligned to model refreshes
We time publication to the expected refresh rhythm of model generations. We calculate when content must appear to still have a chance of entering the sources processed for the next generation.
03Verification and evidence
We check whether the content appears in the Common Crawl open web archive, and whether — per your server logs — the models' data-collection crawlers, such as GPTBot or ClaudeBot, actually visited the page.
Every service runs on an application we built in-house for exactly this job — not manual spreadsheets, not general-purpose tools.
A corpus-presence verifier: Common Crawl checks and server-log AI-crawler evidence, timed to model refresh rhythms.
- A quarterly campaign plan: with concrete, trackable tasks.
- A placement log: every item documented. We work only with real, earned press and professional presence; no fake profiles, no artificial manipulation.
- An AI-crawler visit report: presents your Common Crawl presence and the server-log evidence of AI crawler visits.
- A closing re-measurement: the work's actual impact is re-measured on the new model generation by the standalone Release Audit.
Here, the honest commitment isn't a footnote — it's the foundation of the service
- We don't promise what nobody can guarantee. Offers promising guaranteed inclusion in AI model training are misleading. Entry into training data cannot be directly controlled and cannot be proven in advance.
- We commit to what can actually be influenced. We build the brand's presence on the best-documented channels, then verify the outcome with measurement that can be checked after the fact. We cite AI licensing only with a source, and crawler visits only with measured evidence.
- The time constraints cannot be bypassed. New content can take effect at the earliest in a following model generation. The gap between a model's knowledge cutoff and its release can itself be months.
- An AI crawler visiting your page does not mean the model has learned the content. A crawler visit is an important early signal — not proof. Real proof comes from re-measurement on the new model generation.
Reports are delivered in Hungarian and English.
Industry figures are always reported as ranges, with sources. llms.txt is an information file written for AI crawlers — not a device that would secure entry into training. The work requires a client profile and, where possible, server logs. The process is closed by the Release Audit: that is where the work's effect becomes verifiable. ---
- How much does the corpus campaign cost?
- From €475 (net) per quarter. Media cost is on top, as a transparent pass-through: we neither absorb nor mark it up beyond a 15% management fee — the middle of the industry's 10–20% band.
- What do I receive?
- A quarterly placement plan aligned to the model release calendar, a licensed-outlet registry, and crawl evidence: Common Crawl and AI-crawler proof that your placements were actually reachable by the models' collectors.
- Do you guarantee I get into the models' knowledge?
- No — and anyone who guarantees that is overpromising: nobody controls model training. What we provide is cutoff-aware timing, the right outlets, and evidence that the content was reachable by the collectors. As a matter of principle we take no success fee on outcomes we cannot causally prove.
Curious what AI says about you?
Start with the free entry check — no commitment.
Request your free check