The short answer
The keyword did not disappear. It changed hands. The user no longer types it; the engine writes it, sometimes dozens of them, before composing an answer. Google calls this query fan-out. Several platforms now offer to extract those machine queries and resell them to you as a new keyword list to work on.
We tested whether that list means anything. 224 calls to the Gemini API, 3 commercial intentions, 8 equivalent paraphrases of each, 2 models, and 4 hypotheses whose refutation criteria were fixed before data collection. Two of the four failed. We publish them as failed, which is what makes the rest contestable.
Three results, all measured:
- Two runs of the exact same prompt, word for word, share only 3 to 9% of their machine queries. On one model, 62 calls produced 237 distinct queries.
- On the model you are using today, half the calls issue no query at all. The answer comes straight out of the model's weights. On one of our three intentions, no search happened across 24 calls.
- On local intent, 94% of queries name a specific venue. And when you remove the search tool, the model already names those same venues.
So a keyword list extracted from an AI engine describes one execution and the model's prior beliefs. Not your customer's intent. The full paper, raw data and code are at the bottom of this article.
One intention, two phrasings, no shared words
On 1 October 2026, two searches run seconds apart:
| Query typed | Geography resolved by the engine | Featured answer |
|---|---|---|
meilleure pizza Paris ("best pizza Paris") |
Paris, FR | Peppe Pizzeria, 50 Top Pizza ranking |
Pizzeria numéro 1 dans la capitale de la France ("number 1 pizzeria in the capital of France") |
France | Peppe Pizzeria, Giuseppe Cutraro, 50 Top Pizza |
Not one word in common beyond the "pizz-" root. Two different geographic resolutions. The same address on top.
A string-matching engine, the 2010 kind, would have treated these two lines as two separate markets: two search volumes, two difficulty scores, two pages to write. An engine that encodes the query and the index into the same vector space treats them as two neighbouring points. The unit of work is no longer the string, it is the intention.
That is where keyword thinking starts to cost money, and there is a figure for it. Semrush matched the prompts observed in its clickstream panel, over a billion rows across 200 million US users, against its own base of 27 billion keywords: over most of the October 2024 to February 2026 window, 65 to 85% of prompts match no existing keyword (Semrush). No volume, no CPC, no difficulty. Nothing to buy.
Incidentally, the same study corrects a widely circulated figure. The "AI prompts are 23 words versus 4 on Google" claim (and its "60 versus 3.4" variant) counts all prompts, including coding, translation, writing and brainstorming, none of which ever trigger a web search. Separate the two populations and the reading flips:
| Prompt population | Jan-Feb 2025 | Jan-Feb 2026 |
|---|---|---|
| Prompts that do trigger a search | 4.7 words | 8.7 words |
| Prompts that do not | 24.9 words | 13.5 words |
The high average comes from prompts that search for nothing, and it is falling. The ones that matter for AI visibility are 8.7 words long, and they are getting longer. Building a strategy on "prompts are long, think conversational" therefore repeats a reading error.
What the engine actually does: query fan-out
Definition of query fan-out: the mechanism by which a generative engine breaks the user's question into subtopics and issues several web searches of its own, whose results then feed its answer.
Those are not our words, they are Google's, which describes AI Mode as "breaking down your question into subtopics and issuing a multitude of queries", and Deep Search as able to issue "hundreds of searches" (Google).
On the ChatGPT side, Nectiv extracted those machine queries across more than 8,500 commercial prompts in nine verticals: 31% of prompts trigger a search, 2.17 searches per searching prompt, 5.48 words per query, and local searches most (59%) (Nectiv). A year later, on roughly 4,000 replayed prompts: 7.61 queries per prompt, a maximum of 29, and 64% of queries use a site: operator (Nectiv).
Hold on to that last figure, we come back to it: a site: query interrogates a domain already chosen. The engine is not looking for who could answer, it is checking someone it already has in mind.
The keyword has therefore gone from an input you picked to an output of the engine you never see, and which appears in no Search Console. Hence the offering built around it: platforms promising to "turn an AI query into keywords". That promise assumes two things. That these queries are stable for a given intention, and that they exist at all. We checked both.
Our test: 3 intentions, 8 paraphrases, 224 calls
Why Gemini. It is, to our knowledge, the only mainstream engine whose official API returns its own search queries: generateContent with the google_search tool exposes groundingMetadata.webSearchQueries, the exact list of queries issued, and groundingChunks[], the sources kept (Gemini documentation). ChatGPT's queries are visible in the JSON returned by its web interface, but no official API exposes them. Every figure below is therefore measured on Gemini, and we do not generalise it to other engines.
The material. Three commercial intentions in French, chosen to cover three different search surfaces: a CRM for a small business (B2B SaaS), the best Neapolitan pizzeria in Paris open on Sunday evening (B2C local), an electric bike for a 15 km daily commute at around €2,000 (B2C considered purchase).
Each intention has eight paraphrases of 21 to 25 words (mean 21.4) preserving every constraint: direct question, lexical synonyms, reversed constraint order, informal register, imperative form, narrative context first, negatively phrased constraint, and a paraphrase of the head term itself. The last one is reported separately: it is the hardest test, and a sceptic could argue it is not a true paraphrase.
The control group, which is the whole point. Each prompt is replayed 3 times (5 times in a second power run). Comparing two paraphrases means nothing unless you know how much two runs of the same prompt already differ. Without that baseline, you publish sampling noise and present it as an effect. It is the same requirement that led us to take apart a popular AEO case study: a number with nothing to compare it to says nothing.
The hypotheses, with their refutation criteria, fixed before collection:
| # | Hypothesis | Refuted if | Verdict |
|---|---|---|---|
| H1 | Equivalent paraphrases produce different query sets | paraphrase Jaccard > 0.80 | holds (0.011) |
| H1-bis | That divergence exceeds the engine's non-determinism | gap with the control < 0.10 | refuted |
| H2 | Paraphrasing moves queries well beyond noise, but not cited sources | difference-in-differences < 0.10 | refuted |
| H3 | A substantial share of prompts triggers no search | 0% of calls without search | holds (49%) |
We had committed to publishing nothing from a hypothesis that failed. Two failed. What follows explains why they failed, and why the conclusion comes out stronger.
Result 1: machine queries are not reproducible
We set out to show that two equivalent phrasings produce different queries. That is true, trivially: the Jaccard index between the queries of two paraphrases is 0.011, essentially no overlap.
Except the control tells a different story.
| Measured on calls that searched | gemini-3.8-flash | gemini-2.5-flash |
|---|---|---|
| Query Jaccard, same prompt replayed | 0.085 | 0.030 |
| Query Jaccard, between paraphrases | 0.011 | 0.010 |
| Gap (publication threshold: 0.10) | 0.074 | 0.020 |
Rerun the identical prompt: 91 to 97% of the queries change. On gemini-2.5-flash, 62 searching calls produced 237 distinct queries, which is nearly one fresh query per call. With 70 control pairs on the pizzeria intention, the gap falls to 0.049, 95% confidence interval 0.018 to 0.073. On both models the entire interval sits below our 0.10 threshold.
H1-bis is therefore refuted, and that is a better argument than the one we were after. We wanted to say "rephrase, and the keywords change". The reality is harsher: change nothing at all, and the keywords change anyway. The "AI keyword" is not a property of the query, nor of the intention. It is an artefact of the execution.
The direct operational consequence: there is no "right phrasing" from which to extract truer keywords, and a list extracted from one call is not reproducible by the person you present it to. If you buy a fan-out report, you are buying one draw.
At a different level of measurement, paraphrasing does move something: at the word level rather than whole strings, the gap rises to 0.115 on local intent. What those queries share are the prompt's constraints (Sunday evening, opening hours, price, booking), not shared queries.
Result 2: half the time there is no query to extract
On gemini-3.8-flash, 49% of calls issued no search query at all. The breakdown by intention is telling:
| Intention | Calls with no search |
|---|---|
| CRM for a small business (B2B SaaS) | 24 of 24 |
| Electric bike (considered purchase) | about half |
| Pizzeria in Paris (local) | 0 of 24 |
On the B2B intention the model never searched. It answered from its weights, roughly 4,000 characters of recommendations, naming products, without consulting a single web page. There is no keyword to extract there, no source to influence in the moment, and no trace to audit. Visibility is settled entirely upstream, in what the model memorised during training.
The rate depends on the model: on gemini-2.5-flash only 1.6% of calls skip search. So "half of prompts never search" cannot be stated as a general law. "On the model users have today, one call in two" can. And this specific point converges with what Nectiv observes on ChatGPT, where 31% of commercial prompts trigger a search, and where local is the vertical that searches most.
This is the first conceptual flaw in fan-out offerings: they measure the retrieval path, and on a substantial share of prompts that path does not exist.
Result 3: on local intent, the fan-out verifies a shortlist, it discovers nothing
Here are queries actually issued on the pizzeria intention:
peppe pizzeria paris dimanche soir reservation
popine paris dimanche soir reservation
guillaume grasso pizzeria paris dimanche soir
agata pizzeria paris reservation dimanche soir
"da michele" bastille paris pizza prix dimanche
94% of that intention's queries (66 of 70) name a specific restaurant or a ranking. Strip the name and the remaining words are the prompt's own constraints. On the bike intention, 57% of queries name a brand or model.
These queries are not looking for "where to eat Neapolitan pizza in Paris". They are checking whether Peppe is open on Sunday evening. Which leaves the question of where the names come from.
We replayed the same prompt with no search tool at all, with the model's thought summaries enabled. With no web access, the model already names twelve venues: Peppe, Popine, Guillaume Grasso, Dalmata, La Manifattura, Agata, Iovine's, Popolare, Bricktop, Da Michele and others. Turn search back on: all five venues queried belong to that list, and across the full 72-call run, 9 of the 12 venues queried are in it. One thought summary, during a call with search enabled, reads: "First, I need to consider my top candidates."
So the fan-out formula on local intent is:
machine queries = [the shortlist the model already has] × [the prompt's constraints]
And here the argument becomes structural rather than statistical. Even if machine queries were perfectly stable, extracting them would teach you nothing about demand. You would be reading the candidate list the model already held, filtered by the user's constraints. A ranking of recurring n-grams in those queries, "reviews", the current year, "comparison", describes the model's behaviour. It is not an inventory of intents.
The 64% of site: queries measured on ChatGPT say exactly the same thing by another route: interrogating a named domain presupposes having already chosen it.
Which moves the working question. It is no longer "which keywords to target", it is "how do we get into the shortlist", meaning: exist in what the model memorised, and in the sources it consults to check itself.
What we set out to prove, and what turned out to be false
Our starting hypothesis was elegant: queries diverge but cited sources converge, so the intention is the stable object and the lexical path is just noise. It is false, and how it failed is worth telling.
The naive test seemed to confirm it. Cited domains do overlap far more than queries: around 0.23 to 0.34 against 0.01. But that test discriminates nothing. We ran it on simulated data where every paraphrase draws its queries at random from one shared pool, so with no effect to detect at all: domains come out at 0.47 against 0.05 for queries, and the test declares "confirmed". The reason is mechanical: there are far fewer possible domains than possible query strings, so domains overlap more even under pure noise.
So we replaced the test, before running the study, with a difference-in-differences: the paraphrase effect on queries minus its effect on domains, each measured against the noise of the same prompt replayed. Verdict: −0.018 on one model, −0.068 on the power run, 0.010 on the other model. Paraphrasing moves sources at least as much as queries. H2 is refuted.
Worse for the elegant hypothesis: on local intent, paraphrasing moves the shortlist itself. The overlap of restaurants named per call goes from 0.555 between two runs of the same prompt to 0.395 between paraphrases. The sources follow. This is not "stable answer, unstable path", it is the reverse: query strings are noise, and where phrasing acts, it acts on the answer.
We will therefore never write "sources converge while queries diverge". Our own data refute it. And that nuance has an uncomfortable commercial consequence: on local intent, how your customer phrases the question changes the list you need to be on.
What remains true of SEO
None of the above says SEO is useless. GEO (Generative Engine Optimization) is not a replacement, it is an extension, and two floors remain entirely technical.
The retrieval floor. No crawl, no citation. When the engine checks its shortlist it reads pages, and yours need to be reachable and readable by AI engine crawlers, which do not have the same capabilities as Googlebot (crawler comparison).
The extraction floor. The queries we observed are constraint checks: opening hours, price, availability, booking, compatibility. If the answer to a constraint sits in an image, behind a JavaScript tab, or buried mid-paragraph, the engine does not find it and the candidate drops off the list.
What changes is the object of the work. Not a list of strings to target, but a presence to build and an absence to measure.
Action plan
- Stop buying AI keyword lists. If a report presents fan-out queries as keywords to target, ask for the control: the same measurement replayed twice on the same prompt. Without that number, the report describes one draw.
- Work in intentions, not strings. An intention is a set of constraints (who, where, when, how much, instead of what). List your customers' constraints, not their phrasings.
- Sample instead of measuring once. Since two runs of the same prompt diverge, a one-off measurement of your visibility is unreliable: GEO visibility is a distribution, not a score.
- Measure per model, not on average. Our two models do not behave alike: 49% of calls without search versus 1.6%. A visibility figure aggregated across engines hides the essential.
- Treat entering the shortlist as the objective. Which means: mentions on third-party sources the engine consults, consistent product naming everywhere, presence in your category's rankings and comparisons.
- Make every constraint checkable in one line. Opening hours, price, coverage, compatibility, lead times, in text, on a crawlable page. That is what the engine comes to verify.
- Watch the intentions that never search. On your B2B topics, test whether the engine consults the web at all. If not, your lever is the model's memory, which builds over months.
Traaker measures that presence per engine, per intention and over time, with the sources actually cited. See how the platform works.
Frequently asked questions
What is query fan-out? It is the mechanism by which a generative engine breaks the user's question into subtopics and issues several search queries of its own before answering. Google describes it for AI Mode, and Nectiv measures 7.61 per prompt on average on ChatGPT, with a maximum of 29.
Can you extract the keywords an AI uses and work them like SEO keywords? No, for three measured reasons. Those queries are not reproducible: two runs of the same prompt share only 3 to 9% of them. A substantial share of prompts issues no query at all (49% on the model we tested). And on local intent, 94% of queries name a candidate the model already had in mind, so they describe its beliefs rather than demand.
Does that mean keywords are dead? The keyword as a buyable targeting unit is dead: 65 to 85% of prompts match no keyword in Semrush's 27-billion base. Vocabulary itself stays useful for a different reason: the engine verifies constraints, and those constraints must be written in plain text on your pages.
Why do "best pizza Paris" and "number 1 pizzeria in the capital of France" return the same answer? Because matching no longer happens between strings but between vectors. The engine encodes the query and the indexed content into the same semantic space and compares positions, not words. Two phrasings of one intention therefore land in the same area. One caveat though: this does not make phrasings perfectly interchangeable. Our data show that on local intent, paraphrasing moves the list of candidates kept (overlap 0.555 versus 0.395).
Do you still need technical SEO to be cited by AI? Yes. No crawl means no retrieval, and therefore no citation. The queries we observed are checks on precise constraints: if the answer is not machine-readable, the candidate drops off the list. Technical SEO has become a floor, no longer a differentiator.
Do these results hold for ChatGPT and Perplexity? They are measured on Gemini, the only engine whose official API publishes its own queries. Only two points converge with published ChatGPT data: the frequent absence of search, and local being the vertical that searches most. Existing ChatGPT data measure neither reproducibility nor paraphrase effects, so we do not transpose the rest.
What should you track instead of a keyword ranking? Presence: on what share of a sample of intentions are you cited, by which engine, with which sources, and where are you absent. That is a distribution to sample over time, not a position to check.
Paper, data and code
- Full paper (7 pages: method, verdicts, limitations): Paraphrase and Fan-Out: What Gemini's Queries Reveal (and Don't Reveal) About Intent
- Raw data and scripts: ZIP archive containing the 224 call records (prompt, variant, run, queries, domains, answer, model, timestamp), the no-tool probe, the protocol, the collection script and the analysis script.
Replaying the study takes a Gemini API key, Node, and about four minutes per 72-call run. Verdicts may differ from ours: engine behaviour changes over time, which is exactly why we publish the files.
Method note and limitations. This is a reproducible demonstration of mechanism, not a population estimate: 3 intentions, 24 French prompts, 224 calls, 2 models, one geography and one point in time (1 October 2026). One engine, Gemini, because it is the only one whose official API publishes its queries; no mechanical transposition to ChatGPT or Perplexity.
webSearchQueriesis the engine's own declaration about itself, not a network capture, and the API with thegoogle_searchtool is not the consumer Gemini application. Thought summaries are written by the model, and the API does not timestamp them against searches: the order between candidate selection and search is therefore inferred, not observed. The semantic equivalence of the paraphrases is a human judgement, made auditable by an imposed constraint set and a named transformation grid, not proven. The entity lexicon used to classify queries was built by hand from the raw data. The CRM intention provides no data for H1 and H2 on the primary model since it never searched, and the bike intention has few control pairs. Intervals are 95% confidence intervals obtained by bootstrap over variants. Descriptive measurement, no causality claimed.