Trending
August 17, 2026

Inside ChatGPT's Retrieval Stack: The Index, the Cache, and the 1.2% It Actually Reads

A 26,900-page study maps the three stores ChatGPT grounds answers from. The findings reorder almost every GEO priority list currently in circulation.

13 min read
Updated August 17, 2026

Key Takeaways

  • A study of 1,249 ChatGPT answers, 88,000 search results and 26,900 pages maps the retrieval stack to three tiers: a discovery index, a globally shared read cache, and rare live page opens.
  • Of 61,332 retrieved URLs, only 759 pages — about 1.2% — were actually opened and read. Nine answers in ten open nothing at all.
  • Being retrieved earns a citation roughly 7% of the time. Being opened earns one roughly 74% of the time.
  • Your practical grounding budget in instant mode is the title plus about 150 characters starting at your H1 — not your meta description, and not the depth of your body copy.
  • The cache is keyed to the URL, not the user, ignores noindex, and shows no observed expiry — so a stale copy of your page can serve strangers for months.

What Happened: Someone Finally Measured the Plumbing

The ChatGPT retrieval stack has been a black box that everyone optimized for and nobody had measured. That changed with a study from French SEO consultancy RESONEO, published in July 2026 and updated this month, which captured and analysed 1,249 ChatGPT answers, 88,000 search results, 26,900 distinct pages and 6,400 domains to answer a single question: when ChatGPT cites a page, where did that page come from?

The corpus split 682 answers in instant mode against 567 in thinking mode, tested across free and paid accounts in several countries. Search Engine Land picked the study up on August 17, and it has been circulating in the French SEO press since early August. The work was done with collaboration from Jérôme Salomon at Oncrawl.

The headline finding is structural. ChatGPT does not have “an index” in the way SEOs use that word. It has three separate stores, they hold different things, and they are populated by different bots on different schedules.

The three tiers of ChatGPT's retrieval stackThree stacked layers. Tier one is a discovery index holding a roughly 200 character snippet frozen at crawl time, written by OAI-SearchBot and queried by keyword. Tier two is a read cache holding the full page as Markdown, written by ChatGPT-User and queried by exact URL. Tier three is a live fetch, where the model opens the page itself, which happens in barely more than one case out of eighty.ChatGPT grounds answers from three different storesMost SEO advice is aimed at the layer it reaches least often.TIER 1Discovery index~200-char snippet, frozen at crawl timewritten by OAI-SearchBot · queried by keyword83.6%of snippets start at the H1TIER 2Read cachefull page stored as Markdownkeyed by URL, shared across all usersmonthsretention, no expiry observedTIER 3Live fetchthe model opens your page directlyrare — but cited 74% of the time1 in 80answers open a page at allSource: RESONEO, 1,249 ChatGPT answers · July 2026, updated August 2026

That structure matters because almost all published GEO advice implicitly assumes the model is reading your page. In the overwhelming majority of answers, it is not.

The Three Tiers, and What Each One Actually Stores

Tier 1: the discovery index (“labrador”)

OpenAI's own index, written by the OAI-SearchBot crawler and queried by keyword. It does not hold your page. It holds a snippet of roughly 202 characters, frozen at crawl time, taken from the rendered body rather than from your metadata.

RESONEO reviewed 534 cited pages and found the snippet begins at the H1 in 83.6% of pages that have H1 markup. Publication dates and alt text consume part of the remaining budget. Meta descriptions rarely appear in this pipeline at all.

Tier 2: the read cache

A store of complete pages converted to Markdown, written by the ChatGPT-User agent and queried by exact URL. Three properties of this cache deserve more attention than they have received:

  • It is shared globally. The key is the page address, not the user. RESONEO watched two unconnected paid accounts request the same fingerprinted test page — the second account received the first account's cached copy without the server ever being hit.
  • It does not appear to expire. Retention ran to months in the observed data, with a freshness window of about 30 minutes below which no re-check happens at all.
  • It ignores noindex. In the researchers' words, forbidding indexing does not keep you out of the store.

Tier 3: the live fetch

The model opens and reads your page directly. This is the tier every GEO checklist is written for, and it is by far the rarest. RESONEO's summary is blunt: the page itself is opened in barely more than one case out of 80.

In the July data, free instant mode opened zero pages per answer. Paid thinking mode opened roughly one page per six conversations. The August update, after the Think button reached free accounts, moved that to 0.28 pages per conversation for free Think users — but the median across all modes stayed at zero. Nine answers out of ten open nothing.

Pro Tip

If your AI-visibility tooling reports 'ChatGPT read our page', check what it is actually measuring. A ChatGPT-User hit in your server logs means a cache write, which may then serve hundreds of later answers with no further requests. Counting server hits will dramatically undercount your real exposure.

Why 61,332 URLs Became 759 Reads

The most useful number in the study is not any single tier — it is the drop-off between them. Across the corpus, 61,332 URLs entered grounding. 7,616 (12.4%) were attached to a citation. 5,032 (8.2%) were visible as a primary source. And 759 (1.2%) were opened and read.

From 61,332 retrieved URLs to 759 pages actually readA funnel chart. Of 61,332 URLs retrieved across the study corpus, 7,616 were attached to a citation, which is 12.4 percent. 5,032 were shown as a primary source, which is 8.2 percent. Only 759 were opened and read, which is 1.2 percent.Retrieved is not read, and read is not citedThe drop-off between each stage is where GEO effort is usually wasted.URLs retrieved61,332100%Attached to a citation7,61612.4%Shown as primary source5,0328.2%Opened and read7591.2%Source: RESONEO analysis of 88,000 search results across 26,900 distinct pages and 6,400 domains

Now pair that with the conversion rates at each stage, because this is where the strategy lives. RESONEO found that a page merely present in the grounding URLs ends up cited about 7% of the time. A page the model actually opens ends up cited about 74% of the time.

That is a tenfold difference in outcome between two states most SEO teams cannot currently distinguish in their reporting. “We appeared in ChatGPT's retrieval set” and “ChatGPT opened our page” are not degrees of the same success. They are almost different businesses.

The utm parameter tells you the wrong story

RESONEO observed that the citations worth the most never carry the utm_source=chatgpt.com parameter — it was absent from pages the model opened. If you are measuring ChatGPT performance primarily through that parameter in your analytics, you are looking at a filtered view that systematically excludes your strongest placements.

And It Is Mostly Not Bing — It Is Scraped Google

The second surprise concerns which engine feeds the retrieval set. The long-standing assumption, grounded in the OpenAI–Microsoft commercial relationship, was that ChatGPT search runs on the Bing API. RESONEO's finding on that is unambiguous: they state they never observed the Bing engine, on any account, in any regime.

What they observed instead was a split that shifts by mode and by month. In July, free instant mode ran 99.9% on OpenAI's own index for stable information, dropping to 52.6% index and 46.1% Google for news queries. Paid thinking mode ran 99% on scraped Google results delivered through commercial providers Bright Data and Oxylabs. By the August revision, free Think sat at 74.7% index and paid thinking at 75.3% scraped Google.

This corroborates a test Aleyda Solís ran earlier, in which ChatGPT answered from a Google SERP snippet word for word before Bing had indexed the page at all — and kept using the Google snippet even after Bing caught up.

The practical read: for a large share of ChatGPT queries, your Google ranking and your Google SERP snippet are your ChatGPT visibility. That is an uncomfortable conclusion for anyone selling GEO as a discipline entirely separate from SEO — and a reassuring one for teams who never stopped doing the fundamentals.

What the Researchers Are Saying

The most quotable line from the study is also its most actionable, and it reframes what “optimizing for ChatGPT” means at the page level:

Your grounding budget in instant mode is your full title plus roughly one hundred and fifty useful characters starting at your H1. Mark up an H1, clear the runway between it and the first paragraph, and keep an eye on the alt text of the first image.

RESONEO, ChatGPT retrieval study, July 2026 (updated August 2026)

This work builds on research by Suganthan Mohanadasan, who in June 2026 read ChatGPT's raw network traffic rather than its outputs and found a result_source field tagging every web result as one of four values: serp, labrador, bright or oxylabs. He also estimated that around 30% of current-event queries bypass live search entirely and are answered from training data.

Notably, Mohanadasan initially characterised labrador as a licensed-publisher tier — “largely closed to smaller sites” — and then publicly revised that assessment after receiving user captures showing small Italian publishers appearing through identical pipelines. The larger dataset supports the revision: there is no observed difference in how OpenAI's in-house index serves licensed and unlicensed sites.

Where licensing does bite is in the word budget applied once your content is used. RESONEO documented three tiers: about 200 words per source as the open-web default, about 100 words for licensed publishers including Le Monde, WSJ, Politico and Condé Nast, and about 25 words for a restricted set including the Guardian, ESPN and the Washington Post. Being licensed buys a publisher money and constraints, not visibility.

One honest caveat from the researchers: on July 21 OpenAI removed the result_source field from the response stream, so engine attribution after that date is inferred rather than read directly. They also note that API testing showed 61–71% multi-turn iteration against 0–17% in the product, meaning model behaviour and product behaviour are not the same thing. Treat the percentages as a well-evidenced snapshot of a moving system, not as constants.

Which Layer Does Your GEO Tactic Actually Touch?

This is the part worth arguing about. Once you accept that the three tiers are distinct, every tactic on your GEO list can be sorted by which store it reaches — and several popular ones turn out to reach none of them. Here is that mapping against the study's findings.

TacticLayer reachedVerdict
H1 wording and the first ~150 characters after itDiscovery indexHighest-leverage change available. This is literally the text ChatGPT stores about your page.
Meta description rewritingNone of themRarely appears in this pipeline. Keep it for SERP click-through, not for ChatGPT.
Long-form depth and comprehensive body contentRead cache + live fetchOnly pays off once you are already being opened or cached — which is the minority of retrievals.
Publishing frequency and 'freshness' languageNone reliablyFreshness keywords such as 'today' and 'now' produced no measurable effect in the corpus.
Alt text on the first imageDiscovery indexIt competes for the same ~200-character budget. Sloppy alt text can eat your snippet.
Page weight and template bloatRead cachePast 4MB the page is rejected outright — no partial read, no truncation.
Server-rendered HTML vs client-side renderingAll threeEvery tier stores what comes back in the response. Content that needs JS to appear may never be stored.
noindex to control AI exposureDoes not applyPages carrying noindex still land in the cache. Use authentication if content must stay out.

The pattern is consistent. Work that shapes the first few hundred characters of your rendered page has outsized effect, because that is the unit ChatGPT stores. Work that deepens the body has effect only in the minority of cases where the page is opened or cached. And work aimed at metadata that this pipeline does not read has no effect here at all.

What to Do About It This Week

None of this requires a replatform. It requires re-pointing effort at the layer that actually gets read.

Step 1: Audit the first 200 characters after every H1

Pull your top 50 pages by AI-search relevance and read only the H1 plus the first two sentences of rendered body text. That fragment is what ChatGPT stores about you. If it opens with a breadcrumb, a byline block, a cookie notice, an image caption, or a throat-clearing sentence about how important the topic is, you are spending your entire grounding budget on nothing.

The fix is unglamorous: one H1 per page, a direct declarative first sentence that answers the query, and no interstitial furniture between them. Our free Heading Analyzer will flag pages with missing, duplicated, or badly nested H1s in one pass.

Pro Tip

Check the alt text on your first image while you are in there. It competes for the same ~200-character budget, so a decorative hero image carrying alt text like 'blog-header-final-v3-1920x1080' is actively displacing text you wanted stored.

Step 2: Stop optimizing meta descriptions for AI answers

Meta descriptions rarely enter this pipeline. They remain worth writing for click-through on traditional SERPs, and you can check how yours render with our SERP Preview or generate a clean set with the Meta Tags Generator. Just move them out of the “GEO work” column on your roadmap, where they have been quietly absorbing hours.

Step 3: Treat the cache as a publishing risk

A cache that is keyed by URL, shared across all users, and shows no observed expiry means an outdated version of your pricing page, your product spec, or your correction-worthy claim can serve strangers for months after you fixed it. There is no cache-purge button for this.

The practical mitigations are the ones you already know but may not have prioritised for this reason: change the URL when the substance of a page changes materially, keep time-sensitive claims out of the opening 200 characters where they get frozen into the snippet, and audit your highest-stakes pages for accuracy on a schedule rather than reactively.

Step 4: Confirm your content exists without JavaScript

All three tiers store what comes back in the HTTP response. If your H1 and opening paragraph are injected client-side, there is a real chance none of the three stores ever holds your actual content. Fetch your key templates as raw HTML and confirm the first 200 characters of meaningful text are present in the source. A run through the SEO Content Grader gives you a fast read on structure and on-page signals at the same time.

Step 5: Fix your bot policy on purpose, not by default

The three tiers are written by different agents. OAI-SearchBot populates the discovery index; ChatGPT-User populates the read cache. Blocking the wrong one removes you from a layer you wanted to be in, and blocking either does nothing about training crawlers. Check what your current rules actually say with the Robots.txt Checker before you change anything — and read our companion piece on what AI bots are actually doing on your site for the full agent-by-agent breakdown.

Do not reach for noindex as an AI control

The study found pages carrying meta robots noindex still ending up in the read cache. Noindex governs whether a search engine lists your page in its results — it is not an instruction that AI retrieval systems are bound to honour. Content that genuinely must not be retrievable belongs behind authentication, not behind a meta tag.

Tools to Help You Audit the Layer That Gets Read

The work above is mostly structural inspection, which is exactly what free tooling is good for. These are the ones that map to the steps in this article.

What to Watch Next

The engine mix is the most volatile number here. It moved substantially between the July capture and the August revision, and the removal of the result_source field on July 21 makes it harder for anyone outside OpenAI to track. Expect the index-versus-scraped-Google balance to keep shifting, and treat any single percentage as a snapshot.

The second thing to watch is the live-fetch rate. It rose for free users when the Think button arrived, from zero pages per answer to 0.28 per conversation. If reasoning modes keep expanding to more users by default, the tier that converts at 74% becomes structurally more important — and the balance of this advice would shift toward body content again. That is a change worth re-measuring rather than assuming.

Finally, watch the cache. A globally shared, URL-keyed store with no observed expiry and no purge mechanism is an unusual thing to have in the content supply chain. If publishers begin pressing for cache controls the way they pressed for crawler controls, that is where the next round of standards fights will happen.

Frequently Asked Questions

Key Takeaways

The useful thing about this study is not that it makes ChatGPT visibility easier. It is that it makes it legible. For two years, GEO advice has been assembled from output observation and inference. This is a mechanism — three stores, different bots, measurable conversion rates between stages — and a mechanism lets you rank your work by expected return instead of by plausibility.

Your Action Plan:

  • Read the H1 plus first 200 rendered characters of your top pages, and rewrite anything that wastes that budget.
  • Move meta-description work out of your GEO column and into your SERP click-through column, where it belongs.
  • Verify your opening content renders server-side, and check no critical template approaches the 4MB rejection threshold.
  • Change the URL when a page's substance materially changes, since the cache has no purge and no observed expiry.
  • Separate OAI-SearchBot and ChatGPT-User in your bot policy, and stop treating noindex as an AI control.

One closing note on the biggest finding of all: for a large share of queries, ChatGPT is reading scraped Google results. Which means the least fashionable answer is also the best supported one — ranking well on Google remains the most reliable way to show up in ChatGPT. Browse our free SEO tools or the rest of our AI search coverage to keep that side of the work sharp.

Related Free SEO Tools

More Articles