Research · GEO Measurement

AI Answers Change 74% Day-to-Day. Here's the 13% That Doesn't.

Glen Allsopp's 70K-response study at Detailed is the most useful GEO measurement research published in 2026. It quietly settles a debate we've been having with clients for a year: whether AI-search visibility should be measured like rankings or like brand awareness. Neither. The right unit is core-vs-tail, and it reshapes what a GEO program should optimize for. Here's our reading of the study, the numbers we think matter most, and how our GEO KPI framework changes because of it.

Quick answer. In a 28-day study of 70,000+ ChatGPT and Google AI Overview responses to 1,300+ prompts, Detailed’s Glen Allsopp found that the top brand for a typical prompt shows up on 82–89% of days, while the tail brands turn over at 74–78% per day. Only 0.3% of day-to-day comparisons produced an identical brand list on ChatGPT; only 2.4% produced an identical set of cited domains. The right takeaway for anyone running a GEO program is not “AI answers are too noisy to measure.” It is: stop measuring the tail. The metric that matters is whether you are inside the 1–2 name core that gets cited on 80%+ of days for a specific prompt. Everything else is noise engineered by the model’s randomness. This piece walks through the study, the numbers we think matter most, three findings that reshape how GEO programs should be budgeted, and how our own 5-KPI GEO framework changes because of it.

Table of contents

  1. What Detailed measured, in one paragraph
  2. The core-vs-tail rule and why it settles a long-running debate
  3. Three findings that reshape GEO budgets
  4. The category effect: one AI-visibility strategy doesn’t work across industries
  5. ChatGPT and AI Overviews are different channels, not one channel
  6. The brand-renaming problem: most tools are undercounting your presence
  7. What we’re changing in our own GEO KPI framework
  8. The honest limits

What Detailed measured, in one paragraph

Glen Allsopp of Detailed (now an Ahrefs company) ran the same 1,300+ prompts once per day, from a US, non-logged-in account, against ChatGPT and Google AI Overviews for 28 days between 26 August and 22 September 2026. That works out to about 70,000 total responses. Prompts were commercial-leaning, spread across ten categories (Agency, Beauty, Best Products, Fashion, Finance, Health, Marketing, Software, Tech Stack, plus a mixed “AI Visibility” bucket). Named brands were extracted with a custom fine-tuned OpenAI model with manual sampling for accuracy. All measurements are on that daily-snapshot basis.

This is the largest publicly-documented cross-platform AI answer volatility study we know about. Everything below is our reading of Glen’s data through the lens of running GEO programs for enterprise brands.

The core-vs-tail rule and why it settles a long-running debate

The single most useful finding in the study is not the volatility number. It is the discovery that AI answers are not uniformly volatile. They have a core and a tail, and the two behave completely differently.

Glen sorted every brand named across the 28 days into two groups per prompt: core brands, named on 80%+ of days, and the tail, everything else. The day-to-day change rate for each:

13%
Core brand change · ChatGPT
Percentage of core brands that dropped or newly appeared between adjacent days.
78%
Tail brand change · ChatGPT
Almost every tail brand rotates in and out day to day.
11%
Core brand change · AI Overviews
Even more stable than ChatGPT.
65%
Tail brand change · AI Overviews
Slightly more stable than ChatGPT's tail, but still churning fast.

Cited domains follow the same pattern (14% vs 74% on ChatGPT, 12% vs 61% on AI Overviews).

This is the number that settles a debate we’ve been having with clients for a year.

Camp A says AI visibility should be measured like search rankings: position, count of mentions per day, moves up and down. The problem: on the exact-list-match test, ChatGPT returned identical brand lists on only 0.3% of day-to-day pairs. If you’re measuring positions on daily snapshots, 99.7% of your “movements” are model randomness.

Camp B says the volatility means AI visibility is unmeasurable, so you should track it lightly if at all, and focus on organic search. That’s a defensible position for a small program. But at enterprise scale it’s an abdication — AI Overviews now render on a majority of commercial queries and ChatGPT alone has hundreds of millions of monthly active users.

Detailed’s data suggests the third answer, which is what we now recommend: measure whether your brand is in the core, not where it ranks. Core brands change at 11–13% per day, which is a signal-to-noise ratio you can actually work with. Tail brands change at 65–78% per day, which is essentially noise. If you’re not in the core for a prompt today, the right strategic question is not “how do we move up in the tail,” it’s “what would it take to enter the core for this prompt.”

This is a different mental model than SEO. In rankings you can be at position 15 and know that position 8 is your realistic next stop. In AI answers there is a small stable core, a very large noisy tail, and almost no meaningful positions in between.

Try the tool · Free · No signup

Are you in the core, in the tail, or missing entirely?

We built a free tool that operationalises the framework above. Enter your brand, name variants, category, and 1–3 buyer-language prompts. Claude simulates what an AI assistant would answer for each. In about 20 seconds you get a scorecard showing presence rate, primacy rate, category benchmark, and a core-vs-tail prediction — one prompt at a time, with the brands Claude would actually list next to yours.

Check my visibility

Three findings that reshape GEO budgets

1. The core is small, so being in it is worth more than it looks

Detailed reports that only about half of ChatGPT prompts that named brands had any core brand at 80%+ presence. When a core did exist, it was usually one or two names. For AI Overviews the numbers were slightly better (~60% of prompts had a core, still typically one or two names).

Read that carefully. For roughly half of commercial prompts, there is no brand cited stably enough to count as core. The category is wide open. For the other half, there is usually one and sometimes two brands. If you’re one of them, you own the answer. If you’re not, no amount of tail optimization changes the outcome.

The implication for budgeting is unpleasant: GEO investment does not scale linearly with keyword count. Investing across 200 prompts where you can never enter a core of one produces less commercial return than concentrating investment on 20 prompts where you can realistically become that core name in 12 months. This is the opposite of how most SEO programs are budgeted, where broad long-tail coverage often does compound.

2. Being cited first is a different metric than being cited at all

For each prompt, Detailed also tracked how often the usual leader — the brand most often named first — actually held that first position. On ChatGPT it was named first on only 50% of days but appeared somewhere in the answer on 67%. On AI Overviews it was named first on 60% of days and appeared on 78%.

That’s a gap of 15–20 percentage points between “leader appears” and “leader is first.”

For B2B categories, this matters because the first-named brand gets a disproportionate share of click-through when the user goes on to check any of the recommendations. For consumer categories, users are more likely to scan multiple options.

Our advice to clients has been to report AI-search visibility as two metrics per platform: presence (in the answer at all) and primacy (named first). Detailed’s data is the first public confirmation we’ve seen that these two behave differently enough to be reported separately.

3. The typical prompt shows exact-list match under 3%, but that’s not the right threshold to worry about

Detailed’s strictest test: pick two random days for a prompt, and check whether the full brand list matches exactly.

0.3%
ChatGPT · exact brand-list match
Two random days out of 28 · same brands, same order.
2.4%
ChatGPT · exact domain-set match
Two random days · identical set of cited URLs.
1.1%
AI Overviews · exact brand-list match
More consistent than ChatGPT on brands.
41%
Prompts · never repeated brand list once
4 in 10 ChatGPT prompts never once returned the same brand list in a month.

The right reading of these numbers is not that AI answers are broken. It is that exact list match is the wrong stability metric. If ChatGPT names 10 brands and 9 of them match the next day, that counts as a non-match, which understates real stability. A more useful stability metric is core-brand overlap rate, which is what the 11–13% change figures actually represent.

For anyone building or evaluating a GEO tracking tool, this is a design decision worth making explicit: report core overlap, not exact match.

The category effect: one AI-visibility strategy doesn’t work across industries

The category-level table in Detailed’s study is where the practical implications get sharpest. Consistency of the first-named brand across the 28 days varied by category:

CategoryFirst brand named first (ChatGPT, median)Read
Tech Stack56.7%Very consistent — Zapier / HubSpot / Salesforce dominate the same slots.
Marketing49.5%Consistent — Ahrefs / SEMrush / HubSpot own the frame.
Software44.6%Consistent within sub-categories, unstable across them.
Beauty29.1%Moderate — beauty brands rotate at the top.
Best Products18.9%Noisy — publisher recommendations shift often.
Fashion12.3%Very noisy — the first-named brand almost never persists.

Two takeaways. First, the ROI of a GEO program varies by category more than most agencies (including us, until recently) have been willing to admit. If you sell financial services software, becoming one of the 1–2 core brands for “best financial services software” is a defensible, measurable, compounding investment. If you sell fashion, the same investment against “best cocktail dresses under $200” will be systematically dismantled by the model every 3 days.

Second, Health is a category on its own. Detailed notes that Health prompts almost never name brands but do produce the most stable domain citations of any category (55.8% same first domain on ChatGPT, 15% exact domain-set match). This is the model treating health information as a domain-authority question rather than a brand-recommendation question, and that behavior is remarkably stable. For healthcare marketers, the practical implication is that GEO strategy in health is not about being named — it is about being cited. Different metric, different tactics, different content.

We’re now recommending that GEO retainer scopes be sized by category, not by revenue. A $50K/quarter GEO program in Tech Stack buys more than a $50K/quarter GEO program in Fashion, because the model rewards the investment differently.

ChatGPT and AI Overviews are different channels, not one channel

This one lands hardest for the “one AI SEO strategy” pitches we still hear from generalist agencies.

The two platforms agreed on the most-present brand for the same prompt only 27% of the time.

They cite different sources by structural preference:

  • ChatGPT leans on articles (~60% of citations), category / listing pages, homepages, and documentation. Video is essentially absent (~0%).
  • AI Overviews also lean on articles (~60%), but the second-largest citation type is video at 11% (mostly YouTube), and Reddit / Quora / forums appear far more often than on ChatGPT.

Look at the top-cited domain lists Detailed publishes:

  • ChatGPT top 5: clutch.co, ahrefs.com, developers.google.com, nerdwallet.com, mayoclinic.org
  • AI Overviews top 5: youtube.com, reddit.com, semrush.com, mayoclinic.org, my.clevelandclinic.org

YouTube alone was cited 12,299 times in AI Overviews vs essentially never on ChatGPT. Reddit was cited 8,216 times in AI Overviews vs a fraction of that on ChatGPT.

The practical read: if your brand’s answer-shaped content lives only in written articles, you are optimizing for ChatGPT presence and quietly forfeiting AI Overviews presence. If you have serious AI Overviews ambitions, you need a YouTube channel with structured, extractable content — not brand videos, extraction-friendly explainer content that Google’s system can pull a 20-second answer from. This is the biggest single content-strategy difference between the two platforms and most marketing plans don’t reflect it.

The two platforms also disagree on how many brands to name and how many sources to cite:

  • ChatGPT: 5.9 brands per answer, 5.5 links from 3.7 domains
  • AI Overviews: 5.1 brands per answer, 6.8 links from 6.0 domains

ChatGPT concentrates citations on fewer domains but names more brands. AI Overviews spreads citations across more domains but names fewer brands. If you’re a brand that primarily wins through inclusion in third-party listicles (Clutch, G2, Capterra, Forbes rankings), ChatGPT rewards that pattern more heavily. If you’re a brand that primarily wins through your own owned content getting cited directly, AI Overviews rewards that.

We’re now scoping AI-search work as two parallel programs when the client cares about both, not one program that serves both. The overlap is real (both platforms cite articles heavily) but the divergence is bigger than any single content strategy can bridge.

The brand-renaming problem: most tools are undercounting your presence

This is the finding that changes how we now audit competing GEO tools.

Glen counted brand names exactly as AI wrote them. When he then manually merged obvious name variants for 55 brands (411 total names), he found that 36% of apparent brand drop-offs on ChatGPT and 30% on AI Overviews were actually the same company under a different name the next day.

“Ahrefs” and “Ahrefs Brand Radar.” “HubSpot” and “HubSpot CRM.” “Grow and Convert” and “Grow & Convert.” “Anaplan” and “Anaplan Financial Planning.”

If your GEO tool is not doing name-variant matching well, your reported presence in AI answers is likely understated by 25–35%. Which sounds like a nice thing until you realize you might be over-buying remediation work to solve a phantom decline.

Our operational recommendation to clients: at the start of any AI-visibility tracking engagement, spend a session compiling every legitimate way your brand is named — legal entity, common shortening, sub-brand, product, product+company, common misspellings that AI actually uses. That list is now a required deliverable in our onboarding, not an afterthought. Every tracked mention gets matched against the alias list before being reported as absent.

Detailed also notes that this varies by brand. HubSpot and Wells Fargo produced most of their apparent drop-offs from renaming (they have many sub-brand and product variants). Freshdesk and Anaplan were almost always named the same way (single product, single brand). Know which one you are.

What we’re changing in our own GEO KPI framework

Six months ago we published our 5-KPI GEO framework: AI-Surface Share of Voice, Citation Authority Score, Entity Recognition Coverage, Branded AI Traffic, and Assisted Conversion Lift. That framework survives Detailed’s study. But three of the five need refinement based on this data.

KPI 1, AI-Surface Share of Voice. We were reporting this as a single presence percentage per platform. We’re now splitting it into core presence (in the 80%+ stable set for a prompt) and tail presence (appeared at least once in 28 days). Only core presence gets weighted in the composite AI-SOV score. Tail presence is reported as an early-warning signal — a brand that appears in the tail but not the core is either about to enter the core or about to disappear entirely, and either signal is worth acting on.

KPI 2, Citation Authority Score. We were reporting this as an absolute weighted count. We’re now reporting it per platform — ChatGPT and AI Overviews receive separate scores that are not composited into a single number. The 27% cross-platform agreement rate means treating them as one signal averages away meaningful strategic direction.

KPI 3, Entity Recognition Coverage. Unchanged. If anything, Detailed’s study reinforces this KPI — the “different names for the same brand” problem is fundamentally an entity-recognition problem, and improving entity recognition (Wikidata, Organization schema, consistent brand naming across owned properties) directly attacks it.

KPI 4, Branded AI Traffic. Unchanged in definition, but we’re now reporting the volatility of the traffic alongside the volume. A brand pulling stable AI-referred traffic day-to-day is in the core for something valuable; a brand pulling spikey traffic is in the tail and should not be planning on the current run rate holding.

KPI 5, Assisted Conversion Lift. Unchanged. This one is downstream of all the others and Detailed’s study doesn’t directly touch it.

The bigger change is what we no longer chase. We used to include rank-style position tracking (are you cited 1st, 2nd, 3rd) as a secondary metric. Detailed’s data confirms this is mostly noise. We’ve moved it to a diagnostic view, not a reported KPI.

The honest limits

A few caveats we want to state explicitly, because they affect how confidently anyone should generalize from Detailed’s study to their own brand.

One US region, one non-logged-in account, 28 days. Personalization, location, chat history, and browsing session context all shift AI answers. The reported volatility is the un-personalized baseline. For a logged-in user with a purchase history in your category, the answer they see is likely more stable than these numbers suggest. Our clients often see internal AI-monitoring numbers that look calmer than Detailed’s, which is consistent with this.

Prompts were “commercial-leaning” but not purely commercial. Some informational queries got mixed in. Detailed acknowledges this and flags it as a future refinement. Commercial-only prompts would likely show slightly more stability, since transactional intent narrows the acceptable answer set.

Model updates during the window. ChatGPT and AI Overviews both update their retrieval and generation systems on non-public schedules. Twenty-eight days is long enough to see model behavior during an update but not long enough to isolate the effect of one. This is why Detailed’s own recommendation is to track your own space over longer windows.

28 days is not 12 months. Category leadership can shift meaningfully across quarters as new products launch and new content indexes. The “core” for a prompt today may not be the core in six months. Detailed’s data captures a snapshot, not a trend.

Brand-name automation was imperfect. As Detailed notes, name-variant deduplication was tested but not fully automated across the whole dataset. Real-world presence is likely 25–35% higher than the reported baseline for brands with meaningful sub-brand or product variants.

What we recommend now, in one line

Stop measuring AI-search visibility by rank and daily count. Start measuring by core presence: whether your brand is inside the 1–2 name group that gets cited on 80%+ of days for a specific prompt, on a specific platform, over a rolling 28-day window. Everything else is model noise engineered to look like signal. Detailed’s data is the strongest public evidence for this reframe that we’ve seen in 2026, and our KPI framework, our client audit templates, and our internal weekly reporting all now reflect it.

Full credit to Glen Allsopp and the Detailed / Ahrefs team for putting the numbers in the public domain. The original study is here — read it, don’t just take our reading of it.


This piece is written by Klara & Nadia, Resocial’s GEO measurement lead and AI-search analyst. It builds on data published by Glen Allsopp at Detailed on 23 September 2026, and on our own GEO KPI framework from May. For the strategic 10-year view that frames why this measurement work matters now, see our group’s AI Search vs Google decade research at House of 47. Want a first read of your own brand right now? Try the free AI Visibility Checker — brand, variants, category, three prompts, 20 seconds. If you want the real 28-day baseline across ChatGPT, AI Overviews, Perplexity, and Gemini, that’s the first deliverable of any AI Search & GEO engagement we take on.

Want strategy like this for your brand?

Not sure where to start?

Describe your problem and our AI maps it to the closest Resocial service.

Describe my problem

Get a free SEO audit

60+ dimensions, 48-hour turnaround.

Get a Free SEO Audit

Submit an enterprise RFP

Tailored proposal in 5 business days.

Submit an Enterprise RFP