Evidence.

Numbers on brand, AI search, and agents that you can cite in a presentation without having to take them back later. Every entry names the primary source, the sample, the date, and the limit of what it proves. Ready to copy, with attribution.

Most numbers in this field circulate without a source, with the wrong denominator, or a year too late. I need them for my own writing, so I check them against the primary source. The result is here, including where it contradicts my own position.

A number is not the same as its interpretation. So every entry states what it explicitly does not prove. Four entries exist for that reason alone: they correct figures that circulate in the field in overreached form, among them the 78-to-42-percent agent figure and the 40 percent of generative engine optimization. If you read a number differently or find an error, write to me and I will correct it and note the date.

Topic

Verified study

Of roughly 38,000 domains that have an llms.txt, 97 percent saw no request for the file at all in May 2026.

What the number does not say: It measures requests, not effect. The file remains useful for coding and browser agents. What is refuted is only the claim that AI search reads llms.txt for its recommendations. Mind the denominator: the 97 percent refer to the roughly 38,000 domains that have a file, not to all 137,210 studied. And the sample is not a cross-section of the web but the domains of one analytics vendor that had traffic in May.

Ahrefs, 137,210 domains, 28 percent of them with an llms.txt · June 2026 · Source

Permalink

Verified study

Across nearly 300,000 domains studied, no relationship was found between having an llms.txt and how often a domain appeared as a source in AI answers.

What the number does not say: No relationship is not a measurement of effect. The study compares, it does not experiment, and it limits itself to the model and dataset tested. The comparison group is also smaller than the headline number suggests: only 10.13 percent of the domains had a file at all. Its weight comes from being the second independent study with the same result as the finding above.

SE Ranking, nearly 300,000 domains, 10.13 percent of them with an llms.txt · November 2025 · Source

Permalink

Vendor documentation

Google explicitly states that llms.txt, special markup, and custom chunking are not needed for AI search.

What the statement does not say: It applies to Google Search including its generative features, explicitly not to other systems — for services that do use such files, the same text calls them harmless. And on structured data Google does not say "useless" but keeps recommending it, because it qualifies pages for rich results in classic search. From Google’s perspective, optimizing for AI search is still SEO.

Google Search Central, official guide · as of 10 July 2026 · Source

Permalink

Single test

1,885 pages that added JSON-LD barely moved against 4,000 control pages: no effect distinguishable from zero in ChatGPT and Google AI Mode, and a statistically significant 4.6 percent decline in AI Overviews.

What the test does not say: Only pages that were already heavily cited by AI were studied — each had over a hundred AI Overview citations in February 2025. The study says nothing about pages that do not appear at all, and the authors explicitly allow that markup may help there. It is an observational study with matched controls, not an experiment. For rich results in classic search the markup remains uncontested.

Ahrefs, 1,885 pages against 4,000 matched controls, markup added between August 2025 and March 2026 · May 2026 · Source

Permalink

Verified study

Ask the identical question twice and the chance of getting the same list of brands is under one in a hundred.

What follows and what does not: Anyone measuring AI visibility with a single run is mostly measuring noise. The study explicitly does not conclude that measuring is pointless: across dozens to hundreds of prompts, run repeatedly, it considers a visibility share a reasonable metric. Citing it as "tracking is useless" goes further than the evidence.

SparkToro with Gumshoe.ai, 600 participants, 12 prompts, 2,961 runs across ChatGPT, Claude, and Google AI · January 2026 · Source

Permalink

Verified study

When an AI summary appears, users click a traditional search result on 8 percent of visits. Without a summary it is 15 percent.

What the number does not say: It establishes no cause. Queries that trigger an AI summary are systematically different from those that do not, so the comparison runs between query types, not within the same query. It also shows no revenue loss and says nothing about commercial queries. The harder number sits beside it: a link inside the summary was clicked on one percent of visits. And the mix of sources in the summaries resembled ordinary search, which contradicts the popular story about concentration on Reddit and Wikipedia.

Pew Research Center, passive browser tracking of 900 US adults, 68,879 Google searches in March 2025 · July 2025 · Source

Permalink

Usually miscited

The most quoted figure in the AI visibility business, "up to 40 percent more visibility", comes from a lab setup running a since-retired model and measures a purpose-built metric.

What the number actually covers: It measures a position-weighted share of words a source occupies in a generated answer. The "generative engine" was a two-stage build of the authors’ own: fetch the top five Google results, then generate an answer with GPT-3.5. No commercial product, no current model. "Up to 40 percent" is also a maximum across the best-performing methods, not an average. It does not evidence more clicks, more revenue, or more mentions in ChatGPT or Google. The paper itself is careful and states its limits; the overreach happens in the citing.

Aggarwal et al., GEO: Generative Engine Optimization, arXiv 2311.09735, KDD 2024 · submitted November 2023 · Source

Permalink

Verified study

45 percent of answers from four AI assistants to news questions had at least one significant issue. In 31 percent it concerned the handling of sources.

What the number does not say: What was tested was news content, not brands. Citing it as evidence for how often AI misrepresents companies transfers it illegitimately. The assessments came from journalists at the participating organisations, so not an independent body, and the models date from mid-2025. The sourcing finding is the brand-relevant part: almost a third of answers attributed statements to a source that does not support them. That is precisely the mechanism by which a brand gets miscited too.

European Broadcasting Union and BBC, 22 media organisations across 18 countries and 14 languages (including ARD, ZDF, Deutsche Welle, and SRF), over 3,000 answers assessed · October 2025 · Source

Permalink

Verified study

On 30.6 percent of one million home pages surveyed, buttons had no accessible name.

What the number does not say: It was made as an accessibility survey, not an agent test. The connection is direct nonetheless: a button without a name has no label in the accessibility tree, and agents working from that tree cannot name it. Home pages were measured, not entire sites, and the share refers to pages carrying at least one such button, not to the share of all buttons.

WebAIM Million, eighth edition, one million home pages · February 2026 (29.6 percent the year before) · Source

Permalink

Preliminary, prototype

In a controlled experiment, three browser agents reached a strict success rate of 89.3 percent on the agent-friendly version against 49.3 percent on the original.

What the experiment does not say: What varied was machine clarity, not accessibility, and these are two versions of a purpose-built shop prototype, not a real website. The authors state explicitly that this is a proof of concept whose results "should not be generalized to all domains, websites, or agent systems". The figure quoted is also the stricter of two success rates measured. It is the best available indication, not a proof.

Elnaffar and Rashidi, Designing Agent-Ready Websites, arXiv 2607.12056, 300 runs across three models · July 2026 · Source

Permalink

Usually miscited

The widely cited drop in agent success from 78 to 42 percent comes from a study that changed no website at all.

What actually varied: how the agent operates, not the accessibility of the pages. The study took the agent’s mouse away, restricting it to the keyboard. As evidence that accessible websites help agents it does not hold, although it is cited for exactly that everywhere. A clean comparison of accessible against inaccessible is still missing. The numbers are narrower than their citations too: they describe a single model (Claude Sonnet 4.5, precisely 78.33 to 41.67 percent), and the tasks span desktop applications, not only websites. A second, open model fell from 20 to 0 percent.

A11y-CUA, arXiv 2602.09310, presented at CHI 2026 · February 2026 · Source

Permalink

Controlled test

Seven widely used US AI assistants execute no JavaScript on a user-triggered fetch and read only the raw HTML. Five others do execute it.

What the test shows and what it does not: The setup put a decoy value in the raw HTML and the real value behind JavaScript. ChatGPT, Claude, Gemini, Perplexity, Meta AI, Copilot, and Grok returned the decoy; DeepSeek, ERNIE, Qwen, Kimi, and Mistral returned the real value. The dividing line runs by vendor, not by technology — it is a decision, not a limit. The finding covers the direct fetch, not content that reaches an answer through the Google index. An earlier test in December 2025 still measured Gemini as the only system that rendered, so the picture moves.

Search Engine World, 12 assistants compared · June 2026 · earlier test: searchVIU, December 2025 · Source

Permalink

Vendor documentation

Since version 13.3.0 of 7 May 2026, Lighthouse ships an "Agentic Browsing" category in its default configuration.

What the score does not say: Google labels the category explicitly as experimental and based on proposed standards; it requires Chrome 150 or later, and the WebMCP audits require registering for the origin trial. It reports a pass rate, not a 0-to-100 score like performance or SEO. It measures agent readiness, explicitly not visibility in Google Search. What is notable is the direction: machine readability turns from a claim into a measured property.

Lighthouse release 13.3.0 of 7 May 2026, category in the default config · Google’s documentation · Source

Permalink

Vendor documentation

Google names three ways agents perceive a page: screenshots, raw HTML, and the accessibility tree. Modern agents combine them.

What does not follow: that agents work from the accessibility tree alone. That shortcut is exactly what circulates. Google describes the tree as a high-fidelity map that ignores visual noise, but says in the same text that agents cross-reference tree and DOM with a visual rendering. For practice this changes little: an element without an accessible name is missing from two of the three routes.

Google, Build agent-friendly websites (web.dev) · as of 1 April 2026 · Source

Permalink

State of standardization

A common standard for how websites permit or refuse AI use of their content still does not exist.

How far the work has come: The IETF working group AIPREF is developing two building blocks, a vocabulary and an attachment mechanism. At the end of August 2026 both sit at draft status "I-D Exists", meaning they have not even been submitted for approval; there is no RFC, although the group’s own milestones are dated 31 August 2026. In practice: anyone steering AI crawlers today does so through robots.txt and vendor-specific tokens that can change at any time. In parallel, Cloudflare runs its own vendor-bound mechanism.

IETF working group AIPREF, drafts of 18 August 2026 · checked 28 August 2026 · Source

Permalink

Market observation

Buying directly inside the chat was rolled back a good five months after launch. Live at that point were either a dozen or close to thirty Shopify merchants, depending on the source.

What follows and what does not: The thesis that commerce moves wholesale into the chat is refuted for now. What is not refuted is the shift in discovery: found at the agent, bought at the brand. On the number: the source gives two conflicting figures, a dozen per a press report and "closer to 30 and climbing" per Shopify itself. Both cover Shopify merchants only, not all partners.

Forrester, analysis of the instant checkout rollback · 7 March 2026 (launch: 29 September 2025) · Source

Permalink

Controlled test

In a flight-search test, language models went to the airline’s own website directly in only about five percent of cases. They preferred booking portals.

Why this matters for brands: The direct channel and the loyalty program, the most expensive distribution assets many brands own, are bypassed in the agent layer. The study’s explanation is remarkably unglamorous: portals deliver cleaner, more structured, more agent-readable data. Limits: three European carriers, their ten most relevant markets, three language models, 60 non-branded prompts. One industry, one point in time, no claim about other categories.

Bain & Company, test across three carriers, three language models, and 60 prompts · 16 March 2026 · Source

Permalink

Vendor and company figures

Walmart has supplier negotiations run by an AI system: 2,000 negotiations at once, around three percent average savings, and three out of four suppliers preferring to negotiate with the machine over a person.

Where the numbers come from: the system’s vendor and Walmart itself, reported via Bloomberg. There is no independent audit. The closing rate of 68 percent often quoted alongside comes from an earlier Harvard Business Review case study. The third number remains the striking one: the preference for the machine comes from the other side of the table, not from the operator. That is mandate in procurement, quantified and in production.

Pactum figures via Bloomberg, April 2023 · closing rate: Harvard Business Review, November 2022 · Source

Permalink

Survey

30 to 45 percent of US consumers use generative AI for product research and comparison. 17 percent said they would begin their holiday shopping on an AI platform.

What the numbers do not say: They rest on self-reporting and cover the US, not the German-speaking market. And the two figures measure different things: research and comparison is not the same as beginning the shopping journey. Merging them produces a number that does not exist. A journey begun is also not a purchase. As an order of magnitude for the shift in discovery they are usable; as a revenue statement they are not.

Bain & Company, Consumer Lab survey, US · November 2025 · the May 2026 insights page states around 30 percent · Source

Permalink

Survey

Only 11 percent of respondents would let an AI make the purchase decision, and only in low-stakes categories such as personal care and household supplies.

What puts the number in perspective: Even willingness to accept mere help is not a majority: 31 percent would let AI narrow choices for household supplies, 28 percent for electronics. It is the counterweight to forecasts about autonomously buying agents. Stated preference and later behaviour do diverge regularly, especially where experience is missing. Sample: 322 US consumers.

Gartner, survey of 322 US consumers in January 2026 · published 27 May 2026 · Source

Permalink

Tribunal decision

Air Canada is liable for what its chatbot promised. The defence that the bot was responsible for its own actions was called "a remarkable submission" by the decision-maker.

What the case is and is not: It concerns a bereavement fare and 650.88 Canadian dollars in damages, not a general refund practice. It was decided by the Civil Resolution Tribunal in British Columbia, not a court, and it binds no one in Europe. The load-bearing sentence still reaches far beyond the case: it makes no difference whether information comes from a static page or a chatbot. The pointed phrase "separate legal entity" is the decision-maker’s summary, not the airline’s wording.

Moffatt v. Air Canada, 2024 BCCRT 149, Civil Resolution Tribunal of British Columbia · 14 February 2024 · full decision · Source

Permalink

Documented incident

Cursor’s support agent invented a usage rule that never existed and replied under the name "Sam".

What the case shows: The damage came not from a wrong answer alone but from the fact that it sounded like a binding company rule. The co-founder publicly clarified that no such rule existed and apologized; users then reported cancelling subscriptions. No figure for that exists. A single incident, but with the same pattern as the others: the agent speaks as the brand.

The Register, 18 April 2025 (invented rule, apology) · Fortune, 19 April 2025 (the name "Sam", cancellations) · Source

Permalink

Legal status

The EU AI Act transparency duties for chatbots and AI-generated content have applied since 2 August 2026. The high-risk duties were postponed to December 2027 and August 2028.

What this means in practice: Anyone interacting with an AI system must be able to tell, unless it is obvious anyway, and synthetic content must carry machine-readable marking. Systems already in use get a four-month grace period for the marking requirement. On the calendar: deadlines have been moved repeatedly, most recently by the Digital Omnibus in May 2026 — though not this one. The direction is stable; the calendar was not.

EU AI Act, Article 50, applicable since 2 August 2026 · postponements via the Digital Omnibus, May 2026 · Source

Permalink

Court ruling, not final

A German higher regional court has ruled that a company is directly liable for the misleading statements of its own AI chatbot. The chatbot is not a third party in the meaning of the law.

What the case is and is not: A clinic’s chatbot attributed medical specialist titles to its directors that they did not hold and that in part do not exist. The court held that it does not matter whether the company had the chatbot programmed with correct data only. Important limits: the judgment is not final; appeal to the Federal Court of Justice was allowed. It is competition law, so it concerns injunctive relief, not damages. And it covers the self-operated chatbot. Who is liable when a third-party system makes false claims about a brand is not decided by it.

Higher Regional Court of Hamm, judgment of 12 May 2026, case 4 UKl 3/25 · Source

Permalink

Self-assessment, forecast

In a Gartner survey, chief executives estimate that by 2030, 15 to 20 percent of their revenue will come from machine customers.

What this is: a self-assessment by surveyed executives about the future, not a measurement and not a Gartner house forecast. Citations routinely compress both into "Gartner expects". Estimates like this for new categories are often wrong, usually in the timing rather than the direction. On sourcing: the page blocks automated retrieval; the wording is evidenced through an archive and Gartner’s own video title, not through a direct fetch.

Gartner, CEO survey, cited in the Think Again series · archive snapshot July 2026 · Source

Permalink

Vendor figures

The large user numbers for AI systems come from the vendors themselves and are not comparable with each other: around 900 million weekly active users for ChatGPT, over one billion monthly users for Google AI Mode.

Why the comparison limps: One number counts weekly, the other monthly. Placed side by side they still read as equivalent, and that is exactly how they travel through presentations. Neither is independently audited. The Google figure is at least documented by the vendor directly; the ChatGPT figure circulates as a company statement in reports about it. Usable as an order of magnitude, not as evidence.

Google, AI Mode blog post, 19 May 2026 · ChatGPT figure: OpenAI statement, February 2026 · Source

Permalink

Official statistics

26 percent of German companies with ten or more employees used AI technologies in 2025. Among large companies with 250 or more employees it is 57 percent, among small ones 23 percent.

What the number does not say: It is collected as a yes-or-no and says nothing about intensity, maturity, or effect, and it is not broken down by function — so it is no evidence about marketing or brand management. Why it belongs here anyway: it is the most methodologically rigorous German figure available and a sober anchor against industry-association surveys reporting markedly higher numbers for the same year. When two figures on the same question diverge widely, the cause lies in population and method, not in reality.

German Federal Statistical Office, ICT usage survey, companies with 10 or more employees · November 2025 · Source

Permalink
Verified study, vendor documentation, or court decision Preliminary: prototype, single test, forecast, or vendor figure Status, case report, or market observation

How this collection is built

An entry is added only after I have checked it against the primary source for a text of my own. No entry rests on a summary of a summary. Where only a report about a study was available to me, the report is named as the source, not the study.

The limit belongs to the number. The most common error in this field is not the wrong number but the right number carrying a claim that reaches further than the evidence. That is why every entry has two parts, and the second one matters more.

This page ages. Every entry carries its date, the page carries its state. Numbers that are superseded get replaced, not quietly deleted.

Free to use with attribution. When in doubt link the primary source rather than this page: it is the route there, not the evidence itself.

Two small things about using this page: sources open in a new tab so your filter and search term stay put here. And "Permalink" puts the address of the single entry on your clipboard, so you do not have to fish it out of the address bar.

Next: what machines see of a page · Glossary · Articles

Contact

LinkedIn · Email · Home