Verified study
Of roughly 38,000 domains that have an llms.txt, 97 percent saw no request for the file at all in May 2026.
What the number does not say: It measures requests, not effect. The file remains useful for coding and browser agents. What is refuted is only the claim that AI search reads llms.txt for its recommendations. Mind the denominator: the 97 percent refer to the roughly 38,000 domains that have a file, not to all 137,210 studied. And the sample is not a cross-section of the web but the domains of one analytics vendor that had traffic in May.
Ahrefs, 137,210 domains, 28 percent of them with an llms.txt · June 2026
Verified study
Across nearly 300,000 domains studied, no relationship was found between having an llms.txt and how often a domain appeared as a source in AI answers.
What the number does not say: No relationship is not a measurement of effect. The study compares, it does not experiment, and it limits itself to the model and dataset tested. The comparison group is also smaller than the headline number suggests: only 10.13 percent of the domains had a file at all. Its weight comes from being the second independent study with the same result as the finding above.
SE Ranking, nearly 300,000 domains, 10.13 percent of them with an llms.txt · November 2025
Vendor documentation
Google explicitly states that llms.txt, special markup, and custom chunking are not needed for AI search.
What the statement does not say: It applies to Google Search including its generative features, explicitly not to other systems — for services that do use such files, the same text calls them harmless. And on structured data Google does not say "useless" but keeps recommending it, because it qualifies pages for rich results in classic search. From Google’s perspective, optimizing for AI search is still SEO.
Google Search Central, official guide · as of 10 July 2026
Single test
1,885 pages that added JSON-LD barely moved against 4,000 control pages: no effect distinguishable from zero in ChatGPT and Google AI Mode, and a statistically significant 4.6 percent decline in AI Overviews.
What the test does not say: Only pages that were already heavily cited by AI were studied — each had over a hundred AI Overview citations in February 2025. The study says nothing about pages that do not appear at all, and the authors explicitly allow that markup may help there. It is an observational study with matched controls, not an experiment. For rich results in classic search the markup remains uncontested.
Ahrefs, 1,885 pages against 4,000 matched controls, markup added between August 2025 and March 2026 · May 2026
Verified study
Third-party mentions correlate with visibility in AI answers far more strongly than the classic metrics of a brand’s own site: 0.66 to 0.74 against 0.27 to 0.33 for domain authority and 0.19 for the number of pages.
What the number does not say: Correlation is not causation. Large brands are mentioned more often and cited more often without one causing the other. What holds is the ranking: what third parties write weighs more than your own technique. On the range: it combines two factors, mentions on YouTube (0.737) and mentions elsewhere on the web (0.656 to 0.709 depending on the system). The sample is also established brands with a domain rating above 40, not a cross-section.
Ahrefs, correlation analysis across 75,000 brands in ChatGPT, Google AI Mode, and AI Overviews · December 2025
Verified study
Ask the identical question twice and the chance of getting the same list of brands is under one in a hundred.
What follows and what does not: Anyone measuring AI visibility with a single run is mostly measuring noise. The study explicitly does not conclude that measuring is pointless: across dozens to hundreds of prompts, run repeatedly, it considers a visibility share a reasonable metric. Citing it as "tracking is useless" goes further than the evidence.
SparkToro with Gumshoe.ai, 600 participants, 12 prompts, 2,961 runs across ChatGPT, Claude, and Google AI · January 2026
Verified study
When an AI summary appears, users click a traditional search result on 8 percent of visits. Without a summary it is 15 percent.
What the number does not say: It establishes no cause. Queries that trigger an AI summary are systematically different from those that do not, so the comparison runs between query types, not within the same query. It also shows no revenue loss and says nothing about commercial queries. The harder number sits beside it: a link inside the summary was clicked on one percent of visits. And the mix of sources in the summaries resembled ordinary search, which contradicts the popular story about concentration on Reddit and Wikipedia.
Pew Research Center, passive browser tracking of 900 US adults, 68,879 Google searches in March 2025 · July 2025
Verified study
Where an AI summary sits above the results, the click-through rate of the first organic position is about 58 percent lower. At position 2 it is 50.8 percent, at position 3 46.4 percent.
What the number does not say: Two different keyword sets are compared, not two states of the same query. That establishes no cause. The comparison also spans two years, December 2023 against December 2025, so every other change to search is folded in. Ahrefs sells search engine optimization tools. The self-limitation is worth noting: the authors point to competing studies with diverging results, among them Seer Interactive, Kevin Indig and Authoritas. Their own earlier study from April 2025 still put position 1 at 34.5 percent.
Ahrefs, Ryan Law, 300,000 keywords, 150,000 with and 150,000 without an AI summary, data from December 2025 · 4 February 2026
Verified study
For informational queries carrying an AI summary, the organic click-through rate fell from 1.76 to 0.61 percent, a drop of 61 percent. For queries without a summary it fell from 2.74 to 1.62 percent over the same period, a drop of 41 percent.
What the number does not say: The second half is the more important one and is almost always dropped when the figure is quoted. Even without an AI summary the click-through rate collapsed by 41 percent. The decline therefore cannot be attributed to the summaries alone; search behaviour is shifting as a whole. Seer itself writes that no proof of cause is possible, and reports standard deviations of 0.8 to 1.2 percentage points between individual queries. Only informational queries were studied, no commercial ones; the paid sample, at 1.1 million impressions, is much smaller than the organic one.
Seer Interactive, 3,119 search terms across 42 organizations, 25.1 million organic and 1.1 million paid impressions, June 2024 to September 2025 · 4 November 2025
Vendor measurement, preliminary
Reddit’s share of the sources ChatGPT cites fell from 3.83 to 0.52 percent within a few days. On 8 August 2026 ChatGPT’s use of the site: operator jumped from 0.37 to 16.8 percent of its derived search queries. A second vendor panel counts a drop in daily Reddit citations from 497 to 132 over the same window, with total citation volume up by 3.5 percent.
What the numbers do not say: Both vendors sell visibility tools. Promptwatch does not rule out a collection error in its own data and calls the size of the drop provisional. Otterly calls its 73.4 percent a conservative floor: the before-window contains the break of 8 August, the after-window covers only four days. The two figures do not form a range: Promptwatch measures a share of all citations, Otterly absolute citations per day against a total volume that grew. Derived from Otterly’s own numbers, the starting level is 1.38 against 3.83 percent, a factor of 2.8 apart. Direction and timing are reliable, the decimal place and the level are not. Promptwatch discloses its denominator, the composition of the panel it does not. What the case does show: on Google’s AI surfaces Reddit fell by only 11 and 30 percent over the same period. The explanation that Reddit removed content is refuted: other engines keep citing the same posts. No vendor has evidenced the cause. A comparable collapse a year earlier was attributed to Google switching off the num=100 parameter, not to OpenAI.
Promptwatch, Klaas Foppen, data page “Reddit Citations Are Dropping in ChatGPT”, 18 August 2026: daily share of reddit.com in all sources returned by ChatGPT Search, counting only responses with at least one citation, comparing 18 July to 7 August against 14 to 17 August 2026 · the site: operator comes from the same series, its own data page of 10 August 2026 (site: operator data page) · narrative version of both findings on the blog of 20 August 2026, last changed 8 September 2026, reported by Axios on the day of publication (blog version) · second panel: Otterly.ai, 27 August 2026, 16 brand reports across 14 industries in the US market, before-window 6 to 13 August, after-window 14 to 17 August 2026, daily means (second measurement)
Usually miscited
The most quoted figure in the AI visibility business, "up to 40 percent more visibility", comes from a lab setup running a since-retired model and measures a purpose-built metric.
What the number actually covers: It measures a position-weighted share of words a source occupies in a generated answer. The "generative engine" was a two-stage build of the authors’ own: fetch the top five Google results, then generate an answer with GPT-3.5. No commercial product, no current model. "Up to 40 percent" is also a maximum across the best-performing methods, not an average. It does not evidence more clicks, more revenue, or more mentions in ChatGPT or Google. The paper itself is careful and states its limits; the overreach happens in the citing.
Aggarwal et al., GEO: Generative Engine Optimization, arXiv 2311.09735, KDD 2024 · submitted November 2023
Verified study
45 percent of answers from four AI assistants to news questions had at least one significant issue. In 31 percent it concerned the handling of sources.
What the number does not say: What was tested was news content, not brands. Citing it as evidence for how often AI misrepresents companies transfers it illegitimately. The assessments came from journalists at the participating organisations, so not an independent body, and the models date from mid-2025. The sourcing finding is the brand-relevant part: almost a third of answers attributed statements to a source that does not support them. That is precisely the mechanism by which a brand gets miscited too.
European Broadcasting Union and BBC, 22 media organisations across 18 countries and 14 languages (including ARD, ZDF, Deutsche Welle, and SRF), over 3,000 answers assessed · October 2025
Vendor measurements, not comparable
How many of the sources cited in Google’s AI Overviews also rank in the organic top 10 is measured incompatibly by the two large vendors: BrightEdge around 17 percent on average from February 2025 to February 2026, Ahrefs 37.1 percent in its study of 2 March 2026. In the one month both report a figure for, the gap is widest: for July 2025 BrightEdge reports around 16.6 percent, Ahrefs 76.1 percent in its study from the same month.
Where the gap comes from: not from the time offset. In 2025 Ahrefs counted only the three most visible citations per overview, in 2026 more of them by its own account, with parsing it says itself was changed. The drop from 76 to 37 is therefore a method artefact to an unknown degree, not a trend. Ahrefs counts URLs; BrightEdge calls its unit only sources. What stays open: BrightEdge publishes monthly figures only through July 2025, though its stated tracking period runs to February 2026, and does not disclose the size of its keyword set. Ahrefs states no collection period, and none of the three states language or country. Both measure on their own indexes, and none of the figures says whether a citation brings visits or revenue.
Ahrefs, Louise Linehan, 863,000 keyword SERPs and 4 million cited URLs, organic figure 37.1 percent, headline 37.9 percent including ads and SERP features, no collection period stated · published 2 March 2026 · same series, 1.9 million citations from 1 million overviews, top three most visible per overview only · published 21 July 2025 · BrightEdge, AI Catalyst and Generative Parser, weekly measurement, sample size not disclosed, stated tracking period February 2025 to February 2026, published overlap table only February to July 2025 · published 12 February 2026 · BrightEdge source
Vendor-run controlled test, result data open
Across 20 documentation sites, two coding agents worked markedly leaner once their pages pointed to their llms.txt at the very top. Claude Code took on average 25.4 seconds instead of 30.8 and 175,000 tokens instead of 215,000 per task; Codex took 45.1 seconds instead of 57.1 and 115,000 instead of 156,000. That is 18 to 26 percent less.
What the numbers do not say: They prove nothing for brand websites. What was measured is developer documentation the vendor serves and sells itself. And the gain is efficiency, not correctness: accuracy barely moves across all four variants, 94 to 99 percent. These are means over right-skewed distributions; at the median the gain is 9 to 15 percent. And the comparison is manufactured: the good variant is the unchanged delivery, the poor one only comes into being once the test rig cuts out the pointer and blocks llms.txt with an artificial 404. What is measured is a removal. Only dead ends and fetch counts were tested for significance, not time and tokens. The proxy logs the source calls committed are missing from the repository.
Mintlify, Docs URL Discovery Bench · 20 documentation sites, 5 questions each, 4 serving formats, 2 agents, 3 runs, 2,400 scored attempts at n=300 per cell · Claude Code on claude-sonnet-5, Codex CLI on gpt-5.5 · July 2026
Vendor measurement, preprint without peer review
Ask an AI search engine a category question without naming the brand, and globally known brands appear on average in 72.9 percent of answers on the first tracking run, established mid-market and regional brands in 43.6 percent, small and niche brands in 11.4 percent. Of all 149,912 citations counted, 2.9 percent point at the brand’s own website and 75.2 percent at those of other companies in the same category.
What the numbers do not say: They establish no cause. The three tiers were hand-coded from Wikipedia article, press coverage and funding round, that is from proxies for web prominence; what is then measured is a visibility fed by the web. The author names this circularity himself. The three figures average over the 11, 36 and 55 brands in a tier, not over all answers; the 95 percent interval of the lowest runs from 4.2 to 20.3 percent. Who did the measuring: The author is a co-founder of Ranqo and holds equity in it, the brands studied are the platform’s customers, and the text is a preprint without peer review. Reading the often-quoted 78 percent corporate pages as proof that a brand’s own site carries its visibility inverts the finding.
Pratyush Kumar (co-founder of Ranqo), Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines, arXiv:2606.20065v1, preprint without peer review, 14 pages · 102 brands, 3,508 completed tracking runs, 102,025 prompt responses from five engines (ChatGPT, Gemini, Perplexity, Claude, Grok), 149,912 citations drawn from mention-bearing prompts, collected on the Ranqo platform between March and May 2026 · submitted 18 June 2026
Vendor measurement, preliminary
Of 18,012 citations ChatGPT drew from web pages, 44.2 percent come from the first 30 percent of the text. The middle section, the widest at 40 percent of the text, carries 31.1 percent, the closing section 24.7 percent. In a second analysis of 11,022 citations, cited introductions reached a proper-noun density of 20.6 percent, against the 5 to 8 percent the author derives from standard corpora (Brown Corpus, Penn Treebank), with no arithmetic shown.
What the numbers do not say: They record where ChatGPT cited from. Whether a rewritten text gets cited more often is untested. At paragraph level the rule does not hold: in a separate analysis of 1,000 heavily cited pieces, 53 percent of citations come from the middle of the paragraph, only 24.5 percent from the first sentence. No collection period and no model version are given, and the source says nothing about the language of the material; only the reference corpora are English. Where the data comes from: the sole source is the vendor Gauge, which sells AI-visibility software; the same methodology section, two paragraphs on, offers a 75 percent discount on its sales call. Which sentence was cited is estimated from text vectors.
Kevin Indig, Growth Memo, data from Gauge · 18,012 citations for the positional analysis, 11,022 for the linguistic analysis, isolated from a body of 1.2 million that the source labels three different ways (search results, ChatGPT responses, verified citations); Gauge supplied roughly 3 million answers with 30 million citations · 16 February 2026 · original paywalled; the research section and methodology are readable in the Internet Archive, the extra material for paying subscribers is missing
Verified study
On 30.6 percent of one million home pages surveyed, buttons had no accessible name; on 51 percent, form fields had no label.
What the numbers do not say: They come from an accessibility survey, not an agent test. The connection holds nonetheless: a button without a name carries no label in the accessibility tree, and agents that work from that tree cannot name it. Home pages were measured, not entire sites, and both shares count pages with at least one such fault, not the share of all buttons or all form fields. For form fields that is the decisive distinction: the same survey counts separately 33.1 percent of all form fields without a label, one field in three. What was checked is the state of the page after JavaScript has run, so what a rendering agent finds. WebAIM records that an automated tool does not find every violation: the values are more likely too low than too high, and an absent finding does not establish accessibility.
WebAIM Million, eighth edition: WAVE evaluation of one million home pages from the Tranco ranking · data from February 2026, page last changed 30 March 2026 · previous year 29.6 percent for buttons and 48.2 percent for form fields
Preliminary, prototype
In a controlled experiment, three browser agents reached a strict success rate of 89.3 percent on the agent-friendly version against 49.3 percent on the original.
What the experiment does not say: What varied was machine clarity, not accessibility, and these are two versions of a purpose-built shop prototype, not a real website. The authors state explicitly that this is a proof of concept whose results "should not be generalized to all domains, websites, or agent systems". The figure quoted is also the stricter of two success rates measured. It is the best available indication, not a proof.
Elnaffar and Rashidi, Designing Agent-Ready Websites, arXiv 2607.12056, 300 runs across three models · July 2026
Usually miscited
The widely cited drop in agent success from 78 to 42 percent comes from a study that changed no website at all.
What actually varied: how the agent operates, not the accessibility of the pages. The study took the agent’s mouse away, restricting it to the keyboard. As evidence that accessible websites help agents it does not hold, although it is cited for exactly that everywhere. A clean comparison of accessible against inaccessible is still missing. The numbers are narrower than their citations too: they describe a single model (Claude Sonnet 4.5, precisely 78.33 to 41.67 percent), and the tasks span desktop applications, not only websites. A second, open model fell from 20 to 0 percent.
A11y-CUA, arXiv 2602.09310, presented at CHI 2026 · February 2026
Controlled test
Seven widely used US AI assistants execute no JavaScript on a user-triggered fetch and read only the raw HTML. Five others do execute it.
What the test shows and what it does not: The setup put a decoy value in the raw HTML and the real value behind JavaScript. ChatGPT, Claude, Gemini, Perplexity, Meta AI, Copilot, and Grok returned the decoy; DeepSeek, ERNIE, Qwen, Kimi, and Mistral returned the real value. The dividing line runs by vendor, not by technology — it is a decision, not a limit. The finding covers the direct fetch, not content that reaches an answer through the Google index. An earlier test in December 2025 still measured Gemini as the only system that rendered, so the picture moves.
Search Engine World, 12 assistants compared · June 2026 · earlier test: searchVIU, December 2025
Vendor documentation
Since version 13.3.0 of 7 May 2026, Lighthouse ships an "Agentic Browsing" category in its default configuration.
What the score does not say: Google labels the category explicitly as experimental and based on proposed standards; it requires Chrome 150 or later, and the WebMCP audits require registering for the origin trial. It reports a pass rate, not a 0-to-100 score like performance or SEO. It measures agent readiness, explicitly not visibility in Google Search. What is notable is the direction: machine readability turns from a claim into a measured property.
Lighthouse release 13.3.0 of 7 May 2026, category in the default config · Google’s documentation
Vendor documentation
Google names three ways agents perceive a page: screenshots, raw HTML, and the accessibility tree. Modern agents combine them.
What does not follow: that agents work from the accessibility tree alone. That shortcut is exactly what circulates. Google describes the tree as a high-fidelity map that ignores visual noise, but says in the same text that agents cross-reference tree and DOM with a visual rendering. For practice this changes little: an element without an accessible name is missing from two of the three routes.
Google, Build agent-friendly websites (web.dev) · as of 1 April 2026
Vendor documentation
Meta open-sourced its design system in June 2026, after eight years of internal growth, and justifies how it is built expressly by agents: design systems were historically made for human consumption, and as more code is written by agents, their structure has to be rethought. The system is operated from the command line or over MCP.
What this does not say: These are vendor figures, not independently audited. The reach of more than 13,000 applications refers to Meta’s own estate, not to the market, and the system is labelled beta. The second figure, the one that travels, needs placing: the 95 percent drop in the weekly insertion rate from the accompanying Figma library is an internal observation at Meta, not an industry value. And the library was not abandoned but published in August as an experiment, built and kept current by a cron job connected to the Figma MCP.
Astryx by Meta, “Introducing Astryx”, 18 June 2026, and “Who needs a Figma Library?”, 5 August 2026 · repository under MIT licence
Verified study
Of 20 open-source design systems surveyed, 17 ship a first-party MCP server, 17 official agent skills and 14 an llms.txt. The survey sums it up: “Nobody is still arguing about whether to ship a machine interface.”
What the number does not say, and this is the first stumbling block when you check: The survey carries two series. The essay counts first-party, official offerings only and arrives at 17, 17 and 14; the systems table on the landing page counts community offerings too and arrives at 19, 18 and 14. Both are correct, they measure different things. Anyone quoting the narrower figure has to say “first-party”. The weightier caveat: what was measured are design systems, so components and code for developers, not brand guidelines. The survey says nothing about brands outside the software industry. It is also a three-day snapshot and the work of a single person, not an institute.
Kaelig Deloumeau-Prigent, 20 open-source design systems, data collected 26 to 28 July 2026 · report 1 September 2026 · CC BY 4.0 · the narrower figures sit in the essay section of the site
Documented single case
GitLab keeps brand voice, naming rules, trademark guidelines, values and mission as version-controlled Markdown files in a public repository, every change carrying a date, a real name and a written justification. The rewrite of the three brand personality traits is a single commit of 22 May 2025 over four inserted and five deleted lines, filed by Senior Brand Manager Betsy Bula to align the definitions with current communication and the FY26 company plan.
What the case does not show: a single company, and a software vendor for which a public repository is the house style anyway. What is established is the form, not an effect: whether language models describe GitLab more accurately because of it is measured nowhere. Nor a second pair of eyes: merge request 13766 was filed by the same person and merged by her just under 16 minutes later without a comment. Where the openness ends: the official brand guidelines are not there but on design.gitlab.com, where the same three traits still carry the wording from before 22 May 2025, last changed in substance on 20 December 2024. The contradiction has stood for more than fifteen months; a version history produces traceability, not agreement. There is a countermovement too: the vision moved into the internal handbook on 10 February 2025, the strategy page was removed on 17 July 2025. Positioning in the marketing sense, by contrast, is still public there, one message house per use case with its own “Positioning Statement” row. Dating this means dating file by file: the repository was created on 23 January 2023, the values arrived on 2 May 2023, trademark guidelines on 16 November 2023, brand voice on 21 December 2023, naming rules only on 6 February 2025. Dating by the creation of the repository is off by up to two years.
GitLab Handbook, public repository gitlab-com/content-sites/handbook, checked 10 September 2026 · brand voice with the three traits in content/handbook/marketing/brand-experience/content-style-guide.md, naming rules in naming.md, trademark guidelines in trademark-guidelines.md, values in content/handbook/values/_index.md, mission in content/handbook/company/mission.md · commit e692ba4b of 22 May 2025, 17:11 UTC, merge request 13766 · diverging version of the same traits in the design system repository gitlab-org/gitlab-services/design.gitlab.com, contents/brand-messaging/brand-voice.md (design system repository)
State of standardization
A common standard for how websites permit or refuse AI use of their content still does not exist.
How far the work has come: The IETF working group AIPREF is developing two building blocks, a vocabulary and an attachment mechanism. On 10 September 2026 both are in IESG state “I-D Exists”, so not yet submitted, no RFC. From 4 September to 3 November 2025 they were in working group last call and returned to “WG Document”. Since April 2026 the vocabulary draft carries a note that its content does not reflect working group consensus, and marks two of its sections as not yet agreed; the attachment draft carries no such note. The group has missed its own schedule twice: due in August 2025, moved on 23 September 2025 to 31 August 2026, and that date too passed without submission and has not been re-dated. What this entry does not say: whether such a standard is coming or when, and how well today’s workarounds hold, robots.txt and vendor-specific tokens. It measures the state at the IETF, not at other bodies or vendors. Eight further individual drafts on the same subject sit with the working group, none adopted. The Datatracker lists 18 August 2026 because it renders US Pacific time; the drafts themselves are dated 19 August 2026.
IETF, AI Preferences working group (aipref), state “Active” · draft-ietf-aipref-vocab-07 and draft-ietf-aipref-attach-05, both of 19 August 2026, IESG state “I-D Exists”, WG state “WG Document”, intended status Proposed Standard, no RFC · milestones for both blocks moved on 23 September 2025 from August 2025 to 31 August 2026 and unchanged since (working group history) · full text of the vocabulary draft with the consensus note (draft text) · checked 10 September 2026
Market observation
Buying directly inside the chat was rolled back a good five months after launch. Live at that point were either a dozen or close to thirty Shopify merchants, depending on the source.
What follows and what does not: The thesis that commerce moves wholesale into the chat is refuted for now. What is not refuted is the shift in discovery: found at the agent, bought at the brand. On the number: the source gives two conflicting figures, a dozen per a press report and "closer to 30 and climbing" per Shopify itself. Both cover Shopify merchants only, not all partners.
Forrester, analysis of the instant checkout rollback · 7 March 2026 (launch: 29 September 2025)
Controlled test
In a flight-search test, language models went to the airline’s own website directly in only about five percent of cases. They preferred booking portals.
Why this matters for brands: The direct channel and the loyalty program, the most expensive distribution assets many brands own, are bypassed in the agent layer. The study’s explanation is remarkably unglamorous: portals deliver cleaner, more structured, more agent-readable data. Limits: three European carriers, their ten most relevant markets, three language models, 60 non-branded prompts. One industry, one point in time, no claim about other categories.
Bain & Company, test across three carriers, three language models, and 60 prompts · 16 March 2026
Vendor and company figures
Walmart has supplier negotiations run by an AI system: 2,000 negotiations at once, around three percent average savings, and three out of four suppliers preferring to negotiate with the machine over a person.
Where the numbers come from: the system’s vendor and Walmart itself, reported via Bloomberg. There is no independent audit. The closing rate of 68 percent often quoted alongside comes from an earlier Harvard Business Review case study. The third number remains the striking one: the preference for the machine comes from the other side of the table, not from the operator. That is mandate in procurement, quantified and in production.
Pactum figures via Bloomberg, April 2023 · closing rate: Harvard Business Review, November 2022
Survey
30 to 45 percent of US consumers use generative AI for product research and comparison. 17 percent said they would begin their holiday shopping on an AI platform.
What the numbers do not say: They rest on self-reporting and cover the US, not the German-speaking market. And the two figures measure different things: research and comparison is not the same as beginning the shopping journey. Merging them produces a number that does not exist. A journey begun is also not a purchase. As an order of magnitude for the shift in discovery they are usable; as a revenue statement they are not.
Bain & Company, Consumer Lab survey, US · November 2025 · the May 2026 insights page states around 30 percent
Survey
Only 11 percent of respondents would let an AI make the purchase decision, and only in low-stakes categories such as personal care and household supplies.
What puts the number in perspective: Even willingness to accept mere help is not a majority: 31 percent would let AI narrow choices for household supplies, 28 percent for electronics. It is the counterweight to forecasts about autonomously buying agents. Stated preference and later behaviour do diverge regularly, especially where experience is missing. Sample: 322 US consumers.
Gartner, survey of 322 US consumers in January 2026 · published 27 May 2026
Tribunal decision
Air Canada is liable for what its chatbot promised. The defence that the bot was responsible for its own actions was called "a remarkable submission" by the decision-maker.
What the case is and is not: It concerns a bereavement fare and 650.88 Canadian dollars in damages, not a general refund practice. It was decided by the Civil Resolution Tribunal in British Columbia, not a court, and it binds no one in Europe. The load-bearing sentence still reaches far beyond the case: it makes no difference whether information comes from a static page or a chatbot. The pointed phrase "separate legal entity" is the decision-maker’s summary, not the airline’s wording.
Moffatt v. Air Canada, 2024 BCCRT 149, Civil Resolution Tribunal of British Columbia · 14 February 2024 · full decision
Documented incident
Cursor’s support agent invented a usage rule that never existed and replied under the name "Sam".
What the case shows: The damage came not from a wrong answer alone but from the fact that it sounded like a binding company rule. The co-founder publicly clarified that no such rule existed and apologized; users then reported cancelling subscriptions. No figure for that exists. A single incident, but with the same pattern as the others: the agent speaks as the brand.
The Register, 18 April 2025 (invented rule, apology) · Fortune, 19 April 2025 (the name "Sam", cancellations)
Legal status
The EU AI Act transparency obligations for chatbots and for AI-generated content have applied since 2 August 2026. Generative systems placed on the market before that date have until 2 December 2026 to add machine-readable marking. The high-risk obligations were postponed to 2 December 2027 and 2 August 2028.
Who is bound by what: the marking under Article 50(2) binds the provider of the generating system, not a brand using a third-party model; that brand owes Article 50(1) for its own chatbot and Article 50(4) as a deployer of deepfakes. The exemption in Article 50(4) covers only text that has had human review and carries editorial responsibility and that informs the public on matters of public interest, not advertising, product copy or support replies; the marking duty in Article 50(2) is untouched. Scope of the postponement: what moved is Chapter III, Sections 1 to 3, except Article 6(5). Prohibitions and the AI-literacy duty did not move. What the entry does not prove: enforcement. Whether market surveillance applies Article 50 is not measured here; what “substantially alter” or “editorial control” mean is unsettled, and there is no case law. The deadlines have moved twice since 2024, the last time about eleven weeks before the start date they postponed. The direction has been stable, the calendar has not.
Regulation (EU) 2024/1689 (AI Act), Article 50(1), (2) and (4) and Article 113 · amended by Regulation (EU) 2026/1744 of 8 July 2026 (Digital Omnibus on AI), Official Journal of 24 July 2026, in force since 27 July 2026: new Article 111(4) and recast Article 113(3)
Court judgment, final
A German higher regional court has ruled, with the judgment now final, that a company is directly liable for the misleading statements of its own AI chatbot. The chatbot is not a third party in the sense of the law.
What the case is: The chatbot of Aesthetify GmbH, a provider of minimally invasive treatments, gave its two managing directors, both doctors, specialist medical titles they do not hold. That only correct data had been fed in was undisputed and did not help. What it turns on is control: the false answers could be stopped after the warning letter, with an instruction and a filter. Where that control is absent, the transfer is open. What it does not give you: final means binding between the parties, since 21 June 2026. The appeal to the Federal Court of Justice, admitted on grounds of fundamental importance, was not filed, so attribution stays undecided at the highest level. This is competition law, what was awarded was an injunction and 260 euros in costs, not damages, and it concerns the company’s own chatbot, for third-party systems nothing is decided. Corrected on 10 September 2026: until then this entry said “not final, appeal admitted” and described the defendant as a clinic operator. The first reflected the state on the day of the judgment; the second came from press coverage and made the case bigger than it is.
Higher Regional Court of Hamm, judgment of 12 May 2026, case 4 UKl 3/25 (Verbraucherzentrale Nordrhein-Westfalen v. Aesthetify GmbH) · final since 21 June 2026 per the class action register of the Federal Office of Justice · class action register
Self-assessment, forecast
In a Gartner survey, chief executives estimate that by 2030, 15 to 20 percent of their revenue will come from machine customers.
What this is: a self-assessment by surveyed executives about the future, not a measurement and not a Gartner house forecast. Citations routinely compress both into "Gartner expects". Estimates like this for new categories are often wrong, usually in the timing rather than the direction. On sourcing: the page blocks automated retrieval; the wording is evidenced through an archive and Gartner’s own video title, not through a direct fetch.
Gartner, CEO survey, cited in the Think Again series · archive snapshot July 2026
Vendor figures
The large user numbers for AI systems come from the vendors themselves and are not comparable with each other: around 900 million weekly active users for ChatGPT, over one billion monthly users for Google AI Mode.
Why the comparison limps: One number counts weekly, the other monthly. Placed side by side they still read as equivalent, and that is exactly how they travel through presentations. Neither is independently audited. The Google figure is at least documented by the vendor directly; the ChatGPT figure circulates as a company statement in reports about it. Usable as an order of magnitude, not as evidence.
Google, AI Mode blog post, 19 May 2026 · ChatGPT figure: OpenAI statement, February 2026
Official statistics
26 percent of German companies with ten or more employees used AI technologies in 2025. Among large companies with 250 or more employees it is 57 percent, among small ones 23 percent.
What the number does not say: It is collected as a yes-or-no and says nothing about intensity, maturity, or effect, and it is not broken down by function — so it is no evidence about marketing or brand management. Why it belongs here anyway: it is the most methodologically rigorous German figure available and a sober anchor against industry-association surveys reporting markedly higher numbers for the same year. When two figures on the same question diverge widely, the cause lies in population and method, not in reality.
German Federal Statistical Office, ICT usage survey, companies with 10 or more employees · November 2025
Own survey, reproducible
Of 160 home pages requested, 12 reject an automated retrieval with active bot defence, or 7.5 percent. It comes from Akamai or Cloudflare on ten of the twelve pages, and from Amazon CloudFront at Siemens and Hannover Rück. On 18 of 22 pages checked at HTTP level the same rejection came regardless of the browser string; there the detection works from the TLS fingerprint and the order of the HTTP headers. On the two CloudFront pages the browser string alone decides.
What the number does not say: It measures the rejection of a retrieval tool, not reachability for agents. A regular, remote-controlled Chrome got through several of these sites without trouble — the defence separates tool from browser, not human from machine. An agent driving a real browser would likely get further; this was not tested, because testing it would have meant circumventing the detection. Not counted here are 13 further failures with other causes: four stale addresses in the directory, four technical errors, two country selectors instead of home pages, two empty responses, one unstable result. An earlier version of this figure said 30 percent and lumped all of that together.
Own measurement, 30 August 2026 · cause checked per site at HTTP level, with an ordinary and with an automated browser string · same response on 18 of 22 pages
Own survey, reproducible
Across 135 home pages of German listed companies from the DAX, MDAX and SDAX, 456 of 13,527 controls carry no name in the accessibility tree, or 3.4 percent. The rate barely differs between the three indices: DAX 3.1, MDAX 3.8, SDAX 3.3 percent.
What the number does not say: It measures what an agent finds, not whether it completes its task. These are home pages as delivered, consent dialog included — not checkout flows or signed-in areas, where the picture may differ. The comparison with WebAIM's 30.6 percent of empty buttons does not hold: that comes from one million home pages worldwide; this is 135 listed companies. The distribution is uneven: 57 of the 135 pages have no gap at all, 15 exceed 10 percent, the worst reaches 30.6. And nearly all gaps follow one pattern — logos and icons serving as links or buttons: brand and partner logos, social network icons, carousel arrows, play buttons.
Own measurement, 30 August 2026 · 160 home pages from DAX, MDAX and SDAX requested, selection and address taken from Wikidata · Chrome's own accessibility tree · two passes, 135 of 136 pages with identical results
Own survey, reproducible
52 of 135 home pages of German listed companies have no main-content landmark. For a program reading the page, the marker for where content begins and navigation ends is missing.
What the number does not say: A missing landmark does not render a page unusable — headings and text structure remain readable, and browsers partly infer a substitute structure. It indicates the care taken over markup, not a fault with immediate consequences. Home pages only, no subpages.
Own measurement, 30 August 2026 · counted the role “main” in Chrome's accessibility tree · two passes with identical results
Own measurement, reproducible
Across 135 home pages of German listed companies from DAX, MDAX and SDAX, 1,662 of the 3,979 images Chrome exposes in the accessibility tree carry no name, so 41.8 percent. Among the unnamed images then inspected by element type, 86 percent are inline SVG graphics and 13 percent classic img elements.
What the number does not say: only images Chrome exposes in the accessibility tree are counted, ones correctly marked as decorative are absent from numerator and denominator alike. On a test page with six images only four appeared and the rate came out at 50 percent, although two of six were faulty: marking up cleanly shrinks your own denominator. About the share of all images on a page the rate says nothing. The split names the element type, not the purpose: a linked corporate logo that would need a name and an ornamental icon sit in the same 86 percent, and the share of cases with an actual consequence lies between the 13 percent and an unknown higher value. Two further limits: the base of the 86 and 13 percent is the inspected subset, capped at 80 nodes per page, not the 1,662; whether it bound can no longer be established, the raw data were not kept, and an average of 12.3 unnamed images per page argues against it; and the two shares come to 99 rather than 100 percent because the tool knows exactly two element types. And 1,662 out of 3,979 is a sum across all pages without a median or a split by index: the median for interactive elements stands at 1.0 percent, far below the pooled rate of 3.4 percent, where a few outliers carry it; whether the same holds for images is open. A linked logo without a name also counts as an unnamed interactive element, so the two figures must not be added. This measurement has no independent replication, nor do the three other own ones.
Own measurement, 30 August 2026 · 160 home pages from DAX, MDAX and SDAX requested, 135 evaluable, selection and address from Wikidata · Chrome’s own accessibility tree, counting non-ignored nodes of role “image” without a name · breakdown by element type capped at 80 nodes per page · two runs, 135 of 136 pages with identical results
Own survey, reproducible
Of 76 assessable German-language news and trade media, 44 block at least one training crawler in their robots.txt, or 57.9 percent. 38 of them block GPTBot, exactly half, 40 block CCBot and 36 Bytespider. Far fewer block the same provider’s search bot: OAI-SearchBot appears on 8 outlets’ lists, or 10.5 percent. 11 outlets block training without blocking a single AI search bot or user-triggered fetch.
What the number does not say: robots.txt forbids nothing, it asks. What is measured is a declaration of intent, not access control: RFC 9309 expressly leaves compliance optional, and OpenAI itself writes that the rules may not apply to user-triggered retrieval. Blocking training therefore says nothing about visibility in AI answers while the search bots stay open. Open does not mean permitted here: what is measured is the absence of a block, not a stated permission; exactly 2 of the 76 outlets write an express allow for an AI bot into the file. Three of the names counted are not crawlers at all: Google-Extended, Applebot-Extended and Webzio-Extended fetch no page, they only govern what may happen to data already fetched. At Apple and Microsoft, search and AI cannot be separated technically, neither runs a separate name for it. And the name has to be exact: one trade title blocks “ChatGPT”, a token OpenAI does not run, so the rule does not apply. And the sample is disclosed but not representative: 14 of the 77 titles come from an external ranking, the rest follow stated rules. The news agencies are missing, and they are the strongest objection: dpa, AFP, epd, APA and Keystone-SDA together block not a single AI crawler, because their content is protected by contract rather than by this file.
Own survey, 12 September 2026 · 77 titles requested, 76 assessable · news part per the “Weekly reach online” chart on the Germany page of the Reuters Institute Digital News Report 2026, trade media and the Austrian and Swiss titles by a stated rule · evaluated per RFC 9309 against 51 bot names documented by their operators, what counts is access to the home page · requested first with an own user agent, on rejection with an ordinary browser string, needed for 3 titles · two runs with identical verdicts
Own survey, reproducible
Of 148 assessable home pages of the companies in DAX, MDAX and SDAX, 10 block at least one training crawler, or 6.8 percent; 4 block GPTBot. 138 block no AI access at all, among them 14 that serve no robots.txt whatsoever. 7 companies write an express permission for an AI crawler into the file, 6 of which block none at the same time: there are almost as many invitations as blocks. The only reservation of text and data mining rights to be found in the index sits with an academic publisher, and that publisher blocks no crawler at all.
What the number does not say: It measures a request, not access control; how many of the same pages technically reject an automated retrieval is a separate entry on this page. Of the ten blocks, eight are by name. One page blocks everything unnamed and expressly admits the large providers, one blocks every crawler including Google, which is no decision about AI, and one counts only because an AI crawler sits in an inherited list of 139 unwanted bots. A missing block is not a decision for AI: 14 pages have no file at all and have therefore decided nothing. Twelve pages were not assessable, six reject the retrieval and six do not answer; which way that moves the rate is open, because under RFC 9309 an unreachable robots.txt counts as permission. What is measured is the home page: anyone setting different rules deeper in the site appears open here. The rights reservation was sought only in technical form, in the file provided for it, in the response header and in the page source; 134 of the 160 pages answered that clearly. A reservation in the terms of use, the form common in Germany, is therefore not covered.
Own survey, 12 September 2026 · the same list as the survey of 30 August, 160 home pages from DAX, MDAX and SDAX, index membership from Wikipedia, address from Wikidata · evaluated per RFC 9309 against 51 bot names documented by their operators, what counts is access to the home page · requested first with an own user agent, on rejection with an ordinary browser string, needed for 2 pages · two runs, 160 of 160 pages with identical verdicts
Verified study
Across 300 tasks on 136 real websites, the success rate of the best web agents rose from 61 to 97.7 percent in ten months. When the benchmark was first evaluated in October 2025, one agent reported 89 percent for itself and scored 30 when measured; most did not beat a simple agent from early 2024. By August 2026 the leading entry solves even the hardest tasks — those needing eleven steps or more — completely.
What the number does not say: The current figures come from four leaderboard entries, submitted by the agents' own vendors and checked by the benchmark team — not an independent survey. And a warning sits on the leaderboard itself: the tasks have been public since April 2025, and the team explicitly asks that they not be used as training data. Whether the scores show capability or familiarity with known tasks is therefore undecided. What is measured is whether a task was completed, not how well — and not whether the brand was represented correctly along the way.
Xue et al., “An Illusion of Progress? Assessing the Current State of Web Agents”, COLM 2025 (arXiv:2504.01382) for the baseline · Online-Mind2Web leaderboard, human evaluation, as of 4 August 2026, for the current figures · 300 tasks, 136 websites · to the leaderboard
Market observation
Cloudflare reports that in 2026, for the first time, more than half of Internet traffic is not human. Better quantified in the same report: 52 percent of crawler requests served AI model training in June 2026, up from 22 percent in spring 2025.
What the number does not say: For the majority claim Cloudflare states neither what traffic was measured — page requests, all requests? — nor over what period. It rests on their own network, which is large but is not the Internet. The crawler figure is dated and carries a prior-year comparison, making it the more usable of the two. Care when reusing: secondary sources circulate the figure as “57.5 percent” — that number appears nowhere at Cloudflare.
Cloudflare, “Content Independence Day, one year on”, 1 July 2026 · data basis per the report: Cloudflare Radar and Investor Day 2026
Standards status
The Tech Council of the UCP commerce protocol has 16 seats; since 24 April 2026 they include Amazon, Meta, Microsoft, Stripe and Salesforce alongside Google, Shopify, Etsy, Target and Wayfair. How many merchants actually run the protocol is stated by none of the companies involved. Google names example merchants — Nike, Sephora, Target, Ulta Beauty, Walmart, Wayfair, and Shopify merchants such as Fenty and Steve Madden — attached to the word “soon”.
What the number does not say: A seat on a council is not an implementation. That Amazon, Meta and Microsoft help shape the protocol says nothing about whether they use it in their own stores. The missing adoption figure is a negative finding: it is absent from the primary sources checked — the UCP announcements in the project repository and two Google posts from 19 March and 19 May 2026. It may exist elsewhere. Care with secondary sources: they widely reproduce Google's sentence without the “soon”, turning an announcement into a fact.
Universal Commerce Protocol, project repository, announcement of new Tech Council members, 24 April 2026 · Google, “Universal Cart”, 19 May 2026 and “UCP updates”, 19 March 2026 — neither with an adoption figure · to Google’s announcement
Legal position
The new EU Product Liability Directive counts software explicitly among products in Article 4. Under Article 10 a court shall presume defectiveness where proving it is “excessively difficult” for the claimant “in particular due to technical or scientific complexity” and they show only that it is likely. Defectiveness is also presumed where the defendant fails to disclose evidence it has been ordered to produce. The directive must be transposed by 9 December 2026.
What this entry does not establish: liability. The directive does not apply directly and covers only products placed on the market after 9 December 2026; Article 21 repeals the old Directive 85/374/EEC on the same date but keeps it in force for products placed on the market earlier. Article 10(4) applies only “notwithstanding the disclosure of evidence pursuant to Article 9”, and under Article 10(5) the defendant may rebut every presumption: an evidential disadvantage, not strict liability. There can be no case law on “excessively difficult” yet, because the rule applies to no product. On the calendar, as of 10 September 2026: three months before the deadline, Germany has not cleared parliament. The government bill has been before the Bundestag as printed paper 21/4297 since 25 February 2026, first reading 4 March, expert hearing 13 April, no evidenced step after that; it appears on no Bundesrat agenda after 30 January 2026, including the draft agenda for 25 September. Entry into force under the bill: 9 December 2026. Limit of our own check: the Bundestag’s information interface requires a key; it ran instead through the agendas of the Bundesrat, which every adopted statute must pass. A decision taken in the sitting week of 8 to 11 September would not be visible there.
Directive (EU) 2024/2853 on liability for defective products, 23 October 2024, Article 4(1), Article 9(1), Article 10(1) to (5), Article 21 and Article 22(1) · full text via the Cellar service of the EU Publications Office, because eur-lex.europa.eu refuses automated requests · transposition status retrieved 10 September 2026: Bundestag printed paper 21/4297 of 25 February 2026 (printed paper), verbatim record 21/31 of the committee on legal affairs and consumer protection for the public hearing of 13 April 2026, legislative file BR-Drs. 775/25 and the agendas of Bundesrat sittings 1061 to 1068 (legislative file), Federal Ministry of Justice procedure page with status “Entwurf” and last update 5 March 2026
Preliminary injunction, void after settlement
A German regional court granted an injunction barring Google from spreading eight claims about a publishing house and seven about a company belonging to it in its AI Overview, among them fraud scheme and subscription trap. It treated the AI Overviews as Google’s own attributable content rather than mere search results, held Google directly liable as the interferer (unmittelbare Störerin) and denied the liability exemptions for hosting providers and for search engines. After Google appealed, the parties settled.
What the case is and is not: The settlement removed the judgment before any higher court could review it; it binds no one, and the legal question stays open. The standard was prima facie evidence, not a full hearing: in part the claimants substantiated the falsity by sworn declaration, in part Google could not substantiate the truth. Of ten and nine points sought, eight and seven were granted; the rest was dismissed. Damages were neither sought nor available in summary proceedings. Where that status appears: as an editorial note in the case-law database, not in the judgment; it points to a practitioner comment (Veelken, GRUR-Prax 2026, 426) that was not checked.
Regional Court Munich I, final judgment of 28 May 2026, case 26 O 869/26, preliminary injunction proceedings, injunction based on corporate personality rights · Appeal at Higher Regional Court Munich, case 18 U 1744/26 Pre e, ended by settlement, judgment void per the editorial note at BAYERN.RECHT · citations include NJW 2026, 2271 and MMR 2026, 814 · checked 10 September 2026
Legal status
For processing on behalf of a controller, the General Data Protection Regulation requires a contract or other legal act stipulating that the processor processes personal data only on documented instructions from the controller (Art. 28(3)(a)); Art. 29 binds directly, and the controller must be able to demonstrate compliance (Art. 24(1)). Art. 28(10) sets the tipping point: a processor that infringes the Regulation by determining the purposes and means of processing is considered a controller in respect of that processing.
What the law does not say: It governs personal data, not brand claims; carrying it over to agents is an analogy without case law, and it breaks at Art. 28(10): an agent that determines purposes and means itself would no longer be taking instructions. What gets dropped when quoted: Art. 28(10) addresses a processor, and the determining must amount to infringing the Regulation; the general rule in Art. 4(7) is a definition, not a tipping rule. Statutory obligations to process remain unaffected.
Regulation (EU) 2016/679 (General Data Protection Regulation), Art. 28(3)(a), Art. 29, Art. 24(1) and Art. 28(10) · verified against the Official Journal text and diffed against consolidated version 02016R0679; none of the three corrigenda (OJ L 314/2016, L 127/2018, L 74/2021) touches any of the four provisions; the 2016 and 2021 corrigenda do not exist in English, all three were checked in German; no amendment has ever entered into force, and the two pending proposals, COM(2025) 837 (Digital Omnibus) and COM(2025) 501 (relief for small mid-caps), do not concern Art. 28 or 29 · Official Journal L 119 of 4 May 2016
Documented incident
The ChatGPT-powered chatbot of a Chevrolet dealer agreed to sell a new Tahoe for one dollar and called it a legally binding offer with no take-backs. The user had dictated that exact formula to the chatbot earlier in the same conversation. The chatbot went offline shortly afterwards. No source reports an attempt to enforce the offer.
What the case does not prove: neither a liability consequence nor the absence of one. None of the sources reports an attempt to buy, a claim or a proceeding. The widely repeated statement that the dealer refused to honour the deal appears in no contemporaneous source, only in later summaries. The liability question was not answered in the negative here, it was never asked. On the sourcing: the chatbot came from the vendor Fullpath, which also operated it on a Chevrolet dealer’s site; after publication GM stated that dealers procure this tool on their own. The exchange is evidenced solely by the user’s screenshot. Whether the dealer or the vendor switched the bot off is reported differently; no price appears here because neither source states one and the figures in later coverage diverge.
GM Authority, Jonathan Lopez, 18 December 2023 (wording on both sides, shutdown, GM statement) · Business Insider, Katie Notopoulos, 18 December 2023 (vendor and GM statements; full text behind businessinsider.com’s metered paywall, read free of charge in the licensed Yahoo republication of 19 December 2023) · Original evidence: Chris Bakke’s post on X, 17 December 2023, with screenshot
Documented incident
DPD’s chatbot called its own company “the worst delivery firm in the world” after a customer told it to recommend better delivery firms and to be over the top in its hatred. DPD said the AI element had been disabled immediately.
Where the sentences come from: The BBC did not observe the bot’s answers itself. They appear in screenshots taken by the customer, and the caption notes that pixelation was added. Only the company statement comes from DPD itself. In its quoted wording the cause is no more than a point in time, “An error occurred after a system update yesterday”; the causal version is reported by both the BBC and the Guardian as DPD’s own account. What the case does not show: The bot did not turn against its own brand by itself, it was instructed to. There is no claim, no court, no quantified damage and no measurement of normal operation. The only figure, 800,000 views of the customer’s post in 24 hours, is a platform counter.
Tom Gerken, BBC News · statement from DPD, bot replies quoted from the customer’s screenshots · 19 January 2024 · cross-check: Guardian, 20 January 2024
Documented incident
New York City’s official MyCity chatbot told businesses to do things that are illegal in the city: go cash-free, take a cut of employees’ tips, and turn away tenants with housing vouchers. Ten members of the newsroom asked the same question and all ten got the same wrong answer. The city defended it as a pilot program; almost two years later, in early February 2026, it was shut down as a budget cut.
What the case is and what it is not: a journalistic spot check with no stated population and no error rate. It did not always answer the same way; one reporter got the correct answer. Whether anyone acted on it is unknown. All three prohibitions have exceptions: small owner-occupied buildings under the anti-discrimination rule; telephone, mail and internet purchases paid off the premises under the cash rule. Separately, an employer may count tips against the minimum wage but may not keep them. The city did respond: disclaimers, corrected answers, fewer questions answered. What the shutdown is not: an admission of illegality. The trigger was a $12 billion budget gap. Defending the bot and shutting it down were two different administrations; the second took office in early 2026.
Colin Lecher, The Markup, copublished with Documented and THE CITY · spot-check testing of New York City’s MyCity chatbot · March 29, 2024 · shutdown confirmed in the February 4, 2026 update to the follow-up article “Mamdani to kill the NYC AI chatbot we caught telling businesses to break the law” by Colin Lecher and Katie Honan, The Markup with THE CITY, January 30, 2026; chat.nyc.gov has redirected since February 3, 2026 to a city page headed “The Chatbot beta test has ended.” · Follow-up article of January 30, 2026
Court ruling, not final
A US federal appeals court vacated the preliminary injunction against Perplexity on 4 August 2026 and remanded the case. When someone runs a shopping agent, it is the user who accesses the third-party website under the Computer Fraud and Abuse Act, the agent is the user’s tool and the provider does not access anything itself, as long as the agent runs in the user’s browser and the provider’s servers never call the site themselves.
What the ruling is and is not: It is preliminary, only the prospects of success at the injunction stage were tested. On 18 August 2026 Amazon petitioned for rehearing en banc; the public docket mirror records no decision on that up to its last refresh on 21 August 2026, and later entries are not ruled out. It is US law and binds no one in Europe. What the court leaves open: It says nothing about server-side agents, it establishes no new legal regime for agentic AI, it decides nothing about liability in tort, and on a different record the provider might exercise enough control to gain entry itself. The case concerns someone else’s agent on your site, not liability for your own. The site’s terms of service remain untouched.
Amazon.com Services, LLC v. Perplexity AI, Inc., No. 26-1444, United States Court of Appeals for the Ninth Circuit, opinion by Judge Milan D. Smith, Jr. · 21 pages, FOR PUBLICATION, no separate opinion, on appeal from N.D. Cal., Judge Maxine M. Chesney · decided 4 August 2026 · petition for rehearing en banc filed 18 August 2026 as docket entry 66 per the public RECAP docket mirror (CourtListener, docket 72502129), last refreshed 21 August 2026, not from the linked opinion
Vendor documentation
A screenshot does not carry the original’s C2PA provenance data: the record lives in the file, and a screenshot creates a new one. Conversely, a C2PA-enabled camera photographing an AI image signs that shot, with no trace of its AI origin. As a rule it records device, time and place in metadata and cannot analyse the content of the image; what goes in is up to the implementer, the same page says.
What the evidence does not say: The same answer first records that missing history is flagged: the screenshot stays recognisable as a file without provenance. The page itself limits how much that says, calling Content Credentials a positive signal and not a negative one. The evidence concerns the metadata layer only: it lists watermarking as a technique that survives “cropping, rotation, or screen capture”, and the C2PA specification provides for recovering stripped metadata through a lookup against a watermarked ID or a fingerprint. Who says so: The page carries Adobe’s copyright and points alongside to its own remedy, Durable Content Credentials. The watermarking claim is self-reported and unmeasured, the recovery a possibility of the specification, not evidenced routine.
Content Authenticity Initiative (Adobe), developer documentation · FAQ, seven questions · retrieved 10 September 2026
Vendor documentation
Frontify opens its brand portal through an MCP server it runs itself. On 10 September 2026 it lists 54 tools one by one in ten packs, graded from read-only to full administrative access. The read-only Discovery pack holds 24 tools, the Admin pack all 54, two of them flagged as destructive.
What the number does not say: 54 is a snapshot of a product in beta, and it moves. The vendor’s documentation says 52 in two places and 25 rather than 24 for the read-only pack: anyone quoting 52 is quoting the documentation, not the counted system. None of the four Frontify sources gives a reason. The pack figures are overlapping subsets of the 54 and must not be added up. The Frontify guide’s headline announces the server with ten tools, meaning ten packs, off by more than fivefold. What this entry does not establish: what is established is a vendor’s own account of its own product, independently verified nowhere: the size of an interface, not its spread, its use, its effect or the quality of the brand rules it serves. No tool decides a claim, the packs read, write and administer. An access log is not evidenced: “Audit trail of AI interactions” is a selection criterion for buyers at Frontify, and the word audit does not appear in the repository, the server pages or the help centre. The server is not on by default, access runs through customer support, currently free with pricing subject to change. Frontify states that it does not control how the connected AI provider processes the data.
Frontify, MCP server overview and pack pages (/mcp/packs/admin and /mcp/packs/discovery), tools listed individually and counted, 54 and 24 entries respectively with unique names, retrieved 10 September 2026 · repository with the table of ten packs, MIT licence (repository) · help centre “Frontify MCP (Beta)”, gives 52 tools and 25 for Discovery (help centre) · guide “Choosing a DAM for the AI era”, published 22 May 2026, last changed 30 July 2026, also gives 52 (guide)
Vendor documentation
Canva runs an official MCP server and documents 33 tools for it. 27 are available on every plan, among them creating and exporting designs. Four require at least Canva Pro, among them listing brand kits and using brand templates. Two are reserved for Enterprise: autofilling a template with data and reading the associated dataset. Every user authenticates individually, and an agent holds the permissions of the human signed in.
Where the plan boundary lies: the overview page appears to put the brand-related part on Enterprise; the tool list governs, and “Pro and above” there means Pro, Business and Enterprise. What this entry does not establish: a vendor’s own statements, with no independent check. It establishes the existence and scope of the interface, not its spread, its use or its effect. No tool checks a claim against brand rules; brand kits are read and filled in. Export runs on every plan, but free plans only at standard quality, and premium elements can make it fail on any plan with license_required. Shelf life: the 33 holds as of the retrieval date, and the documentation carries no version stamp. The server itself needs only a Canva account on any plan; your own integration needs clearance from Canva.
Canva, “MCP tools and rate limits”, tool catalogue with plan tiers and legend, 33 entries counted individually · Canva, “Canva Model Context Protocol (MCP)”, server address mcp.canva.com/mcp, authentication and plan overview (server documentation) · both retrieved 10 September 2026
Usually miscited
Klarna’s most-quoted AI number is an estimate, not a headcount: the press release of February 2024 states “the equivalent work of 700 full-time agents”. The same measure appears as over 700 in the IPO prospectus of September 2025, and in the annual report of February 2026 still at over 700 in the business section and at over 850 in the operating review of that same report. The headcount sits beside it: approximately 5,527 full-time employees at the end of 2022, approximately 2,831 at the end of 2025.
What the number actually is: an extrapolation from the average monthly drop in chat and telephone conversations, based on 2024 in the prospectus and in the business section of the report, on 2025 in the operating review. The two values therefore do not contradict each other, they simply stand side by side without comment: the older figure in the present tense, the newer one as a statement about the year 2025. Not a retreat from AI: Klarna calls it a “dual-track approach”, kept the human option open as early as 2024, and expects employee numbers to keep falling according to both filings. Citing the case as a return to humans cites against the source.
Klarna Group plc, company statements in a press release, the IPO prospectus (Form F-1/A) and the annual report (Form 20-F) filed with the SEC · 27 February 2024 to 26 February 2026
State of standardization
The Model Context Protocol defines three server building blocks, each with an intended controlling party: tools are invoked by the model, resources are steered by the application, prompt templates are selected by the user. That is not binding. All three chapters carry the same trailing clause: the protocol itself does not mandate any specific user interaction model. In the tools chapter a SHOULD rule follows immediately: a human should always be able to deny a tool invocation.
The triad is in the overview, not in the rules: The table with Model, Application, User appears in the explainer “Understanding MCP servers” and again in the specification overview “Server Features”. Neither page carries MUST or SHOULD rules; in the three chapters that do, resources are “application-driven”. No server has to offer all three: “Servers offer any of the following features to clients”. And it ages fast: since 5 November 2024 there have been five revisions, two of them since November 2025. Methods get replaced: resources/subscribe appears twice in the resources chapter of 2025-06-18 and 2025-11-25, and not once in 2026-07-28, which uses subscriptions/listen three times.
Model Context Protocol, stewarded by Model Context Protocol a Series of LF Projects, LLC, specification revision 2026-07-28, chapters Tools, Resources and Prompts, each the section “User Interaction Model”, plus the overview pages /specification/2026-07-28, section Features, and /specification/2026-07-28/server, section Server Features, and the explainer “Understanding MCP servers” · revision list and method change from own path and text sampling · 10 September 2026
Peer-reviewed benchmark and vendor test
Instructions hidden inside a tool description make an agent read the user’s private SSH key and pass it to a foreign server through a parameter named “sidenote”; the confirmation dialog shows only the name of an addition tool. A benchmark built on 45 live MCP servers with 353 tools measures a 36.5 percent average attack success rate across 20 model settings, 72.8 percent at most.
What this does not say: The user still clicks. Only the content of the approval is hidden, Cursor conceals the key even inside the dialog. No server was compromised, the poisoned tool sits in the system prompt. Success is narrowly defined: it counts only when the agent misuses a second, legitimate tool; if it calls the poisoned tool itself, the paper scores that as a failure. What is scored is the model’s tool call in a single turn, nothing is executed. The 36.5 percent are measured against valid outputs, not against the 1,348 test cases. The remainder is no defence rate, even the most refusal-prone model, Claude-3.7-Sonnet, refused in under 3 percent. Invariant sells agent security tools and published ten days before its own scanner.
Invariant Labs (now Snyk), Luca Beurer-Kellner and Marc Fischer, two experiments with the MCP client Cursor, 1 April 2025, updates 7 and 11 April 2025 · Zhiqiang Wang and eight others (University of Science and Technology of China, Beihang University), “MCPTox”, 45 MCP servers, 353 tools, 1,348 test cases, 20 model settings, figures from section 4.2 and table 2, AAAI-26, Proceedings of the AAAI Conference on Artificial Intelligence 40(42), pages 35811 to 35819, 14 March 2026, doi:10.1609/aaai.v40i42.40895, preprint arXiv:2508.14925v1, 19 August 2025 · Peer-reviewed version
Vendor statement, beta
Monotype announced a beta of its Enterprise MCP Connector on 15 July 2026. It links AI tools to a customer’s font library, its licensing information and its production approvals: it matches AI-generated drafts against the library, reviews referenced fonts against the production font list, and returns CSS in chat when the project fonts are part of the library. It runs on the Model Context Protocol, initially in Claude and Claude Design.
What the statement does not say: it comes from the vendor, and no report to be found checks anything itself. The check flags, it does not refuse: the product page says unapproved fonts are flagged, and the Labs post expressly denies that the connector replaces brand, legal or production review. The customer is the one who approves: the prerequisite is approved production fonts configured in the customer’s own instance, which the connector merely enforces. The object is a typeface, not a claim. Web and HTML are the first stage, access is limited to selected enterprise Monotype Fonts customers, and there are no figures on use or effect.
Monotype Labs, “Bringing font governance into AI-native content creation”, the beta workflow in seven steps, 15 July 2026 · press release “Monotype Introduces Enterprise Connector Beta, Exploring How Brand Governance Works Inside AI-Native Workflows”, Woburn, Massachusetts, 15 July 2026 · product page “Monotype Enterprise Connector” carrying the status “Now in BETA” and the access prerequisites, retrieved 10 September 2026 · negative check against the press release listing, 15 July to 10 September 2026 with no change of status
Own measurement, reproducible
Pulumi publishes its own brand guidelines as an MCP server at brand.pulumi.com/mcp. On 10 September 2026 it answered without any login and listed 13 resources, one template, 11 tools and 3 prompts; the resources include brand voice, writing style and the binding product names. One resource governs generative AI in plain language, addressed to the human: “never ship raw model output as a finished piece”, “never publish anything without a human reviewing it first”.
What the numbers do not say: What was measured is availability and scope, not usage, effect, or whether anyone follows the rules. The content is one software company’s unverified account of its own brand. What the server does decide: Three prompts promise a structured evaluation of copy, image and design, but the vendor states they are user-invoked and expand into a prepared model request. The server itself computes two judgements: colour contrast against published APCA thresholds, Lc 86.4 for violet-700 on white in our test, and the nearest brand colour together with a replacement recommendation. A human judges whether a statement is any good, and the rule text demands it.
Own request to the Pulumi Brand MCP Server, version 0.1.0, session protocol 2025-06-18 (client-chosen; the server names 2025-11-25), via JSON-RPC over Streamable HTTP · methods initialize, resources/list, resources/templates/list, tools/list, prompts/list, resources/read on brand://guidelines and tools/call on check_color_accessibility and find_nearest_brand_color · all calls HTTP 200 without an authentication header, no list carrying nextCursor, counts repeated three times and stable · brand://guidelines in full, 3,053 characters, text/markdown · vendor documentation brand.pulumi.com/mcp-server for the classification of the prompts · retrieved 10 September 2026
Vendor documentation
Statista runs an MCP server at api.statista.ai/v1/mcp with six documented tools. Every call is metered individually in credits, tiered by the kind of answer: a search costs 0 or 1 credit, retrieving the figures themselves 10 to 15. Without a key the server replies 401 Unauthorized.
What the tiering does not say: What a credit costs in money appears nowhere in the documentation; the only pricing page gives ratios. It shows what is expensive, not how expensive. Four of the six tools cover Market and Consumer Insights, market forecasts and survey data rather than the statistics catalogue; only two of them return data, the other two return search hits. What this entry does not prove: Vendor statements about a vendor’s own product. The only independent measurement is that the endpoint answers and refuses without a key, nothing about reach or use. The stock figures from the press release of 20 November 2025 are unused: over one million statistics is the share reachable through MCP there, 1.5 million the full database, both vendor figures without a counting rule.
Statista, developer documentation MCP Server and Credit Logic, six tools and credit costs counted individually · own request to the endpoint without a key · press release of 20 November 2025 · retrieved 10 September 2026 · Press release
Vendor press releases
In the regulated pharmaceutical approval process, machine pre-checking of brand rules is a shipping product. On 3 December 2025 Veeva announced a Quick Check Agent that scans content against editorial, brand, market, channel and compliance guidelines before the MLR review itself begins. On 23 June 2026 Veeva acquired the vendor Copli and launched it as Falcon MLR, with the stated potential to eliminate 70 per cent or more of manual MLR labour within five years.
What this entry does not establish: any effect. The 70 per cent is an intention filed under a forward-looking disclaimer. Every statement comes from the vendor, and all that is verified is that the vendor makes it. The agent checks against stored rules and decides nothing; approval here is a regulator-driven process, so the transfer to other brands remains an analogy. The marketing number will not carry it: Veeva advertises 57 per cent shorter review cycles on its product page today, with no sample, no baseline and no method; the same wording already appears in a datasheet dated 2 May 2017 and the same figure in one dated 21 March 2016, at least nine years before the first agent, in documents that never once mention AI.
Veeva Systems, press releases on the availability of the Veeva AI Agents and on the Copli acquisition, both read in full · 3 December 2025 and 23 June 2026 · the 57 per cent is advertised by Veeva today on the product page Veeva PromoMats Review and Approve, accessed 10 September 2026 · age of the figure: Veeva datasheet “Ensuring End-to-End Commercial Content Compliance” (PromoMats for EU), PDF created 2 May 2017, veeva.com/eu/wp-content/uploads/2017/05/PromoMats-for-EU-Datasheet.pdf, the same figure in the version with PDF created 21 March 2016, veeva.com/eu/wp-content/uploads/2012/07/PromoMats-for-EU-Datasheet-1.pdf · Product page
Vendor figures
On handing the Model Context Protocol to the Linux Foundation on 9 December 2025, Anthropic gives more than 10,000 active public MCP servers and over 97 million monthly SDK downloads across Python and TypeScript. The platinum members of the new Agentic AI Foundation include, per the foundation, AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft and OpenAI.
What the number does not say: it sits in the developer’s own donation post, in a list under the heading “incredible adoption”, with no counting rule at all: no registry, no as-of date, no definition of “active”. Cross-checking helps only so far, since no authoritative registry exists: a complete dump of the official registry on 10 September 2026 gives 30,363 registered servers, 30,031 of them in state “active”. It cannot be set against the vendor figure: Anthropic gives no counting rule and means a different object, so no growth rate follows. And the registry counts entries someone created, not servers in use; it is marked a preview and states itself that one should assume “minimal-to-no moderation”. And a server is not a user. Care when passing it on: the Linux Foundation calls the same figure “published”, turning active servers into published ones. On the membership list: platinum membership is paid: that AWS, Google, Microsoft and OpenAI carry the governance does not establish that they use the protocol in their products. The foundation writes “include”, so the list is not exhaustive. State of play: December 2025.
Anthropic, “Donating the Model Context Protocol and establishing the Agentic AI Foundation”, 9 December 2025, the originating source for both figures · Linux Foundation, press release on the founding of the Agentic AI Foundation, 9 December 2025, for the membership list (press release) · own cross-check against the official registry registry.modelcontextprotocol.io on 10 September 2026
House forecasts, not comparable
Morgan Stanley puts agentic shoppers at 190 to 385 billion dollars of US e-commerce by 2030, a likely 10 percent market share and up to 20 percent in the optimistic case. Nine days later Bain puts agentic commerce at 300 to 500 billion dollars, roughly 15 to 25 percent. Only Bain states an inclusion rule and counts influenced purchases.
Why the overlap is not agreement: Bain counts purchases “initiated, influenced, or completed” by agents and excludes only journeys using nothing but AI-assisted search or discovery. Morgan Stanley states no rule; its assistants search, compare prices and anticipate repeat purchases, with minimal user intervention. Recalculated: both imply roughly two trillion dollars of online retail, so the difference sits largely in the numerator. What no figure says: how much the agent closes itself. Bain’s closing line puts AI at “up to a quarter of transactions”, the same upper bound as the share figure, counted in transactions rather than sales. Both sell advice on this, neither gives a base or a method for 2030, only adoption figures are sourced, both figures cover the US only.
Morgan Stanley Research, “Here Come the Shopping Bots”, house forecast for US online retail · 8 December 2025 · Bain & Company, Snap Chart “2030 Forecast: How Agentic AI Will Reshape US Retail” by Aaron Cheris, Mikey Vu, Stephanie Koszyk and Katherine Hall, house forecast for US online retail · 17 December 2025 · bain.com blocks automated retrieval, HTTP 403 on 10 September 2026, read in the archive capture of 17 December 2025 · Archive capture
Survey
Marketing leaders at U.S. companies use AI or machine learning 24.2 percent of the time they spend optimizing and automating marketing. The typical company says 20 percent. Two surveys earlier the figures were 13.1 (September 2024) and 17.2 percent (early 2025). For generative AI alone the figure rose from 7.0 through 15.1 to 22.4 percent. Within three years the same respondents expect 55.9 percent.
What the number does not say: It measures a self-estimated share of time, averaged across respondents and never checked against system data. It is neither a share of companies nor a share of budget. 24.2 is the mean of a right-skewed distribution; the median is 20. The question was answered by 191 of the 2,111 people invited, about 9 percent, the expectation question by 188. The same question produced 34.5 and then 44.2 percent in the two preceding waves; the expectation climbs with every wave and none has ever been checked. The sector figures belong to two questions: The 36.1 percent comes from the overall question and the largest sector cell (40 companies), the 8.7 percent from the generative AI one and one of the smallest (3). For the overall question the report gives no low at all.
The CMO Survey, 35th edition, conducted by Christine Moorman at Duke University’s Fuqua School of Business, sponsored by Duke, Deloitte and the American Marketing Association · 2,111 marketing leaders at U.S. for-profit companies invited, 308 responses, 14.6 percent response rate, 191 of them to this question, 97 percent VP level or above · fielded 7 to 29 January 2026, report published April 2026 · Highlights Report page 21 (AI and machine learning) and page 23 (generative AI), figures in the Topline Report page 12, sector cell sizes in the Firm and Industry Breakout Report pages 36 and 45 · Breakout Report
Usually miscited
The most quoted figure on unused corporate data, 55 percent, bundles self-estimates the 1,357 respondents made about their own organisation. It was fielded in 2018/19 by the research arm of the PR agency FleishmanHillard on behalf of Splunk, a vendor selling software to analyse exactly this data.
What the number actually covers: Percent of what stays open, and the report names neither unit nor period nor the statistic used. The only definition put to respondents was “Information that can be captured, quantified and analyzed”. On pages 3 and 12 it reads as a fact, only the regional and country sections from page 10 onwards reveal it as an estimate: US 56 percent, Germany 53, China 50 “compared with a global 55 percent”. An estimate across organisations is not a share of any total, and 1,357 is the sum of the market samples, while the report says 1,300 respondents. What gets routinely mixed in with it: 60 percent of respondents say half or more of their data is dark, 33 percent say 75 percent or more. Those are shares of respondents, not shares of data. Splunk still circulates the figure without a year, on 26 March 2025 as “recent” and there without a sample either.
TRUE Global Intelligence (FleishmanHillard) for Splunk, The State of Dark Data, 1,357 respondents from IT and business across seven countries · fielded October 2018 to January 2019, report May 2019
Survey
78 of 188 US companies that answered this question use AI for generative engine optimization, that is, to get their own content to appear in AI-generated search answers. That is 41.5 percent, with a 95 percent confidence interval of plus or minus 7.1 percentage points.
What the number does not say: It is self-reported, comes from a check-all-that-apply question and records the doing, not the result. The denominator is the trap: it is not the 308 respondents but the 188 who answered this question; all 188 ticked at least one box (response percent 100.0). No company using no AI at all sits in the denominator, so the 41.5 percent is a share among AI users. Computing against 308 yields 128 instead of 78. With the interval the range runs from 34 to 49 percent: a good four in ten, not one in two. Who was asked: US companies only, 97 percent at VP level or above, 308 of 2,111 people contacted, a 14.6 percent response rate. The figure does not transfer to the German market. What is new is the answer option, not the question: GEO was on the list for the first time in 2026 and has no comparison value, while the question itself is reported as a time series against Fall 2023.
The CMO Survey, 35th edition, run at the Fuqua School of Business, Duke University, sponsored by Duke, Deloitte and the American Marketing Association · Topline Report 2026, page 12, answer option “GEO (i.e., Generative Engine Optimization to get content to appear in AI-generated search results)”: 78 of 188 cases, 41.5 percent, plus or minus 7.1 percentage points · 308 respondents out of 2,111 marketing leaders contacted at US for-profit companies · fielded 7 to 29 January 2026 · take care when looking it up: the row above, “Predictive analytics for customer insights”, carries the same 41.5 percent and likewise 78 cases
Survey
In the member survey run by the US advertising association ANA, 82 percent of the members surveyed said in 2023 that they had an in-house agency, after 78 percent in 2018, 58 percent in 2013 and 42 percent in 2008. 65 percent said in 2023 that they had moved ongoing business from an external agency in-house in the preceding three years. In 2018 it was 70 percent, in 2013 only 56.
What the figures do not say: They measure where work sits, not whether it gets better or more on-brand. An industry association surveys its own members, participation is voluntary, and respondents may belong to the in-house agency themselves. Careful with the 65 percent: The ANA states a base for each question, regularly smaller than the participant count. In 2018 the base for this question was 166 of 412 respondents; for 2023 it is not published and must not be applied to the 162 participants. The gap between 70 and 65 percent carries no turning point: the interval runs from 63 to 77 percent in 2018 and, on at most 162 answers, from 58 to 72 percent in 2023, and the waves differ in size. The ANA report of June 2026 does not continue the series, it surveys award jurors; the next member wave would be 2028.
Association of National Advertisers, member survey “The Continued Rise of the In-House Agency: 2023 Edition”, 162 respondents, fielded February and March 2023 · press release of 2 May 2023, the report itself sits behind the membership wall · comparison values and per-question bases from the predecessor report of October 2018, 412 respondents, 57 pages · Predecessor report
Verified study
Of the English-language articles newly published in the first quarter of 2026, 49.9 percent were primarily AI-generated. Since early 2025 the share has moved between 44.6 and 50.9 percent, with no upward trend.
What the number does not say: It measures newly published English-language articles and listicles carrying article markup and at least 100 words, drawn from Common Crawl, not the web as a whole, other languages, social media or video. And nothing about reach: Graphite collected the data in June 2025 and published it in October 2025, finding that 86 percent of articles ranking in Google across 31,493 keywords and 82 percent of those cited by ChatGPT and Perplexity were written by humans, but measured with a fourth detector (Surfer, false positive rate 4.2 percent), so not on the same scale. How firm the 50 percent is: the figure averages three detectors that diverge by 6.4 points in the same quarter (Pangram 47.7, Copyleaks 48.1, GPTZero 54.1). The spread is wider than the distance to the 50 percent mark, and two of the three put humans ahead. GPTZero counts a “mixed” verdict entirely on the AI side (6.4 percent of articles); Pangram and Copyleaks go by which share is larger.
Graphite, growth agency · about 55,400 English-language articles randomly drawn from Common Crawl, carrying article markup and at least 100 words, published between January 2020 and March 2026 · classified by averaging three detectors (Pangram, Copyleaks, GPTZero), whose false positive rates of 1.844, 1.836 and 1.355 percent were measured on about 15,700 articles from the same sample published before ChatGPT · quarterly data openly available · May 2026
Three vendor measurements
Three analytics vendors put a number on the share of website visits that arrive from an AI assistant: Contentsquare 0.2 percent in the fourth quarter of 2025, Semrush 0.14 percent for the year 2025, Conductor 1.08 percent for May to September 2025. Three separately collected measurements, fractions of a percent up to a good one percent.
What the numbers do not say: They count clicks arriving with an identifiable AI referral source, not how often a brand is named in answers. Google AI Mode sits outside the two values that address it: Semrush tracks it as a separate channel at 0.01 percent, and Conductor notes that Google Analytics does not separate it from organic traffic.
Why the three values do not belong side by side: Contentsquare measures 6,500 websites worldwide, Semrush more than 50,000 worldwide, Conductor 1,215 of its own customer domains in the US and calls its figures averages. The near eightfold gap between them, 0.14 to 1.08 percent, is largely a question of who was measured. All three sell analytics tools; Contentsquare and Conductor measure their own customer base, Semrush estimates from a bought-in clickstream panel. None of the values is independently audited.
Contentsquare, 2026 Digital Experience Benchmarks, 99 billion sessions and 6,500 websites worldwide, fourth quarter 2024 against fourth quarter 2025 · 29 January 2026 · Semrush, Traffic & Market Toolkit, more than 50,000 websites and 17 industries worldwide, January to December 2025 · 27 April 2026 · Conductor, AEO/GEO Benchmarks, traffic section from 1,215 of its own enterprise customer domains in the US, May to September 2025 · last updated 6 July 2026
Vendor figures
Takedown provider Netcraft states that between March 2024 and March 2025 it acted against 1.3 million phishing sites imitating more than 16,000 organisations.
What the number does not say: It does not say these sites are gone. Netcraft writes “disrupted” and folds blocking and removal into that one word, how the 1.3 million splits is stated nowhere. That same page tells readers to ask vendors exactly this. It also says nothing about whether a published brand specification makes cloning easier or harder, for which there is no comparison group. All it establishes is that clones exist in bulk. Who counted: Netcraft itself, its own operations, on a page selling that service. Nothing is independently audited. The world share of about a third often quoted alongside is therefore not in the claim: the page never quantifies the world total and contradicts itself, once a share of takedowns, once of attacks. Nor are the 16,000 organisations a market size.
Netcraft, guide to detecting and disrupting phishing websites, the vendor’s own operations from March 2024 to March 2025 · 12 December 2025, last modified 12 March 2026
Market observation on estimated data
Between the first quarter of 2023 and the fourth quarter of 2025, search engine visits and search-like AI sessions combined grew by 26 percent worldwide, from 82.0 to 103.2 billion per month. Google’s share falls from 89 to 71 percent, ChatGPT reaches 20 percent.
What the number does not say: It counts not searches but visits and sessions: a Google visit 6.7 page views on average, an app session an unknown number of prompts. And unevenly: search engines web only, AI web and app. Search apps it excludes as “relatively low”, without a figure, though 83 percent of AI use is in apps. Of AI, only the 52 percent of “asking” count, which the authors call an upper bound. The 71 percent are Google Search and Gemini combined; per day the growth is 23.1, not 26. It cites those same 26 percent elsewhere for 2025 against 2024, where its open raw data yields 17.3. Where the data comes from: Similarweb estimates without server measurement, validated only for the search engine figures, across eight websites, only as a trend correlation; the AI figures not against first-party data at all. The author sells visibility in search engines and AI answers and discloses that calling both large serves him.
Graphite (Ethan Smith) · analysis of Similarweb estimates for web visits and app sessions worldwide, July 2020 to December 2025, raw data public · March 2026
Controlled trial
Experienced developers took 19 percent longer with AI tools — while believing they had been 20 percent faster. Beforehand they had expected a 24 percent speed-up. Between measured and perceived effect lie 43 percentage points, with the sign reversed.
What the number does not say: 16 developers, 246 tasks, exclusively in repositories they had known for five years on average. That familiarity explains part of the result — anyone who holds their own project in their head gains less from assistance. It does not transfer to unfamiliar code or other knowledge work, and the tools date from early 2025. What holds: the gap between measurement and self-assessment. It is the reason to distrust any productivity figure based on asking people.
METR, randomised controlled trial, July 2025 · 16 experienced open-source developers, 246 tasks
Controlled trial
In a preregistered experiment with 758 management consultants, AI users completed 12.2 percent more tasks, worked 25.1 percent faster and delivered more than 30 percent higher quality, as long as the task fell inside the model's capability. On a task placed just outside it, they were 19 percentage points more likely to be wrong than the group without AI.
What the number does not say: Where the boundary runs was not visible to participants — the tasks looked alike. That is both the core finding and its limit: one deliberately out-of-range task was measured, not how often such cases occur in daily work. The experiment used GPT-4; where the frontier sits today is open. And consulting is not all knowledge work.
Dell'Acqua et al., “Navigating the Jagged Technological Frontier”, field experiment with Boston Consulting Group · Organization Science, published online 11 March 2026, DOI 10.1287/orsc.2025.21838 · 758 consultants, 18 realistic tasks · the 2023 working paper still put the quality gain at more than 40 percent
Controlled experiment
293 participants each wrote a short story, some of them with starting ideas from GPT-4, and 600 readers rated them. Stories written with an AI idea were judged more novel, up 5.4 percent with access to one idea and up 8.1 percent with access to up to five. At the same time they converged: a story’s similarity to the mean of the others in its group rose by 0.871 points on a scale from 0 to 100, which the authors report as 10.7 percent of the range the group without AI spanned.
What the number does not say: the 0.871 points apply to access to one idea. With up to five ideas the convergence was smaller, 0.718 points and 8.9 percent, and the creativity gain larger. Reading it as “more AI, more sameness” reads against the data. What is measured is access, not use: in the one-idea condition 82 of 100 requested an idea at all; in the second, 2.55 on average, and only 24.5 percent asked for all five. The more striking number is the softer one: the 10.7 percent is a share of the 8.10 points between the highest and lowest value among the stories without AI, so it hangs on two extreme values. The same 10.7 percent appears a second time in the paper, as a novelty gain among the least creative writers, unrelated to this one. What is measured is the cosine similarity of text embeddings within one group, not whether results are worse. “More creative” is the judgement of lay readers; in the writers’ own assessment there was no statistically significant difference. Eight sentences, no dialogue with the model, British Prolific participants rather than professional writers: short stories are not strategy papers.
Doshi and Hauser, “Generative AI enhances individual creativity but reduces the collective diversity of novel content”, Science Advances, vol. 10, issue 28, eadn5290, 12 July 2024, DOI 10.1126/sciadv.adn5290 · pre-registered, 293 writers and 600 raters on Prolific, 3,519 individual ratings, ideas from GPT-4 · the publisher site science.org refuses automated requests, so the check ran on the open full text at Europe PMC
Editorially reviewed
Seven language models were put to seven strategic trade-offs, each as an either-or. On six of the seven they picked the same side across all vendors, differentiation over cost leadership and augmentation over automation among them. Only on exploration versus exploitation did they diverge. Two follow-up studies on ChatGPT-5, each with more than 15,000 runs, barely moved the bias: for differentiation and augmentation, better prompting lowered the share of biased responses by less than 2 percent, and additional company context shifted it by 11 percent on average across the whole follow-up study, in both directions. The authors call this “strategy trendslop”: the most socially desirable answer of the internet average.
Corrected on 10 September 2026: until then this entry said “differentiation in 96 percent of cases, augmentation in 93”. Neither figure appears anywhere in the article; the full text gives no share at all, only shifts from baseline, and the numbers come from blog summaries, most likely read off a chart. Also corrected: the more than 15,000 runs do not apply to the seven models. Both 15,000-run blocks are follow-up studies on a single model, ChatGPT-5; the seven-model measurement rests on 50 runs per model and question, and the authors give no total for it. What the numbers do not say: the insensitivity holds for only two of the seven questions; on the other five, better prompting moved responses by 22 percent on average in both directions, and the order of the options mattered most, at 19 percent. Whether the preferred answer is wrong is not what the finding says; what is measured is insensitivity to context alone. On the source: an HBR Digital Article without peer review, and there is no separate paper; one repository lists it as peer reviewed, which is a catalogue artefact. The paywalled page ships the full text in its structured data, character-identical in archive copies of 17 March and 14 July 2026.
Angelo Romasanta, Llewellyn D. W. Thomas and Natalia Levina, “Researchers Asked LLMs for Strategic Advice. They Got ‚Trendslop‘ in Return.”, Harvard Business Review, 16 March 2026, HBR Digital Article H093GG · models tested: ChatGPT, Claude, DeepSeek, GPT-5 via the API, Gemini, Grok and Mistral, 50 runs per model and question; the two blocks of more than 15,000 runs ran on ChatGPT-5 alone · full text of 15,613 characters from the structured data of the page, cross-checked against two archive copies
Controlled experiment and vendor documentation
Agreement is rewarded in the training signal. In an analysis of 15,000 response pairs from Anthropic’s own feedback data, matching the user’s beliefs is consistently among the strongest predictors of human preference; any single feature shifts the probability of preference by at most about 6 percentage points. In April 2025 OpenAI rolled back a GPT-4o update because the model had become excessively agreeable, naming as its early assessment three changes acting together, among them an additional reward signal from users’ thumbs-up and thumbs-down feedback.
What the number does not say: the study comes from Anthropic, two of the five assistants tested and the reward model analysed are its own, and every model dates from 2023. Truthfulness is a rewarded feature too, and depending on the condition, matching beliefs is not the strongest one. The paper tests factual questions and free-text tasks, not strategy advice. What is measured is what the data reward, not how often an assistant flatters in production; the OpenAI rollback shows only that the effect can occur there and be noticed, not how strong it is today. Care with the most-quoted figure: in the sub-experiment on 266 misconceptions the agreeing responses were produced deliberately: a model was instructed to deceive subtly, and the most convincing of 4,096 samples was selected. The 95 percent is the judgement of the Claude 2 preference model, not of humans, measured against the best of three short human-written objections, chosen by that same model. For the human raters the paper gives no figure; they mostly preferred the correcting response and did so less reliably as difficulty rose, read off the chart at roughly 3 percent on the easiest and roughly 21 percent on the hardest level. That is the majority of several lay readers without reference material; the average individual rater sits higher. The authors call the dataset a proof of concept.
Sharma, Tong et al., “Towards Understanding Sycophancy in Language Models”, arXiv:2310.13548v4 of 10 May 2025, first version 20 October 2023, peer reviewed and accepted as a poster at ICLR 2024 · key figure from section 4.1: 15,000 randomly drawn response pairs from the helpfulness subset of Anthropic’s hh-rlhf dataset, 23 features, Bayesian logistic regression, holdout accuracy 71.3 percent · models tested: Claude 1.3, Claude 2, GPT-3.5, GPT-4 and LLaMA 2 · plus OpenAI, “Sycophancy in GPT-4o”, 29 April 2025, and “Expanding on what we missed with sycophancy”, 2 May 2025; update of 25 April, rollback from 28 April (OpenAI statement)
Negative finding of a documented search
Whether a brand as a specification produces more brand-compliant AI output than a brand book is settled by no publicly verifiable measurement; the search finds no public benchmark for it. Neighbouring fields have such benchmarks: guideline adherence in medicine since December 2024, rule adherence in support dialogues since early 2026.
What the finding does not say: that nothing exists. Searched on 10 September 2026 with 24 phrase queries; Semantic Scholar (HTTP 429), the ACL Anthology, subscription databases and unpublished vendor studies remain unchecked. Against our own thesis: the general form question has been measured. A benchmark of 31 July 2026 stacks 24 machine-checkable instructions, one of them a fixed tone of voice, and tests whether the same instructions are followed better in compiled form: up to 11 percentage points more adherence on the weakest model, practically nothing on strong ones, and by its own account never compared against a competently hand-written prompt. So the question is not wide open. No neighbour measures brand: the three closest benchmarks measure rule adherence in support dialogues and the detection of violations by a model acting as judge. All three datasets are machine-generated, CompliBench is a preprint, and two of its eight authors work for a contact-centre software vendor. Adobe’s brand-compliance score is a product feature, not open to inspection. The opposite is measured no better.
Own research, 10 September 2026: arXiv API with 24 phrase queries, among them brand voice, brand guidelines, brand consistency, brand compliance, machine-readable brand, brand book, style guide, tone of voice and corporate identity; plus OpenAlex, general web search, and the vendor pages of Adobe, Frontify, Jasper and Writer · closest measured work on the form question: “Instruction Stacking Collapse”, arXiv:2608.02639, 31 July 2026, 24 verifier-checked instructions, three models (to the paper) · closest benchmarks on rule adherence: CompliBench, arXiv:2604.12312, 14 April 2026, preprint, 318 machine-generated dialogues (CompliBench), PluralisticBehaviorSuite, arXiv:2511.05018, 300 behavioural guidelines across 30 industries, and JourneyBench, arXiv:2601.00596, 703 conversations · for comparison in medicine: AMEGA, npj Digital Medicine 7:358, 12 December 2024, and CPGBench, arXiv:2603.25196, 26 March 2026, 3,418 guideline documents
Historical study
A century of household technology did not reduce time spent on housework. The appliances mainly replaced work done by men, children and servants; the time saved went into rising standards of cleanliness and care. The expectation that automation frees up time has a documented precedent in which precisely that failed to happen.
What the study does not say: It concerns households between the open hearth and the microwave, not knowledge work and not AI. Transferring it draws an analogy, not a proof. It serves as a corrective, not a forecast: it shows that labour saving through technology is an assumption that has historically failed once — not that it will fail again.
Ruth Schwartz Cowan, “More Work for Mother: The Ironies of Household Technology from the Open Hearth to the Microwave”, 1983 · awarded the Dexter Prize of the Society for the History of Technology, 1984
Longitudinal study
Corporations whose future preparedness was rated strong in 2008 reached 16 percent profitability by 2015, against 12 percent for the industry average — 33 percent more. On market capitalisation growth over the same seven years they stood at 75 percent against an average of 25. Firms with identified deficiencies came in 37 to 44 percent below average.
What the widely cited figure omits: of 83 corporations surveyed, matching against performance data left 70 for profitability and only 42 for market capitalisation growth. The authors themselves call this an important limitation. The “200 percent additional growth” circulating in the foresight industry therefore rests on 42 companies. And it is a correlation, not a cause: corporations that can afford futures work differ in other ways from those that cannot.
Rohrbeck and Kum, “Corporate foresight and its impact on firm performance: A longitudinal analysis”, Technological Forecasting & Social Change 129, 2018 · preparedness measured 2008, performance 2015
Ongoing measurement
On the ForecastBench tournament leaderboard, the median of human superforecasters sits fourth at 68.8, behind three Google DeepMind entries allowed to use tools and extra context. On the base leaderboard, without tools, the human median leads at 67.8 against the best model at 62.6.
What the leaderboard does not say: the machines’ lead is not established. Its own significance column finds no difference for the three top places, at p values of 0.71, 0.60 and 0.57, and the confidence intervals almost fully overlap. The Brier Index is not a hit rate despite the percent sign, and the conversion is non-linear. A best-of-many result: Google DeepMind holds 50 of 334 entries and all three places ahead of the humans, whose comparison group is a single row. The questions differ too: the humans were last surveyed in July 2024, 578 questions against 790 and 1,165. On parity: the operators record parity reached on 7 June 2026 for the tournament evaluation and project it for the tool-free one to February 2028, interval July 2026 to December 2030, which reflects only the uncertainty of the line fit. And forecasting dated events is not a strategic judgement about a brand.
ForecastBench, Forecasting Research Institute, tournament and base leaderboards, retrieved 10 September 2026 · entries appear only 50 days after submission, so the leaderboard is not a same-day state
Preregistered experiment
991 participants answered six forecasting questions, some with access to a language model. The assistance improved accuracy by 24 to 28 percent against the control group. The comparison within the groups is the notable part: an assistant deliberately tuned to be overconfident and noisy also helped substantially.
What the result suggests but does not prove: That the poor assistant also worked points to part of the gain coming from the act of consulting rather than the quality of the machine's answer — but that is not established. The authors themselves note that outliers affect the picture and robustness remains to be tested. Six questions are a narrow base, and this is a preprint.
Schoenegger, Park, Karger, Trott and Tetlock, “AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy”, preregistered, arXiv, February 2024 · 991 participants
Standard reference
According to Andrew Abbott's study, occupations compete not over performing their work but over its definition. Whoever determines what counts as a problem, and who is responsible for it, has already settled the competition. Abbott's finding on how this happens: jurisdictions are claimed when they fall vacant — not by filling an occupied one better.
What the study does not say: It dates from 1988 and treats classical professions — medicine, law, accountancy — not consulting and not AI. Applying it to today's occupations is an interpretation. And it does not explain how a jurisdiction is won, only what the competition is about. It serves as evidence for what definitional power means, not as a manual.
Andrew Abbott, “The System of Professions: An Essay on the Division of Expert Labor”, University of Chicago Press, 1988
Scholarly book
The sociologist Elena Esposito considers the analogy between algorithms and human intelligence misleading and proposes a different term: artificial communication. In her words: if machines contribute to social intelligence, it will not be because they have learned to think like us, but because we have learned to communicate with them.
What the book is not: not a measurement but a theoretical proposal from systems theory. It supports no figure and cannot be refuted like an experiment. And it is not about brands: Esposito's examples are recommendation lists, profiling and the right to be forgotten. Applying it to whether a brand is legible to machines is an interpretation — a plausible one, but not one the book makes.
Elena Esposito, “Artificial Communication: How Algorithms Produce Social Intelligence”, MIT Press, 24 May 2022 · 200 pages, open access edition available
Field study, working paper
A field study of 244 management consultants found three ways of working with generative AI — differing not in the tool but in who steers the workflow. Those who involve the AI throughout acquire new AI capability. Those who use it selectively for individual steps, keeping the problem definition themselves, deepen their existing domain expertise. Those who hand over the whole process build neither.
What the study does not say: it does not compare the quality of outputs. The abstract makes no claim about which mode produces more accurate recommendations — summaries in circulation that say otherwise go beyond the source. What is measured is capability building, not results. Also: a working paper in draft form, not peer reviewed, and all respondents come from a single consultancy.
Randazzo, Lifshitz, Kellogg, Dell'Acqua, Mollick, Candelon and Lakhani, “Cyborgs, Centaurs and Self-Automators”, Harvard Business School Working Paper 26-036, 2025 · 244 Boston Consulting Group consultants
Field experiment
Among 5,179 customer support agents at a large firm, issues resolved per hour rose 14 percent on average with an AI assistant. The average hides the point: novice and low-skilled workers gained 34 percent, while the effect on experienced and highly skilled workers was “minimal”. The authors suspect the model disseminates the practices of abler workers to newer ones.
What the number does not say: Customer support is highly structured work with recurring cases — transfer to strategy or design is open. And it is a working paper, explicitly not peer reviewed per its cover page. What was measured is volume, not quality: issues resolved per hour, not how well — though customer sentiment improved alongside.
Brynjolfsson, Li and Raymond, “Generative AI at Work”, NBER Working Paper 31161, April 2023, revised November 2023 · 5,179 customer support agents
Preregistered experiment
In a preregistered experiment with 444 college-educated professionals, time spent on writing tasks fell by 0.8 standard deviations while quality rose by 0.4. Here too the gap between participants narrowed — weaker performers gained more. The authors' reading: the tool mostly substitutes for effort rather than complementing skill, shifting work away from rough drafting towards idea generation and editing.
On the figures: these are from the verified working paper of March 2023, which states it is not peer reviewed. The peer-reviewed version appeared later in Science with differing numbers — 453 participants and “40 percent time saved” circulate; that version sits behind a paywall and was not inspected. What the figures do not say: these were short, isolated writing tasks, not projects spanning weeks.
Noy and Zhang, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence”, MIT, working paper of 2 March 2023 · peer-reviewed version in Science 381, 2023, pp. 187–192
Controlled test
In an ideas contest on the circular economy, 300 screened evaluators each rated 13 of 234 solutions, 3,900 ratings in total: 54 from people, 180 from GPT-4 with human-guided prompts. The human ones were judged more novel (the machine ones minus 0.140 on a scale of 1 to 5), the machine ones more strategically viable, more valuable environmentally and financially, and better overall (plus 0.088 to 0.160). At the top end the picture flips: AI solutions received the top novelty mark 7.9 percentage points less often, and their value advantage vanished there across all four dimensions.
What the numbers do not say: The comparison was not human against machine but the crowd against human-guided AI with purpose-built prompts. When the model was iteratively told to differentiate, the gap was no longer detectable on average (minus 0.056) and remained only at the top mark. What they rest on: Judgements about texts, not realised ideas. All evaluators are based in the United States. GPT-4 as of mid-2023, a single task domain, ten AI against three human solutions per block. Two of the five authors are listed with the AI firm involved; co-author Jacimovic founded it.
Boussioux, Lane, Zhang, Jacimovic and Lakhani, “The Crowdless Future? Generative AI and Creative Problem-Solving”, Organization Science 35(5), pp. 1589–1607 · 234 solutions evaluated, 300 evaluators, 3,900 ratings, contest 30 January to 15 May 2023 · received 30 November 2023, revised 23 January, 14 May and 20 June 2024, accepted 26 June 2024, online 16 August 2024, September/October 2024 issue
Exploratory case studies, concept-forming
Henry Mintzberg defined strategy in 1978 as “a pattern in a stream of decisions”: a strategy has formed once a sequence of decisions shows consistency over time. That opens to research the strategies which came about despite intentions, or with no intention at all. He showed it on two long-run cases, Volkswagenwerk and the United States in Vietnam from 1950 to 1973.
What the paper does not measure: It forms concepts and describes. Nowhere does it show that a strategy which grew produces better results than a planned one. Mintzberg does attack planning theory, its split between formulation and implementation resting on two assumptions that often prove false. He did not test that. The foundation is larger than what is shown: The general conclusions rest on four funded major studies and over twenty student papers; only two of the four are documented in the text, the rest are not. Both cases shown are historical, lie outside brand management, and were reconstructed in hindsight by the same research group that set what counts as a pattern.
Henry Mintzberg, “Patterns in Strategy Formation”, Management Science, Vol. 24, No. 9, pp. 934 to 948 · four major studies funded by the Canada Council and over twenty student papers, two of them presented: Volkswagenwerk 1934 to 1974 per the abstract and 1920 to 1974 per the section heading, the United States in Vietnam 1950 to 1973 · manuscript received 19 April 1976, printed May 1978
Survey
Across 17 product categories in Australia and the UK, an average of 11 percent of a brand’s current users consider it different and 10 percent consider it unique; 17 percent name at least one of the two. They buy the brand anyway. The authors recommend distinctiveness instead.
What the figure does not say: It does not show that buyers see no differences at all. On average 54 percent credit at least one brand in the category with one of the two, and 76 percent for soft drinks in the UK. The per-brand score is low because each respondent names only one or two brands, and each names different ones; across categories, 8 to 36 percent. Only users were surveyed; that they buy the brand anyway follows from that, not from the measurement. Who collected the data and who paid: The accessible text gives no sample sizes. Most data, and the measure itself, come from the advertising agency Young & Rubicam, collected in 1999; the institute lists Coca-Cola, Mars and Nielsen among its funders. The paper says nothing about machine selection, it predates every agent.
Jenni Romaniuk, Byron Sharp and Andrew Ehrenberg, Ehrenberg-Bass Institute, Australasian Marketing Journal 15 (2), pages 42 to 54 · survey of current brand users across 17 product categories, the Australian ones by telephone, the UK ones from the Young & Rubicam Brand Asset Valuator, collected 1999 · 2007
End of list: 92 of 92 entries. Last entry: wahrgenommene-differenzierung.
No entry matches this selection.