# Evidence: all 108 claims

Overview of the collection “Evidence” by Robert Haase, as of 21 September 2026: every claim on one line, with its grade and a link to its topic file.

A claim holds only together with its limit (“What the number does not say”). Limit and source are given in full in the topic file named on each line.

Page: https://robert-haase.de/en/evidence.html · JSON: https://robert-haase.de/en/evidence.json · Deutsch: https://robert-haase.de/belege.md

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Please cite the primary source, not this page.

This file is generated from the page. Where the two differ, the page applies.

## How this collection is built

Every figure is traced back to the body that measured it, not to the article citing it. On their way through the retellings, figures lose their denominator first, then their caveat, and finally their origin. Where a figure is only accessible through a third party, that intermediary is named in the source line. Own measurements carry their method with them; they have not been independently verified yet.

**The limit belongs to the number.** The most common error is not the wrong number but the right one carrying a claim that reaches further than the evidence. That is why every entry has two parts, and the second one matters more. Above each figure sits what kind of evidence it is, from verified study to single case. That decides how far it carries.

What does not survive the check does not get in, or gets taken out, my own articles included. One of them claimed that 44 percent of US online shoppers begin their purchase journey in a language model, attributed to Bain. Bain gives two other figures, 17 percent and 30 to 45 percent, which had merged into one along the way. Both are here now; the 44 is not.

This page ages. Every entry carries its date; superseded numbers get replaced, not quietly deleted. If you find an error, [write to me](mailto:hallo@robert-haase.de) and I will correct it and note the date.

The collection does not map the state of the research, only the figures I needed for my own texts. Free to use with attribution. When in doubt, link the primary source rather than this page.

## Topic files

- AI search, 17 entries: https://robert-haase.de/en/evidence-ai-search.md
- Agents, 36 entries: https://robert-haase.de/en/evidence-agents.md
- Commerce, 7 entries: https://robert-haase.de/en/evidence-commerce.md
- Liability, 12 entries: https://robert-haase.de/en/evidence-liability.md
- Market size, 16 entries: https://robert-haase.de/en/evidence-market.md
- Judgement, 20 entries: https://robert-haase.de/en/evidence-judgement.md

## AI search (17)

+ Verified study, vendor documentation, or court decision (9) → llmstxt-abrufe, llmstxt-wirkung, google-leitfaden, mentions-vs-backlinks, inkonsistenz, pew-klicks, aio-klickrate, seer-klickrate, ebu-nachrichten
+ Preliminary: prototype, single test, forecast, or vendor figure (7) → json-ld-test, reddit-zitate, geo-40-prozent, aio-top10-uneinig, markenstatur-sichtbarkeit, zitier-position, eigene-seite-selten-zitiert
+ Status, case report, or market observation (1) → llmstxt-nutzen

- **llmstxt-abrufe** · Verified study · Of roughly 38,000 domains that have an llms.txt, 97 percent saw no request for the file at all in May 2026. → https://robert-haase.de/en/evidence-ai-search.md
- **llmstxt-wirkung** · Verified study · Across nearly 300,000 domains studied, no relationship was found between having an llms.txt and how often a domain appeared as a source in AI answers. → https://robert-haase.de/en/evidence-ai-search.md
- **google-leitfaden** · Vendor documentation · Google explicitly states that llms.txt, special markup, and custom chunking are not needed for AI search. → https://robert-haase.de/en/evidence-ai-search.md
- **json-ld-test** · Single test · 1,885 pages that added JSON-LD barely moved against 4,000 control pages: no effect distinguishable from zero in ChatGPT and Google AI Mode, and a statistically significant 4.6 percent decline in AI Overviews. → https://robert-haase.de/en/evidence-ai-search.md
- **mentions-vs-backlinks** · Verified study · Third-party mentions correlate with visibility in AI answers far more strongly than the classic metrics of a brand’s own site: 0.66 to 0.74 against 0.27 to 0.33 for domain authority and 0.19 for the number of pages. → https://robert-haase.de/en/evidence-ai-search.md
- **inkonsistenz** · Verified study · Ask the identical question twice and the chance of getting the same list of brands is under one in a hundred. → https://robert-haase.de/en/evidence-ai-search.md
- **pew-klicks** · Verified study · When an AI summary appears, users click a traditional search result on 8 percent of visits. Without a summary it is 15 percent. → https://robert-haase.de/en/evidence-ai-search.md
- **aio-klickrate** · Verified study · Where an AI summary sits above the results, the click-through rate of the first organic position is about 58 percent lower. At position 2 it is 50.8 percent, at position 3 46.4 percent. → https://robert-haase.de/en/evidence-ai-search.md
- **seer-klickrate** · Verified study · For informational queries carrying an AI summary, the organic click-through rate fell from 1.76 to 0.61 percent, a drop of 61 percent. For queries without a summary it fell from 2.74 to 1.62 percent over the same period, a drop of 41 percent. → https://robert-haase.de/en/evidence-ai-search.md
- **reddit-zitate** · Vendor measurement, preliminary · Reddit’s share of the sources ChatGPT cites fell from 3.83 to 0.52 percent within a few days. On 8 August 2026 ChatGPT’s use of the site: operator jumped from 0.37 to 16.8 percent of its derived search queries. A second vendor panel counts a drop in daily Reddit citations from 497 to 132 over the same window, with total citation volume up by 3.5 percent. → https://robert-haase.de/en/evidence-ai-search.md
- **geo-40-prozent** · Usually miscited · The most quoted figure in the AI visibility business, "up to 40 percent more visibility", comes from a lab setup running a since-retired model and measures a purpose-built metric. → https://robert-haase.de/en/evidence-ai-search.md
- **ebu-nachrichten** · Verified study · 45 percent of answers from four AI assistants to news questions had at least one significant issue. In 31 percent it concerned the handling of sources. → https://robert-haase.de/en/evidence-ai-search.md
- **aio-top10-uneinig** · Vendor measurements, not comparable · How many of the sources cited in Google’s AI Overviews also rank in the organic top 10 is measured incompatibly by the two large vendors: BrightEdge around 17 percent on average from February 2025 to February 2026, Ahrefs 37.1 percent in its study of 2 March 2026. In the one month both report a figure for, the gap is widest: for July 2025 BrightEdge reports around 16.6 percent, Ahrefs 76.1 percent in its study from the same month. → https://robert-haase.de/en/evidence-ai-search.md
- **llmstxt-nutzen** · Vendor-run controlled test, result data open · Across 20 documentation sites, two coding agents worked markedly leaner once their pages pointed to their llms.txt at the very top. Claude Code took on average 25.4 seconds instead of 30.8 and 175,000 tokens instead of 215,000 per task; Codex took 45.1 seconds instead of 57.1 and 115,000 instead of 156,000. That is 18 to 26 percent less. → https://robert-haase.de/en/evidence-ai-search.md
- **markenstatur-sichtbarkeit** · Vendor measurement, preprint without peer review · Ask an AI search engine a category question without naming the brand, and globally known brands appear on average in 72.9 percent of answers on the first tracking run, established mid-market and regional brands in 43.6 percent, small and niche brands in 11.4 percent. Of all 149,912 citations counted, 2.9 percent point at the brand’s own website and 75.2 percent at those of other companies in the same category. → https://robert-haase.de/en/evidence-ai-search.md
- **zitier-position** · Vendor measurement, preliminary · Of 18,012 citations ChatGPT drew from web pages, 44.2 percent come from the first 30 percent of the text. The middle section, the widest at 40 percent of the text, carries 31.1 percent, the closing section 24.7 percent. In a second analysis of 11,022 citations, cited introductions reached a proper-noun density of 20.6 percent, against the 5 to 8 percent the author derives from standard corpora (Brown Corpus, Penn Treebank), with no arithmetic shown. → https://robert-haase.de/en/evidence-ai-search.md
- **eigene-seite-selten-zitiert** · Two vendor measurements, preliminary · When AI search recommends brands, it rarely relies on the brand’s own website. At AirOps, for queries in which users look for and compare vendors, 85 percent of 21,311 brand mentions in ChatGPT, Claude and Perplexity came from third-party sources and 13.2 percent from the brand’s own domain. At Ranqo, only 2.9 percent of 149,912 source citations from five AI search engines point to the brand’s own domain, and 75.2 percent to pages of other companies in the same field. → https://robert-haase.de/en/evidence-ai-search.md

## Agents (36)

+ Verified study, vendor documentation, or court decision (26) → leere-buttons, javascript, lighthouse, a11y-tree, astryx-agenten, designsysteme-maschinenschnittstelle, dax-zutritt, dax-benennung, dax-landmarken, dax-bilder, verlage-robots, marken-robots, agenten-erfolg, frontify-mcp, canva-mcp, mcp-tool-poisoning, pulumi-brand-mcp, statista-mcp, rechtsvorbehalt-kommentar, google-ads-textregeln, microsoft-brand-kit, muse-connectors, adobe-markenpruefung, markup-ai-stilpruefung, lighthouse-agent-discovery, content-signal-selten
+ Preliminary: prototype, single test, forecast, or vendor figure (7) → agent-ready, a11y-cua, klarna-700, monotype-mcp, veeva-mlr, abruf-kuerzung, olivares-access-map
+ Status, case report, or market observation (3) → gitlab-markenrepo, aipref, mcp-primitive

- **leere-buttons** · Verified study · On 30.6 percent of one million home pages surveyed, buttons had no accessible name; on 51 percent, form fields had no label. → https://robert-haase.de/en/evidence-agents.md
- **agent-ready** · Preliminary, prototype · In a controlled experiment, three browser agents reached a strict success rate of 89.3 percent on the agent-friendly version against 49.3 percent on the original. → https://robert-haase.de/en/evidence-agents.md
- **a11y-cua** · Usually miscited · The widely cited drop in agent success from 78 to 42 percent comes from a study that changed no website at all. → https://robert-haase.de/en/evidence-agents.md
- **javascript** · Controlled test · Seven widely used US AI assistants execute no JavaScript on a user-triggered fetch and read only the raw HTML. Five others do execute it. → https://robert-haase.de/en/evidence-agents.md
- **lighthouse** · Vendor documentation · Since version 13.3.0 of 7 May 2026, Lighthouse ships an "Agentic Browsing" category in its default configuration. → https://robert-haase.de/en/evidence-agents.md
- **a11y-tree** · Vendor documentation · Google names three ways agents perceive a page: screenshots, raw HTML, and the accessibility tree. Modern agents combine them. → https://robert-haase.de/en/evidence-agents.md
- **astryx-agenten** · Vendor documentation · Meta open-sourced its design system in June 2026, after eight years of internal growth, and justifies how it is built expressly by agents: design systems were historically made for human consumption, and as more code is written by agents, their structure has to be rethought. The system is operated from the command line or over MCP. → https://robert-haase.de/en/evidence-agents.md
- **designsysteme-maschinenschnittstelle** · Verified study · Of 21 open-source design systems surveyed, 18 ship a first-party MCP server, 18 official agent skills and 15 an llms.txt. The survey sums it up: “Nobody is still arguing about whether to ship a machine interface.” → https://robert-haase.de/en/evidence-agents.md
- **gitlab-markenrepo** · Documented single case · GitLab keeps brand voice, naming rules, trademark guidelines, values and mission as Markdown files in a public repository. Every change carries a date, a real name and a written justification. One example: on 22 May 2025 the three brand personality traits were rewritten, in a single commit of four inserted and five deleted lines, filed by Senior Brand Manager Betsy Bula, justified by aligning the definitions with current communication and the FY26 company plan. → https://robert-haase.de/en/evidence-agents.md
- **aipref** · State of standardization · A common standard for how websites permit or refuse AI use of their content still does not exist. → https://robert-haase.de/en/evidence-agents.md
- **dax-zutritt** · Own survey, reproducible · Of 160 home pages requested, 12 reject an automated retrieval with active bot defence, or 7.5 percent. It comes from Akamai or Cloudflare on ten of the twelve pages, and from Amazon CloudFront at Siemens and Hannover Rück. On 18 of 22 pages checked at HTTP level the same rejection came regardless of the browser string; there the detection works from the TLS fingerprint and the order of the HTTP headers. On the two CloudFront pages the browser string alone decides. → https://robert-haase.de/en/evidence-agents.md
- **dax-benennung** · Own survey, reproducible · Across 135 home pages of German listed companies from the DAX, MDAX and SDAX, 456 of 13,527 controls carry no name in the accessibility tree, or 3.4 percent. The rate barely differs between the three indices: DAX 3.1, MDAX 3.8, SDAX 3.3 percent. → https://robert-haase.de/en/evidence-agents.md
- **dax-landmarken** · Own survey, reproducible · 52 of 135 home pages of German listed companies have no main-content landmark. For a program reading the page, the marker for where content begins and navigation ends is missing. → https://robert-haase.de/en/evidence-agents.md
- **dax-bilder** · Own measurement, reproducible · Across 135 home pages of German listed companies from DAX, MDAX and SDAX, 1,662 of the 3,979 images Chrome exposes in the accessibility tree carry no name, so 41.8 percent. Among the unnamed images then inspected by element type, 86 percent are inline SVG graphics and 13 percent classic img elements. → https://robert-haase.de/en/evidence-agents.md
- **verlage-robots** · Own survey, reproducible · Of 76 assessable German-language news and trade media, 44 block at least one training crawler in their robots.txt, or 57.9 percent. 38 of them block GPTBot, exactly half, 40 block CCBot and 36 Bytespider. Far fewer block the same provider’s search bot: OAI-SearchBot appears on 8 outlets’ lists, or 10.5 percent. 11 outlets block training without blocking a single AI search bot or user-triggered fetch. → https://robert-haase.de/en/evidence-agents.md
- **marken-robots** · Own survey, reproducible · Of 148 assessable home pages of the companies in DAX, MDAX and SDAX, 10 block at least one training crawler, or 6.8 percent; 4 block GPTBot. 138 block no AI access at all, among them 14 that serve no robots.txt whatsoever. 7 companies write an express permission for an AI crawler into the file, 6 of which block none at the same time: there are almost as many invitations as blocks. The only reservation of text and data mining rights to be found in the index sits with an academic publisher, and that publisher blocks no crawler at all. → https://robert-haase.de/en/evidence-agents.md
- **agenten-erfolg** · Verified study · Across 300 tasks on 136 real websites, the success rate of the best web agents rose from 61 to 97.7 percent in just over 16 months. When the benchmark was first evaluated in March 2025, one agent reported 89 percent for itself and scored 30 when measured; most did not beat a simple agent from early 2024. By August 2026 the leading entry solves even the hardest tasks — those needing eleven steps or more — completely. → https://robert-haase.de/en/evidence-agents.md
- **frontify-mcp** · Vendor documentation · Frontify opens its brand portal through an MCP server it runs itself. On 17 September 2026 it lists 53 tools one by one in ten packs, graded from read-only to full administrative access; on 10 September it was 54. The read-only Discovery pack holds 24 tools, the Admin pack all 53, two of them flagged as destructive. → https://robert-haase.de/en/evidence-agents.md
- **canva-mcp** · Vendor documentation · Canva runs an official MCP server and documents 33 tools for it. 27 are available on every plan, among them creating and exporting designs. Four require at least Canva Pro, among them listing brand kits and using brand templates. Two are reserved for Enterprise: autofilling a template with data and reading the associated dataset. Every user authenticates individually, and an agent holds the permissions of the human signed in. → https://robert-haase.de/en/evidence-agents.md
- **klarna-700** · Usually miscited · Klarna’s most-quoted AI number is an estimate, not a headcount: the press release of February 2024 states “the equivalent work of 700 full-time agents”. The same measure appears as over 700 in the IPO prospectus of September 2025, and in the annual report of February 2026 still at over 700 in the business section and at over 850 in the operating review of that same report. The headcount sits beside it: approximately 5,527 full-time employees at the end of 2022, approximately 2,831 at the end of 2025. → https://robert-haase.de/en/evidence-agents.md
- **mcp-primitive** · State of standardization · The Model Context Protocol defines three server building blocks, each with an intended controlling party: tools are invoked by the model, resources are steered by the application, prompt templates are selected by the user. That is not binding. All three chapters carry the same trailing clause: the protocol itself does not mandate any specific user interaction model. In the tools chapter a SHOULD rule follows immediately: a human should always be able to deny a tool invocation. → https://robert-haase.de/en/evidence-agents.md
- **mcp-tool-poisoning** · Peer-reviewed benchmark and vendor test · Instructions hidden inside a tool description make an agent read the user’s private SSH key and pass it to a foreign server through a parameter named “sidenote”; the confirmation dialog shows only the name of an addition tool. A benchmark built on 45 live MCP servers with 353 tools measures a 36.5 percent average attack success rate across 20 model settings, 72.8 percent at most. → https://robert-haase.de/en/evidence-agents.md
- **monotype-mcp** · Vendor statement, beta · Monotype announced a beta of its Enterprise MCP Connector on 15 July 2026. It links AI tools to a customer’s font library, its licensing information and its production approvals: it matches AI-generated drafts against the library, reviews referenced fonts against the production font list, and returns CSS in chat when the project fonts are part of the library. It runs on the Model Context Protocol, initially in Claude and Claude Design. → https://robert-haase.de/en/evidence-agents.md
- **pulumi-brand-mcp** · Own measurement, reproducible · Pulumi publishes its own brand guidelines as an MCP server at brand.pulumi.com/mcp. On 10 and again on 17 September 2026 it answered without any login and listed the same 13 resources, one template, 11 tools and 3 prompts; the resources include brand voice, writing style and the binding product names. One resource governs generative AI in plain language, addressed to the human: “never ship raw model output as a finished piece”, “never publish anything without a human reviewing it first”. → https://robert-haase.de/en/evidence-agents.md
- **statista-mcp** · Vendor documentation · Statista runs an MCP server at api.statista.ai/v1/mcp with six documented tools. Every call is metered individually in credits, tiered by the kind of answer: a search costs 0 or 1 credit, retrieving the figures themselves 10 to 15. Without a key the server replies 401 Unauthorized. → https://robert-haase.de/en/evidence-agents.md
- **veeva-mlr** · Vendor press releases · In the regulated pharmaceutical approval process, machine pre-checking of brand rules is a shipping product. On 3 December 2025 Veeva announced a Quick Check Agent that scans content against editorial, brand, market, channel and compliance guidelines before the MLR review itself begins. On 23 June 2026 Veeva acquired the vendor Copli and launched it as Falcon MLR, with the stated potential to eliminate 70 per cent or more of manual MLR labour within five years. → https://robert-haase.de/en/evidence-agents.md
- **rechtsvorbehalt-kommentar** · Own survey, reproducible · Of 77 German-language news and trade media, 20 declare a reservation of rights against text and data mining in their robots.txt, as a comment line: 16 name section 44b of the German Copyright Act explicitly, four others invoke Austrian or European law or state the reservation without naming a section. Among the home pages of the DAX, MDAX and SDAX companies, not a single one does. One machine-readable form, the *TDM-policy* line in the same file, appears in none of the files examined. → https://robert-haase.de/en/evidence-agents.md
- **abruf-kuerzung** · Own test, one tool · An agent’s standard fetch tool read only the front part of a page with 128,000 characters of visible text. A probe by hand the same morning found the cut at entry 65 of 90; the tool itself reported “about 80” entries and put the cut between 100,000 and 115,000 characters. In an acceptance test with ten fixed questions, each asked twice, it pointed out the truncation for only three of them, although a visible sentence on the page named exactly the marker for detecting it. With an anchor card, questions about a specific entry led to the right file in 6 of 6 cases, counting questions in 4 of 4. For conceptual questions it kept answering from the truncated text without mentioning the cut. → https://robert-haase.de/en/evidence-agents.md
- **google-ads-textregeln** · Vendor documentation · In Google Ads, brand rules can be stored in plain language, in Performance Max campaigns and in Search campaigns with AI Max: up to 25 term exclusions and up to 40 restrictions naming concepts, associations or styles to avoid, each per campaign. Both are exclusions, they prescribe nothing, and they apply only to automatically customized text assets, not to images. On the same help page Google warns that unsuitable guidelines may remove a large number of good text assets and hurt performance. → https://robert-haase.de/en/evidence-agents.md
- **microsoft-brand-kit** · Vendor documentation · Microsoft’s Copilot learns a brand from exactly one PDF. A brand manager uploads the guidelines, Copilot extracts colour palettes, styles, brand voice and the rules for logo and typography. Only a single guideline document is supported: to add a new one the existing one has to be removed, and the uploaded guidelines override the values already in the brand kit, brand voice included. → https://robert-haase.de/en/evidence-agents.md
- **muse-connectors** · Vendor documentation · Meta launched its personal agent Muse in the US on 8 September 2026. People choose which apps it connects to and how much access it gets, and it checks back before sensitive steps such as sending an email or making a purchase. Businesses can submit their own connectors for Muse: Meta tests them against functional, security and legal requirements, after which people find them in Muse. → https://robert-haase.de/en/evidence-agents.md
- **olivares-access-map** · Vendor statement, beta · There is software for the permissions of AI agents. Olivares discovers agents, sessions, models, MCP servers, tools and identities running in an organisation, keeps a map of read and write access and sets it against what was actually observed. Rules are enforced deny-closed at four points, one of them a gate on MCP tool calls. Budgets can deny or throttle spend, and every operation lands in a hash-chained, Ed25519-signed ledger. → https://robert-haase.de/en/evidence-agents.md
- **adobe-markenpruefung** · Vendor documentation · Adobe checks campaign drafts against stored brand guidelines and shows the result as a percentage. It is the share of guidelines a draft passes out of the guidelines tested. Added to it are pass-or-fail results for channel guidelines such as Meta and LinkedIn and for ADA accessibility. The value is recalculated after every edit. → https://robert-haase.de/en/evidence-agents.md
- **markup-ai-stilpruefung** · Vendor documentation · Markup AI checks content against a brand’s own voice and style rules, and is itself reachable over MCP. The tool flags what does not fit, explains why and supplies wording to apply in place; every text is also scored against the stored standards. The vendor runs an MCP server at api.markup.ai that assistants such as Claude or Cursor connect to, for instance with a prompt asking it to check a paragraph for brand-voice drift. → https://robert-haase.de/en/evidence-agents.md
- **lighthouse-agent-discovery** · Vendor documentation · Since version 13.5.0 of 18 September 2026, Google’s audit tool Lighthouse also checks whether a website’s catalogue for agents conforms to the “Agentic Resource Discovery” specification, and groups this check with the llms.txt check under a group of its own, “Agent Discoverability”. According to the release, this ships in the DevTools of Chrome 156 and in PageSpeed Insights within two weeks. → https://robert-haase.de/en/evidence-agents.md
- **content-signal-selten** · Own survey, reproducible · Cloudflare’s machine-readable declaration “Content-Signal”, with which a robots.txt allows or refuses search, AI input and AI training, appears at none of the 77 German-language news and trade media. Among the 139 robots.txt files served by home pages of the DAX, MDAX and SDAX companies, exactly one carries it, that of Heidelberg Materials, and it allows all three uses. → https://robert-haase.de/en/evidence-agents.md

## Commerce (7)

+ Verified study, vendor documentation, or court decision (2) → airline-direktkanal, shopify-knowledge-base
+ Preliminary: prototype, single test, forecast, or vendor figure (0)
+ Status, case report, or market observation (5) → checkout-rueckbau, walmart-verhandlung, journey-start, kaufentscheidung, ucp-gremium-ohne-zahlen

- **checkout-rueckbau** · Market observation · Buying directly inside the chat was rolled back a good five months after launch. Live at that point were either a dozen or close to thirty Shopify merchants, depending on the source. → https://robert-haase.de/en/evidence-commerce.md
- **airline-direktkanal** · Controlled test · In a flight-search test, language models went to the airline’s own website directly in only about five percent of cases. They preferred booking portals. → https://robert-haase.de/en/evidence-commerce.md
- **walmart-verhandlung** · Vendor and company figures · Walmart has supplier negotiations run by an AI system: 2,000 negotiations at once, around three percent average savings, and three out of four suppliers preferring to negotiate with the machine over a person. → https://robert-haase.de/en/evidence-commerce.md
- **journey-start** · Survey · 30 to 45 percent of US consumers use generative AI for product research and comparison. 17 percent said they would begin their holiday shopping on an AI platform. → https://robert-haase.de/en/evidence-commerce.md
- **kaufentscheidung** · Survey · Only 11 percent of respondents would let an AI make the purchase decision, and only in low-stakes categories such as personal care and household supplies. → https://robert-haase.de/en/evidence-commerce.md
- **ucp-gremium-ohne-zahlen** · Standards status · The Shopping Tech Council of the UCP commerce protocol has 16 seats; since 24 April 2026 they include Amazon, Meta, Microsoft, Stripe and Salesforce alongside Google, Shopify, Etsy, Target and Wayfair. How many merchants actually run the protocol is stated by none of the companies involved. Google names example merchants — Nike, Sephora, Target, Ulta Beauty, Walmart, Wayfair, and Shopify merchants such as Fenty and Steve Madden — attached to the word “soon”. → https://robert-haase.de/en/evidence-commerce.md
- **shopify-knowledge-base** · Vendor documentation · Since 16 May 2025, Shopify has offered merchants a free app of its own for deciding what AI shopping agents answer about their store. Merchants see automatically generated facts and common customer questions and can adjust answers or write new ones. The answers do not appear in the store; they serve AI platforms as a data source, and the app shows how many questions come from agents and whether the AI can answer them. → https://robert-haase.de/en/evidence-commerce.md

## Liability (12)

+ Verified study, vendor documentation, or court decision (4) → air-canada, olg-hamm, perplexity-cfaa, screenshot-metadaten
+ Preliminary: prototype, single test, forecast, or vendor figure (1) → ai-overview-muenchen
+ Status, case report, or market observation (7) → cursor-bot, ai-act, produkthaftung-komplexitaet, auftragsverarbeitung-weisung, chevrolet-dollar, dpd-chatbot, nyc-mycity

- **air-canada** · Tribunal decision · Air Canada is liable for what its chatbot promised. The defence that the bot was responsible for its own actions was called "a remarkable submission" by the decision-maker. → https://robert-haase.de/en/evidence-liability.md
- **cursor-bot** · Documented incident · Cursor’s support agent invented a usage rule that never existed and replied under the name "Sam". → https://robert-haase.de/en/evidence-liability.md
- **ai-act** · Legal status · The EU AI Act transparency obligations for chatbots and for AI-generated content have applied since 2 August 2026. Generative systems placed on the market before that date have until 2 December 2026 to add machine-readable marking. The high-risk obligations were postponed to 2 December 2027 and 2 August 2028. → https://robert-haase.de/en/evidence-liability.md
- **olg-hamm** · Court judgment, final · A German higher regional court has ruled, with the judgment now final, that a company is directly liable for the misleading statements of its own AI chatbot. The chatbot is not a third party in the sense of the law. → https://robert-haase.de/en/evidence-liability.md
- **produkthaftung-komplexitaet** · Legal position · The new EU Product Liability Directive counts software explicitly among products in Article 4. Under Article 10 a court *shall presume defectiveness* where proving it is “excessively difficult” for the claimant “in particular due to technical or scientific complexity” and they show only that it is likely. The same presumption applies where the defendant fails to disclose evidence it has been ordered to produce. The directive must be transposed by 9 December 2026. → https://robert-haase.de/en/evidence-liability.md
- **ai-overview-muenchen** · Preliminary injunction, void after settlement · A German regional court granted an injunction barring Google from spreading eight claims about a publishing house and seven about a company belonging to it in its AI Overview, among them fraud scheme and subscription trap. It treated the AI Overviews as Google’s own attributable content rather than mere search results, held Google directly liable as the interferer (unmittelbare Störerin) and denied the liability exemptions for hosting providers and for search engines. After Google appealed, the parties settled. → https://robert-haase.de/en/evidence-liability.md
- **auftragsverarbeitung-weisung** · Legal status · For processing on behalf of a controller, the General Data Protection Regulation requires a contract or other legal act stipulating that the processor processes personal data only on documented instructions from the controller (Art. 28(3)(a)); Art. 29 binds directly, and the controller must be able to demonstrate compliance (Art. 24(1)). Art. 28(10) sets the tipping point: a processor that infringes the Regulation by determining the purposes and means of processing is considered a controller in respect of that processing. → https://robert-haase.de/en/evidence-liability.md
- **chevrolet-dollar** · Documented incident · The ChatGPT-powered chatbot of a Chevrolet dealer agreed to sell a new Tahoe for one dollar and called it a legally binding offer with no take-backs. The user had dictated that exact formula to the chatbot earlier in the same conversation. The chatbot went offline shortly afterwards. No source reports an attempt to enforce the offer. → https://robert-haase.de/en/evidence-liability.md
- **dpd-chatbot** · Documented incident · DPD’s chatbot called its own company “the worst delivery firm in the world” after a customer told it to recommend better delivery firms and to be over the top in its hatred. DPD said the AI element had been disabled immediately. → https://robert-haase.de/en/evidence-liability.md
- **nyc-mycity** · Documented incident · New York City’s official MyCity chatbot told businesses to do things that are illegal in the city: go cash-free, take a cut of employees’ tips, and turn away tenants with housing vouchers. Ten members of the newsroom asked the same question and all ten got the same wrong answer. The city defended it as a pilot program; almost two years later, in early February 2026, it was shut down as a budget cut. → https://robert-haase.de/en/evidence-liability.md
- **perplexity-cfaa** · Court ruling, not final · A US federal appeals court vacated the preliminary injunction against Perplexity on 4 August 2026 and remanded the case. When someone runs a shopping agent, it is the user who accesses the third-party website under the Computer Fraud and Abuse Act, the agent is the user’s tool and the provider does not access anything itself, as long as the agent runs in the user’s browser and the provider’s servers never call the site themselves. → https://robert-haase.de/en/evidence-liability.md
- **screenshot-metadaten** · Vendor documentation · A screenshot does not carry the original’s C2PA provenance data: the record lives in the file, and a screenshot creates a new one. Conversely, a C2PA-enabled camera photographing an AI image signs that shot, with no trace of its AI origin. As a rule it records device, time and place in metadata and cannot analyse the content of the image; what goes in is up to the implementer, the same page says. → https://robert-haase.de/en/evidence-liability.md

## Market size (16)

+ Verified study, vendor documentation, or court decision (2) → ki-nutzung-deutschland, ki-anteil-artikel
+ Preliminary: prototype, single test, forecast, or vendor figure (8) → machine-customers, marktgroesse, mcp-verbreitung, agentenhandel-2030, dark-data-55, ki-verkehrsanteil, markenklone, suchmarkt-wachstum
+ Status, case report, or market observation (6) → nicht-menschlicher-verkehr, cmo-ki-anteil, geo-verbreitung, in-house-verlagerung, insourcing-absicht, agentur-selbstbild

- **machine-customers** · Self-assessment, forecast · In a Gartner survey, chief executives estimate that by 2030, 15 to 20 percent of their revenue will come from machine customers. → https://robert-haase.de/en/evidence-market.md
- **marktgroesse** · Vendor figures · The large user numbers for AI systems come from the vendors themselves and are not comparable with each other: around 900 million weekly active users for ChatGPT, over one billion monthly users for Google AI Mode. → https://robert-haase.de/en/evidence-market.md
- **ki-nutzung-deutschland** · Official statistics · 26 percent of German companies with ten or more employees used AI technologies in 2025. Among large companies with 250 or more employees it is 57 percent, among small ones 23 percent. → https://robert-haase.de/en/evidence-market.md
- **nicht-menschlicher-verkehr** · Market observation · Cloudflare reports that in 2026, for the first time, more than half of Internet traffic is not human. Better quantified in the same report: 52 percent of crawler requests served AI model training in June 2026, up from 22 percent in spring 2025. → https://robert-haase.de/en/evidence-market.md
- **mcp-verbreitung** · Vendor figures · On handing the Model Context Protocol to the Linux Foundation on 9 December 2025, Anthropic gives more than 10,000 active public MCP servers and over 97 million monthly SDK downloads across Python and TypeScript. The platinum members of the new Agentic AI Foundation include, per the foundation, AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft and OpenAI. → https://robert-haase.de/en/evidence-market.md
- **agentenhandel-2030** · House forecasts, not comparable · Morgan Stanley puts agentic shoppers at 190 to 385 billion dollars of US e-commerce by 2030, a likely 10 percent market share and up to 20 percent in the optimistic case. Nine days later Bain puts agentic commerce at 300 to 500 billion dollars, roughly 15 to 25 percent. Only Bain states an inclusion rule and counts influenced purchases. → https://robert-haase.de/en/evidence-market.md
- **cmo-ki-anteil** · Survey · Marketing leaders at U.S. companies use AI or machine learning 24.2 percent of the time they spend optimizing and automating marketing. The typical company says 20 percent. Two surveys earlier the figures were 13.1 (September 2024) and 17.2 percent (early 2025). For generative AI alone the figure rose from 7.0 through 15.1 to 22.4 percent. Within three years the same respondents expect 55.9 percent. → https://robert-haase.de/en/evidence-market.md
- **dark-data-55** · Usually miscited · The most quoted figure on unused corporate data, 55 percent, bundles self-estimates the 1,357 respondents made about their own organisation. It was fielded in 2018/19 by the research arm of the PR agency FleishmanHillard on behalf of Splunk, a vendor selling software to analyse exactly this data. → https://robert-haase.de/en/evidence-market.md
- **geo-verbreitung** · Survey · 78 of 188 US companies that answered this question use AI for generative engine optimization, that is, to get their own content to appear in AI-generated search answers. That is 41.5 percent, with a 95 percent confidence interval of plus or minus 7.1 percentage points. → https://robert-haase.de/en/evidence-market.md
- **in-house-verlagerung** · Survey · In the member survey run by the US advertising association ANA, 82 percent of the members surveyed said in 2023 that they had an in-house agency, after 78 percent in 2018, 58 percent in 2013 and 42 percent in 2008. 65 percent said in 2023 that they had moved ongoing business from an external agency in-house in the preceding three years. In 2018 it was 70 percent, in 2013 only 56. → https://robert-haase.de/en/evidence-market.md
- **ki-anteil-artikel** · Verified study · Of the English-language articles newly published in the first quarter of 2026, 49.9 percent were primarily AI-generated. Since early 2025 the share has moved between 44.6 and 50.9 percent, with no upward trend. → https://robert-haase.de/en/evidence-market.md
- **ki-verkehrsanteil** · Three vendor measurements · Three analytics vendors put a number on the share of website visits that arrive from an AI assistant: Contentsquare 0.2 percent in the fourth quarter of 2025, Semrush 0.14 percent for the year 2025, Conductor 1.08 percent for May to September 2025. Three separately collected measurements, fractions of a percent up to a good one percent. → https://robert-haase.de/en/evidence-market.md
- **markenklone** · Vendor figures · Takedown provider Netcraft states that between March 2024 and March 2025 it acted against 1.3 million phishing sites imitating more than 16,000 organisations. → https://robert-haase.de/en/evidence-market.md
- **suchmarkt-wachstum** · Market observation on estimated data · Between the first quarter of 2023 and the fourth quarter of 2025, search engine visits and search-like AI sessions combined grew by 26 percent worldwide, from 82.0 to 103.2 billion per month. Google’s share falls from 89 to 71 percent, ChatGPT reaches 20 percent. → https://robert-haase.de/en/evidence-market.md
- **insourcing-absicht** · Survey · Asked "Do you plan to cover more marketing services in-house through AI?", 80.0 percent of 170 executives at German companies with budget and decision authority answer yes, 11.2 percent no, and 8.8 percent do not know. Across company sizes the intention is stable: 79 percent at companies with 100 to 999 employees, 81 percent at 1,000 and above. → https://robert-haase.de/en/evidence-market.md
- **agentur-selbstbild** · Survey · 96.2 percent of 78 executives from member agencies of the German agency association GWA rate their own agency's AI maturity as "advanced" (71.8 percent) or "expert" (24.4 percent), 3.8 percent as "beginner". The same respondents rate the average level of AI knowledge across the agency industry in Germany, Austria and Switzerland mostly at 3 on a scale of 1 to 5 (57.7 percent), 19.2 percent at 2 and 23.1 percent at 4; nobody picks the extremes 1 or 5. → https://robert-haase.de/en/evidence-market.md

## Judgement (20)

+ Verified study, vendor documentation, or court decision (8) → metr-selbsteinschaetzung, jagged-frontier, homogenisierung, sykophanz, cowan-standards, kompetenz-nivellierung, boussioux-neuheit-wert, unverwechselbare-markenelemente
+ Preliminary: prototype, single test, forecast, or vendor figure (0)
+ Status, case report, or market observation (12) → strategie-trendslop, markenspezifikation-wirkung, foresight-performance, prognose-mensch-maschine, prognose-assistenz, abbott-zustaendigkeit, esposito-kommunikation, drei-arbeitsweisen, aufwand-statt-koennen, mintzberg-muster, wahrgenommene-differenzierung, de-skilling

- **metr-selbsteinschaetzung** · Controlled trial · Experienced developers took 19 percent *longer* with AI tools — while believing they had been 20 percent faster. Beforehand they had expected a 24 percent speed-up. Between measured and perceived effect lie 43 percentage points, with the sign reversed. → https://robert-haase.de/en/evidence-judgement.md
- **jagged-frontier** · Controlled trial · In a preregistered experiment with 758 management consultants, AI users completed 12.2 percent more tasks, worked 25.1 percent faster and delivered more than 30 percent higher quality, as long as the task fell inside the model's capability. On a task placed just outside it, they were 19 percentage points more likely to be wrong than the group without AI. → https://robert-haase.de/en/evidence-judgement.md
- **homogenisierung** · Controlled experiment · 293 participants each wrote a short story, some of them with starting ideas from GPT-4, and 600 readers rated them. Stories written with an AI idea were judged more novel, up 5.4 percent with access to one idea and up 8.1 percent with access to up to five. At the same time they converged: a story’s similarity to the mean of the others in its group rose by 0.871 points on a scale from 0 to 100, which the authors report as 10.7 percent of the range the group without AI spanned. → https://robert-haase.de/en/evidence-judgement.md
- **strategie-trendslop** · Editorially reviewed · Seven language models were put to seven strategic trade-offs, each as an either-or. On six of the seven they picked the same side across all vendors, differentiation over cost leadership and augmentation over automation among them; only on exploration versus exploitation did they diverge. Two follow-up studies on ChatGPT-5, each with more than 15,000 runs, tested whether the bias can be prompted away. Barely: for differentiation and augmentation, better prompting lowered the share of biased responses by less than 2 percent, and additional company context shifted it by 11 percent on average, in both directions. The authors call this “strategy trendslop”, the most socially desirable answer of the internet average. → https://robert-haase.de/en/evidence-judgement.md
- **sykophanz** · Controlled experiment and vendor documentation · Agreement is rewarded in the training signal. Anthropic analysed 15,000 response pairs from its own feedback data: whether a response matches the user’s beliefs is consistently among the strongest predictors of which response humans prefer. Any single feature shifts that probability by at most about 6 percentage points. In April 2025 OpenAI rolled back a GPT-4o update because the model had become excessively agreeable, naming as its early assessment three changes acting together, among them an additional reward signal from users’ thumbs-up and thumbs-down feedback. → https://robert-haase.de/en/evidence-judgement.md
- **markenspezifikation-wirkung** · Negative finding of a documented search · Whether a brand as a specification produces more brand-compliant AI output than a brand book is settled by no publicly verifiable measurement; the search finds no public benchmark for it. Neighbouring fields have such benchmarks: guideline adherence in medicine since December 2024, rule adherence in support dialogues since early 2026. → https://robert-haase.de/en/evidence-judgement.md
- **cowan-standards** · Historical study · A century of household technology did not reduce time spent on housework. The appliances mainly replaced work done by men, children and servants; the time saved went into rising standards of cleanliness and care. The expectation that automation frees up time has a documented precedent in which precisely that failed to happen. → https://robert-haase.de/en/evidence-judgement.md
- **foresight-performance** · Longitudinal study · Corporations whose future preparedness was rated strong in 2008 reached 16 percent profitability by 2015, against 12 percent for the industry average — 33 percent more. On market capitalisation growth over the same seven years they stood at 75 percent against an average of 25. Firms with identified deficiencies came in 37 to 44 percent below average. → https://robert-haase.de/en/evidence-judgement.md
- **prognose-mensch-maschine** · Ongoing measurement · On the ForecastBench tournament leaderboard, the median of human superforecasters scores 68.8 and shares third place; two Google DeepMind entries allowed to use tools and extra context lead at 69.2 and 69.0, and a third ties. On the base leaderboard, without tools, the human median leads at 67.8 against the best model at 62.0. → https://robert-haase.de/en/evidence-judgement.md
- **prognose-assistenz** · Preregistered experiment · 991 participants answered six forecasting questions, some with access to a language model. The assistance improved accuracy by 24 to 28 percent against the control group. The comparison within the groups is the notable part: an assistant deliberately tuned to be overconfident and noisy also helped substantially. → https://robert-haase.de/en/evidence-judgement.md
- **abbott-zustaendigkeit** · Standard reference · According to Andrew Abbott's study, occupations compete not over performing their work but over its *definition*. Whoever determines what counts as a problem, and who is responsible for it, has already settled the competition. Abbott's finding on how this happens: jurisdictions are claimed when they fall vacant — not by filling an occupied one better. → https://robert-haase.de/en/evidence-judgement.md
- **esposito-kommunikation** · Scholarly book · The sociologist Elena Esposito considers the analogy between algorithms and human intelligence misleading and proposes a different term: artificial *communication*. In her words: if machines contribute to social intelligence, it will not be because they have learned to think like us, but because we have learned to communicate with them. → https://robert-haase.de/en/evidence-judgement.md
- **drei-arbeitsweisen** · Field study, working paper · A field study of 244 management consultants found three ways of working with generative AI — differing not in the tool but in *who steers the workflow*. Those who involve the AI throughout acquire new AI capability. Those who use it selectively for individual steps, keeping the problem definition themselves, deepen their existing domain expertise. Those who hand over the whole process build **neither**. → https://robert-haase.de/en/evidence-judgement.md
- **kompetenz-nivellierung** · Field experiment · Among 5,179 customer support agents at a large firm, issues resolved per hour rose 14 percent on average with an AI assistant. The average hides the point: **novice and low-skilled workers gained 34 percent, while the effect on experienced and highly skilled workers was “minimal”.** The authors suspect the model disseminates the practices of abler workers to newer ones. → https://robert-haase.de/en/evidence-judgement.md
- **aufwand-statt-koennen** · Preregistered experiment · In a preregistered experiment with 444 college-educated professionals, time spent on writing tasks fell by 0.8 standard deviations while quality rose by 0.4. Here too the gap between participants narrowed — weaker performers gained more. The authors' reading: the tool mostly substitutes for *effort* rather than complementing *skill*, shifting work away from rough drafting towards idea generation and editing. → https://robert-haase.de/en/evidence-judgement.md
- **boussioux-neuheit-wert** · Controlled test · In an ideas contest on the circular economy, 300 screened evaluators each rated 13 of 234 solutions, 3,900 ratings in total: 54 from people, 180 from GPT-4 with human-guided prompts. The human ones were judged more novel (the machine ones minus 0.140 on a scale of 1 to 5), the machine ones more strategically viable, more valuable environmentally and financially, and better overall (plus 0.088 to 0.160). At the top end the picture flips: AI solutions received the top novelty mark 7.9 percentage points less often, and their value advantage vanished there across all four dimensions. → https://robert-haase.de/en/evidence-judgement.md
- **mintzberg-muster** · Exploratory case studies, concept-forming · Henry Mintzberg defined strategy in 1978 as “a pattern in a stream of decisions”: a strategy has formed once a sequence of decisions shows consistency over time. That opens to research the strategies which came about despite intentions, or with no intention at all. He showed it on two long-run cases, Volkswagenwerk and the United States in Vietnam from 1950 to 1973. → https://robert-haase.de/en/evidence-judgement.md
- **wahrgenommene-differenzierung** · Survey · Across 17 product categories in Australia and the UK, an average of 11 percent of a brand’s current users consider it different and 10 percent consider it unique; 17 percent name at least one of the two. They buy the brand anyway. The authors recommend distinctiveness instead. → https://robert-haase.de/en/evidence-judgement.md
- **de-skilling** · Survey · In a global survey of 70 C-suite leaders and senior executives, half already observe a loss of skills inside their own organisation, and more than 60 percent consider it a material threat within three to five years. The five skills the same leaders rate as most critical for long-term performance are exactly the five they see as most at risk: judgment and decision making, problem understanding and framing, creative thinking, analysis and causal reasoning, solution generation and evaluation. → https://robert-haase.de/en/evidence-judgement.md
- **unverwechselbare-markenelemente** · Verified study · Researchers at the Ehrenberg-Bass Institute analysed 1,162 distinctive brand assets of 128 brands across 21 categories, four countries and nine years. Shape-based assets such as logos and packaging perform best: on average 40 percent of respondents link them to the brand, and 71 percent of the links go to that brand alone. Colours perform weakest, at 12 and 39 percent. → https://robert-haase.de/en/evidence-judgement.md

End of overview: 108 of 108 claims in 6 topics.
