Full report + evidence appendix
Take the citation audit with you
The page is fully public. Enter your work email only to download the formatted PDF with the methods table, claim-to-evidence registry, coding definitions, and archive references.
Executive Summary
73.0%–75.5% — of the citations behind brand-recommendation answers in our frozen-question baseline were confirmed marketing-ecosystem sources, after classifying every citation and running a four-part provenance review on every candidate news/official domain. (C1 baseline, T0 recommendation-intent questions: 163 citations, 119 confirmed marketing-ecosystem, 4 unreachable — reported as a sensitivity range. API layer, July 2026.) Re-collections at T+6h and T+24h detected no timepoint difference at either sensitivity endpoint (valid n=16 cell clusters): the pattern was stable within a 24-hour window.
$40 — the publicly listed starting price to place an advertorial on a content channel of a type that ends up cited by the AI. Live within 24 hours, ≥85% claimed search-inclusion rate — per the reseller's own rate card. (≈¥300; archived self-published pricing page, July 2026.)
0 — official or complaint-platform sources among all 27 citations, when we asked about a brand the way a normal user would: "is this brand any good?" The AI's answer said complaints were rare. Its citation list contained zero complaint-platform sources — and did contain the brand's own marketing. (Follow-up-chain probe, 4 chains / 5 questions, July 2026.)
Where our own numbers failed verification, we say so in the appendix — including a preregistered expansion that did not replicate one of our own pilot contrasts.
Scale line: based on a July 2026 frozen-question baseline (144 T0 API runs across 8 industries, plus 32 time-expanded recommendation-intent runs; 641 citations individually classified across the full collection), a June 2026 study of 200 structured questions across 9 industries with 18,841 source-type occurrences, and archived evidence packets.
1 · Why this matters outside China
Doubao was the largest AI-native app in QuestMobile's June 2026 ranking, at 382 million monthly active users (Qianwen 167M, DeepSeek 129M; QuestMobile 2026 half-year report, published 2026-07-14, accessed 2026-07-29). For a brand selling in China, its answers can become an early — and unmanaged — brand touchpoint. And the way it selects sources is not a Chinese curiosity: wherever answers are retrieval-fed and placement is purchasable, the same pressures apply. The affiliate-listicle share already visible in Western AI citations is a smaller-scale version of the same supply problem. China runs the experiment at industrial scale, with a public price list — which makes it the clearest place to measure the mechanics.
We are a Beijing-based measurement firm. We test on real consumer surfaces where the layer matters, we publish our methods, and we open what can be opened (some question banks stay sealed to prevent gaming — stated explicitly below). This report is the aggregate picture.
2 · The experiments
Two measurement rounds feed this report; they are labeled separately throughout, and neither is silently pooled with the other.
Round A (June 2026, exploratory — not preregistered): 200 structured questions across 9 industries and 96 scenarios, five intent types (verification 59 / recommendation 52 / risk-avoidance 45 / comparison 26 / weight-replay 18), 636 archived run records. Search triggered in 98.6% of runs. From the assistant's cited material we logged 18,841 source-type occurrences across five source classes (user feedback 6,046; official/regulatory 5,831; news media 3,636; corporate self-media 2,021; industry vertical 1,307).
Round B (July 2026, the gated baseline behind the headline): 8 industries × 6 frozen questions × 3 repeats = 144 API runs at T0; recommendation-intent questions re-run at T+6h and T+24h (32 additional runs). Across the full collection, 641 citations (T0: 504; T+6h: 58; T+24h: 79) were individually classified against a written coding book; 15.08% were double-coded (76/504, agreement 70/76, κ = 0.879). All 62 candidate news/official domains received a four-part provenance review (ownership, content type, resale evidence, commercial signing). No proportion was computed until every citation reached a terminal classification; 20 citations across the full collection (4 within T0 recommendation-intent) remained unreachable and are carried as a sensitivity interval, not silently dropped.
Model epoch for Round B is pinned (doubao-seed-2-1-pro-260628, responses API, web search enabled). Earlier rounds are labeled by date and layer in the method appendix.
Layer discipline: both rounds measure the API layer; consumer-client spot checks are run for interoperability, and our technical notes cover consumer-surface behavior separately — including a 16-pair web-vs-app study in which displayed search terms differed in 15/16 pairs and displayed source sets in 14/16 (TN-01; the mobile account label in that study is unverified, so we report the divergence as observed across surfaces without attributing cause). "Displayed" means what the interface shows; we make no claim about backend retrieval. API results are directional for the consumer experience; we never conflate the layers, and neither should you.
3 · How the AI talks about deciding — and what it actually cites
Two kinds of evidence, kept strictly apart:
What it says. The assistant's own stated reasoning is risk-heavy. In an earlier deep-probe round (V1 dataset: 6 prototypes, 72 probes, 5,215 signals), it invoked its "one-strike veto" concept 170 times in its own explanations — a frequency of the concept in stated reasoning, not 170 executed vetoes. In a regulated-industry case (car insurance qualification checks), it explicitly ranked regulatory official sources as the only admissible basis, above commercial credit and complaint platforms — its stated hierarchy, preserved verbatim in our archive. We treat all such self-description as hypotheses about behavior, never as facts about internals.
What it cites. Round B's classification shows sourcing is strongly intent-dependent: for recommendation-intent questions, confirmed marketing-ecosystem citations run at 73.0%–75.5% (T0); across the full collection regardless of intent — T0 all intents plus the recommendation-intent time-expansion arms — the share is 51.3%–54.5% (329 confirmed / 20 unreachable / 641 total). Trust-shaped questions pull the most marketing-shaped evidence. The two denominators are different scopes and are never mixed.
4 · The pollution economy
The price list. This ecosystem advertises. A placement reseller (Yunmeitong, operated by Xiamen Yunzhu Technology) publicly lists advertorial placement from ¥300 (~$40) per article, live within 24 hours, with a claimed ≥85% search-inclusion rate. We archived the rate card. We are not alleging any specific placed article was paid for; we are reporting that the channel type the AI cites is openly for sale, with a menu.
Anatomy of a trusted-looking source. One commercial site cited in "which brand is reliable" answers carries a name that resembles an e-government portal (its ICP filing is corporate). It hosted a merchant's self-ranked "top brand" advertorial. The question format that most signals caution — "哪家靠谱 / which one is trustworthy" — is precisely the format this content targets, because the phrasing is the SEO target string.
The listicle entrance. In a small June–July probe (9 runs, 43 citations — small-sample, reported as such), two commercial ranking sites (maigoo.com, chinapp.com) together accounted for 11 of 43 citations: the recommendation entrance is substantially intermediated by ranking-site formats.
The follow-up chain: verification theater. We then asked like a real user: category question → "is [named brand] any good?" → price → contact. Across 4 chains and 27 citations: zero official sources, zero complaint-platform sources. One answer asserted "complaints are rare" while citing no complaint platform at all — and while citing the brand's own website. A "4.95/5 rating" surfaced in answers traces to the brand's own advertorial network (2 of 6 sources same-network; the rating is network-published marketing content, not an independently verified score). We call this the self-citation trap: when the AI "verifies" a brand, the evidence it cites can be the brand's own marketing wearing a third-party costume.
5 · The contrast case — and a hypothesis, stated as one
Regulated industries are not automatically safe. In our pharmacy/insurance contrast (6 questions, 23 citations — small-sample), drug-purchase recommendations still drew 5 listicle citations out of 8. But a major insurer's brand-verification answers pulled up complaint-platform records and claims-dispute reporting — the record layer surfaced on its own. That insurer also maintains its own published answer pages ("is X insurance reliable / pitfalls explained") that the assistant cited: answer-shaping done with real records, in the open. (Named positively: Ping An. We make no claim about its products.)
These observations are consistent with a record-layer supply hypothesis: where a brand has a retrievable, genuine record layer (official filings, complaint history, independent coverage), verification-type questions can retrieve records; where that shelf is thin, marketing content is more available to fill the gap. This is a hypothesis about supply, not an established causal rule. We did run our preregistered intervention study (C2: 8 frozen hypotheses, single question-side conditions added/removed, bidirectional arms, July 2026) — and under its preregistered decision rules, no hypothesis earned a causal license: all seven evaluable hypotheses returned "mixed" (three reached the target-direction endpoint in 4/4 pairs but reverse-direction exclusions occurred as well, so none is reported as causal; one hypothesis was not evaluable at 3/4 valid pairs). We publish that outcome as-is.
The reversal test surfaced something arguably more useful: between paired runs differing by a single question condition, the candidate set changed in 31/31 pairs, with mean overlap (Jaccard) of 0.026 — and 0/31 pairs shared even two common candidates, making ranking comparison impossible. Within this API window, a single answer's candidate list is close to a fresh draw. For brands, the practical unit of AI visibility is aggregate presence across a question space — not any one answer.
6 · What holds over time — and across surfaces
- Round B re-collections: recommendation-intent marketing share was 73.0%–75.5% at T0, 77.6%–79.3% at T+6h, and 70.9%–72.2% at T+24h; with valid n = 16 cell clusters, no timepoint difference was detected at either sensitivity endpoint. We describe the pattern as stable within a 24-hour window — and claim nothing beyond that window.
- The two consumer surfaces diverge in what they display: 15/16 pairs differed in displayed search terms, 14/16 in displayed source sets (TN-01; account-label caveat as above).
- In a separate 10-question × 2-surface brand-comparison test (one run per cell), all 20 first answers introduced a third competitor never named in the question; in 9 of 10 web/app pairs, the added competitor's identity differed between surfaces (video walkthrough with annotated evidence on our channel).
7 · What we would expect elsewhere — if the hypothesis holds
Nothing here demonstrates that Western answer engines behave the same way. The transferable claim is conditional: wherever (a) answers are retrieval-fed, (b) placement is purchasable at scale, and (c) genuine record layers are thin, the supply-side pressure we measured has room to operate. Western affiliate/listicle citation patterns are consistent with the early form of that pressure. Whether the mechanics transfer is a measurement question — the same frozen-question, per-citation method applies to any engine, and that is exactly how we intend to test it.
Method appendix
- Boundary statements. No global-proportion claims extend beyond stated scopes; the assistant's stated weights and reasons are its own claims, not measured internals; operational reference ranges are not platform parameters; findings are time-boxed; independent re-verification is invited.
- Negative and mixed results disclosed. Our preregistered causal study returned all-mixed verdicts (§5) and is reported without upgrade. Separately, a pilot contrast suggested targeted high-risk claims lacked located source support more often than random claims. The preregistered expansion (66 + 54 assessable claims) did not replicate the contrast (Fisher p = 0.359; difference CI −8.1 to +27.7pp). The comparison is retired. Absolute absence rates are reported only under the assessable-claim scope, with repaired inter-coder agreement 57/60 (95.0%); absence of located support does not mean a claim is false, and we do not present absence as causing ranking or exclusion.
- Self-test disclosure. Asked "who is [our own then-current product]?" five times on the web surface, the assistant recognized it zero times (one confusion with a same-named older product). We measure a game we are also losing.
- Layer statement. API ≠ consumer client. In consumer-client spot checks, one merchant's claims that the API layer repeated were framed by the consumer surface as "unverified marketing" — the consumer surface can be more defensive. Our managed audits therefore run on the real logged-in consumer client.
- Independence. We sell GEO measurement and consulting services. No measured entity paid for inclusion. Entities that turned out to be clients were removed from public samples. Measurement itself cannot be purchased.
- Data availability, stated precisely. The TN-01 dataset is public (CC BY 4.0). The Round-B index question bank remains permanently sealed (anti-gaming); its aggregate tables, per-citation classifications, and the coding book are released with this report; key cited pages carry external archives; full evidence packets are available on request. Model epochs per dataset are listed in the method table.
Wang Bo — GeoCheckTool (Superlang, Beijing). The China AI measurement firm.
Open release package
Inspect the method behind the headline
The question bank remains sealed to prevent gaming. The coding rules, dataset boundaries, claim registry, evidence hashes, and public archive references are open.
Managed China measurement
Want this run on your brand?
Freeze 20 buyer questions, measure the real logged-in Doubao consumer surface, and receive per-question evidence with explicit limitations.
China Evidence Snapshot — $390