Illustration: a contemporary investigator's evidence wall of printed social-media posts, network diagrams and a pinned world map, in the same dossier style as the rest of the archive.

Corpus poisoning

Illustration, not an archive object

Also called AI story optimization, Retrieval poisoning, SEO propaganda

Seed search indices, encyclopedias, or model-facing text so downstream systems repeat the frame as if it were independent research.

The tell

Question-formatted reports, brand-new 'institutes,' Wikipedia clones, citation-shaped pages with no staff.

Not this

Ordinary publishing that happens to be indexed.

History and effect

An extension of search-engine optimization into the age of retrieval-augmented chatbots; 2025-26 "AI Story Optimization" vendors document it as a formal service line.

It makes a planted frame reappear to later readers with the borrowed authority of independent research, because a model or search index repeats it as if it had checked.

Lineage: SEO; LLM retrieval/training contamination; 2025–26 'AI Story Optimization'

How it works

Corpus poisoning aims at an intermediary rather than a reader: a system people consult as if it had already done the checking, such as a search index, an encyclopedia, a model's training data, or the pages a chatbot retrieves to answer a question. Planted material that gets in reaches people later as a ranking, an encyclopedia sentence or a chatbot answer, without the signals (author, sponsor, date, context) that would have let them judge it at first contact.

Documented cases work at three layers. At the index layer, material is written to rank, often where little else exists: Michael Golebiewski and danah boyd (Data & Society, 2018, updated 2019) call these "data voids", rare or newly coined search terms whose results belong to whoever filled them first. Commercial reputation management works at this layer too. At the reference layer, the aim is an encyclopedia sentence that other sites and models then repeat, through undisclosed paid editing, edits from institutional networks, or control of a smaller-language edition. At the model layer, content is published in bulk for crawlers to collect for training, or for retrieval-augmented chatbots to find and cite. The The Hanover Institute: reports written for chatbots to cite case study separates retrieval capture (a chatbot cites the page at answer time, which is observable) from training contamination (the text changes the model itself, which is hard to show from outside).

Documented material tends to share features: volume, rewriting rather than original reporting, the form of an authority, shaping for machines (question-phrased titles, many language editions, an llms.txt file), and a disclosure, if any, left behind on the original site. Security research has shown that training datasets can be poisoned through the way they are assembled (Carlini et al., 2023) and that a small, roughly fixed number of documents can plant a narrow behaviour in models of very different sizes (Souly et al., 2025); these are proofs of vulnerability published with defences, not campaigns.

Why it works

People give a system trust they would not give an anonymous page. Automation bias, accepting an automated system's output without enough scrutiny, was described by Kathleen Mosier and Linda Skitka (1996) and by Raja Parasuraman and Victor Riley (1997). Bing Pan and colleagues (2007) found that students chose higher-ranked search results even when researchers had reversed the true order of relevance.

Source-monitoring research (Marcia Johnson, Shahin Hashtroudi and Stephen Lindsay, 1993) shows that people remember what they learned more readily than where they learned it, so a claim met through a chatbot is later recalled as simply known; that gap is also the mechanism of the Sleeper effect. Repetition across apparently separate pages feeds the illusory truth effect (Hasher, Goldstein and Toppino, 1977).

Examples across eras and sides

  • The Beria replacement pages (1954, USSR, state publisher). After Beria's arrest and execution, Great Soviet Encyclopedia subscribers were told to cut out his article and paste in supplied pages carrying an expanded entry on the Bering Sea (David King, The Commissar Vanishes, 1997).
  • Congressional office edits (2006, United States, House staff). The Lowell Sun reported that staff in Representative Marty Meehan's office had replaced his Wikipedia article with a staff-written biography omitting his broken term-limit pledge (Lowell Sun, January 2006).
  • Wiki-PR (2013, United States, commercial firm). Wikipedia blocked more than 250 accounts linked to a firm selling article writing to clients through undisclosed accounts; in 2014 the Wikimedia Foundation's terms of use began requiring paid editors to disclose (Wikimedia Foundation, 2013-2014).
  • The MH17 edit (2014, Russia, state broadcaster network). Shortly after the airliner was shot down, an edit from an IP address registered to the state broadcaster VGTRK changed a Russian-language Wikipedia entry to say Ukrainian soldiers had shot it down, as flagged by the automated account @RuGovEdits. See Russian Defence Ministry satellite image, MH17 briefing.
  • Croatian Wikipedia (2010s, Croatia, volunteer administrators). An assessment commissioned by the Wikimedia Foundation (2021) found that a small group of administrators had for years steered the Croatian edition toward a nationalist framing, including on the Ustaše regime and the Jasenovac camp (Wikimedia Foundation, Croatian Wikipedia Disinformation Assessment, 2021).
  • Eliminalia (to 2023, Spain and elsewhere, reputation firm). The Forbidden Stories "Story Killers" consortium reported that the firm flooded search results with positive or neutral content to push critical articles down for clients, and used false copyright complaints to have critical articles delisted (Forbidden Stories, 2023).
  • The Pravda network (2024-2025, Russia, pro-Kremlin network). France's Viginum documented in February 2024 at least 193 sites it named "Portal Kombat", republishing Russian state media and pro-Russian channels without original reporting. The American Sunlight Project (February 2025) pointed to the sites' very low human traffic, described crawlers as the likely audience and called this "LLM grooming"; NewsGuard (March 2025) reported that ten chatbots repeated the network's false narratives about a third of the time in its tests.
  • The Hanover Institute (2026, Israel and United States, state advertising agency via contractors). A site styled as a U.S. think tank, with no named staff, published 124 unsigned, question-titled reports in about a week, which Politico's tests found ChatGPT and Perplexity citing. FARA filings list Israel's Government Advertising Agency as foreign principal; the disclosure sits in the footer, not in the chatbot answers. See The Hanover Institute: reports written for chatbots to cite (Politico Influence, 14 August 2026; the Guardian, 26 August 2026).

Russian Defence Ministry satellite image shown at the MH17 briefing, with two launchers marked

Satellite image shown at the Russian Defence Ministry's briefing of 21 July 2014 on the shooting down of MH17, from the gallery entry Russian Defence Ministry satellite image, MH17 briefing. The Wikipedia edit in the list above concerned the same event. Russian Ministry of Defence.

Summary page of the VIGINUM Portal Kombat technical report, February 2024

Summary page of VIGINUM's technical report Portal Kombat: A structured and coordinated pro-Russian propaganda network (French state agency SGDSN, February 2024), which reports at least 193 sites and says they relay content rather than produce it. Lower blank half of the page cropped. Licence Ouverte (Etalab 2.0).

FRAME diagram of the Hanover Institute case: a site styled as a think tank, reports cited by chatbots, disclosure in the footer

The Hanover Institute case (2026) as drawn for The Hanover Institute: reports written for chatbots to cite. Drawn by FRAME from the sources listed on that page.

How to spot it

  • When a chatbot or search panel cites a source, are there named authors, an address, a legal entity and a funding statement?
  • Is the source new, and did it publish a large body of work in a short burst?
  • Are titles phrased as the questions people type, with machine-facing extras such as an llms.txt file or many near-identical language editions?
  • Does the source report anything first-hand, or only rewrite other outlets?
  • Is the topic a data void? Ask the same question in neutral wording and compare.
  • For an encyclopedia claim, who added it, when and from which account? For U.S.-facing material with a foreign sponsor, is there a FARA filing?

Where it ends: edge cases and legitimate persuasion

  • Ordinary publishing. An organisation that publishes its position under its own name, written so search engines and chatbots can find it, is communicating normally. It becomes corpus poisoning when origin is hidden, independent research is impersonated, or volume crowds out other sources.
  • Disclosed conflict-of-interest editing. Wikipedia lets connected editors propose changes openly on talk pages; the line is the undisclosed account, not the interest.
  • Neighbours. Passing a message through intermediaries until its origin is lost is Narrative laundering; fake voices on social platforms are Astroturfing and, at scale, Manufactured consensus or Flooding the zone; AI text passed off as human work is Synthetic media; a disclosure that never reaches the reader is Disclosure theater.
  • Accident is not poisoning. Systems repeat errors no one planted. The tag needs evidence of deliberate seeding: coordination, concealment, or content made for machines rather than readers.
  • Security research that discloses its method and proposes defences studies the vulnerability rather than using it.

Sources

  • American Sunlight Project. Report on the Pravda network and "LLM grooming". February 2025.
  • Carlini, Nicholas, et al. "Poisoning Web-Scale Training Datasets is Practical." IEEE Symposium on Security and Privacy, 2024 (preprint 2023).
  • Forbidden Stories. "Story Killers" investigation. 2023.
  • Golebiewski, Michael, and danah boyd. Data Voids: Where Missing Data Can Easily Be Exploited. Data & Society, 2018 (updated 2019).
  • Hasher, Lynn, David Goldstein, and Thomas Toppino. "Frequency and the Conference of Referential Validity." Journal of Verbal Learning and Verbal Behavior, 1977.
  • Johnson, Marcia K., Shahin Hashtroudi, and D. Stephen Lindsay. "Source Monitoring." Psychological Bulletin, 1993.
  • King, David. The Commissar Vanishes. Metropolitan Books, 1997.
  • Lowell Sun, reporting on edits to Representative Marty Meehan's Wikipedia article, January 2006.
  • Mosier, Kathleen L., and Linda J. Skitka. "Human Decision Makers and Automated Decision Aids: Made for Each Other?" In Automation and Human Performance. Lawrence Erlbaum, 1996.
  • NewsGuard. Audit of AI chatbots and the Pravda network. March 2025.
  • Pan, Bing, et al. "In Google We Trust: Users' Decisions on Rank, Position, and Relevance." Journal of Computer-Mediated Communication, 2007.
  • Parasuraman, Raja, and Victor Riley. "Humans and Automation: Use, Misuse, Disuse, Abuse." Human Factors, 1997.
  • Politico Influence, 14 August 2026; the Guardian, 26 August 2026 (Hanover Institute reporting).
  • Souly, Alexandra, et al. "Poisoning Attacks on LLMs Require a Near-Constant Number of Poison Samples." Anthropic, UK AI Security Institute and Alan Turing Institute, 2025.
  • Viginum (SGDSN). Portal Kombat: A Structured and Coordinated Pro-Russian Propaganda Network. February 2024.
  • Wikimedia Foundation. Croatian Wikipedia Disinformation Assessment. 2021.
  • Wikimedia Foundation. Cease-and-desist letter to Wiki-PR (2013) and Terms of Use amendment on paid contributions (2014).

The science

Research on the psychology this technique relies on, from the Learn library:

Case studies

Images

Image Source Licence
MH17 briefing image, 21 July 2014 Gallery entry Russian Defence Ministry satellite image, MH17 briefing See the gallery entry
VIGINUM, Portal Kombat technical report, summary page, February 2024 SGDSN Licence Ouverte (Etalab 2.0)

Image gaps

  • The replacement pages actually sent to Great Soviet Encyclopedia subscribers (Bering Sea, 1954): no scan with a verifiable free licence found on Wikimedia Commons. A Commons portrait of Beria tagged PD-Russia-1996 was left out because the photographer and date are unknown, so its US status could not be verified.
  • Screenshots of the Wikipedia edits (MH17, Meehan, Wiki-PR) and of Politico's chatbot tests: web pages or third-party works with no clear open licence, not hosted.
  • Newsguard and American Sunlight Project reports on the Pravda network: copyrighted, linked in Sources only.

In the watch briefs

All 9 briefs

In campaigns

Related techniques