Illustration: a pre-1914 imperial intelligence-dossier collage of colonial telegraph logs, sepia field photography and empire maps.

Platform interventions:labels, nudges, notes, friction and takedowns

Illustration, not an archive object

In one line: social media companies have tried labelling false or state-backed posts, nudging users to think about accuracy, crowd-sourced notes, speed bumps on sharing, banning accounts and removing covert networks; most of these cut sharing or belief somewhat in the studies that exist, several have side effects or null results, and the companies control most of the data needed to check.

Evidence: mixed. Field studies of X's Community Notes and of Twitter's January 2021 bans show reduced spread, and accuracy prompts replicate with small effects; state-media labels show null results when unnoticed, labels can make unlabelled falsehoods look truer, and many headline figures are the platforms' own.

Background

  • December 2016: disputed flags. Facebook began labelling stories that outside fact-checkers had found false. That programme is covered in Fact-checking and corrections; this page covers the other tools.
  • 2017 onward: takedown reports. Meta says its public threat reporting began in 2017 with a Russian operation linked to the Internet Research Agency. It calls the target "coordinated inauthentic behavior" (CIB): coordinated efforts to manipulate public debate in which fake accounts are central.
  • June 2020: state-media labels on Facebook. Facebook began labelling pages, posts and ads of outlets "partially or wholly under the editorial control of a state," starting with about 200 pages including RT, Sputnik, CCTV and Xinhua, and said it would block their ads in the US.
  • June to December 2020: Twitter friction. Twitter tested a prompt asking users to open an article before retweeting it (June 2020). From 9 October to 16 December 2020 it sent users who pressed retweet to the quote-tweet composer first.
  • 6 August 2020: Twitter state-media labels. Twitter labelled "state-affiliated media" (outlets where the state controls editorial content "through financial resources, direct or indirect political pressures, and/or control over production and distribution") and key government officials from the US, China, France, Russia and the UK. It stopped recommending state-affiliated media accounts. It said publicly funded outlets with editorial independence, such as the BBC and NPR, would not be labelled.
  • January 2021: Birdwatch and the post-6 January bans. Twitter launched Birdwatch, a crowd-sourced notes system, on 25 January 2021. It had permanently suspended Donald Trump on 8 January, and said on 11 January that it had removed more than 70,000 QAnon-linked accounts since then.
  • November 2022: Community Notes. Birdwatch was renamed Community Notes and expanded. A note is shown only when raters who usually disagree both rate it helpful (a "bridging" algorithm), not by majority vote.
  • April 2023: labels dropped on X. X labelled NPR "US state-affiliated media," changed it to "government-funded media" (also applied to the BBC), then removed all such labels, including those on RT, Sputnik and Xinhua. NPR stopped posting to the platform.

Dated sequence of ten platform actions from December 2016 to April 2023: labels, friction, crowd notes and removals

The launches listed in Background, in date order. Drawn by FRAME from this page's sources (company announcements); events are spaced evenly, not to scale. CC BY 4.0 (FRAME).

What was tried

Warning labels. A tag ("Disputed", "Rated false", "state-affiliated media") on a post or account. Twitter said it labelled about 300,000 tweets, 0.2 percent of US election tweets, between 27 October and 11 November 2020, and hid 456 of them behind a warning.

Accuracy nudges. Asking users to rate the accuracy of one unrelated headline, on the theory that people share falsehoods because their attention is elsewhere, not because they prefer them.

Crowd notes. Volunteer-written context under a post, shown only after cross-perspective agreement.

Two-row schematic: a note rated helpful by one group of raters only is not shown; a note rated helpful by both that group and a group that usually disagrees with it is shown

The rule described above for crowd notes, in simplified form. The real system predicts each rating with matrix factorization (twitter/communitynotes, "Ranking notes"); the counts here are illustrative. Drawn by FRAME. CC BY 4.0 (FRAME).

Friction. Extra steps before sharing: the read-first prompt and the quote-tweet default.

Deplatforming. Permanent removal of accounts.

Network takedowns. Removal of coordinated fake-persona networks, then public reports naming the origin. This targets Astroturfing (fake grassroots), not the truth of what is posted.

Measured effects

Study Design N Finding Caveats
Clayton et al. (2020), Political Behavior Pre-registered survey experiment on warnings and tags See paper A general warning, "Disputed" or "Rated false" tags each lowered perceived accuracy of false headlines; "Rated false" worked better than "Disputed"; effects similar whether or not the headline suited the reader's politics Modest effects; survey setting
Pennycook, Bear, Collins and Rand (2020), Management Science Two online experiments, Facebook-style headlines 6,739 Implied truth effect: false headlines left without a warning were rated more accurate when other headlines carried warnings Mechanical Turk sample
Pennycook et al. (2021), Nature Four survey experiments plus a Twitter field experiment 5,379 Twitter users A private message asking users to rate one non-political headline raised the quality of news sites they retweeted in the next 24 hours Users were those who had shared links to low-quality sites
Roozenbeek, Freeman and van der Linden (2021), Psychological Science Pre-registered direct replication, COVID-19 headlines 701, then pooled 1,583 First stage failed (p = .67); pooled data gave a small significant effect on discernment Effect smaller than original
Pennycook and Rand (2022), Nature Communications Meta-analysis of the authors' 20 experiments, 2017 to 2020 26,863 Accuracy prompts cut sharing intentions for false headlines by about 10 percent relative to control Authors' own studies; sharing intentions, not behaviour
Nassetta and Gross (2020), HKS Misinformation Review Experiment with RT YouTube videos and state-media labels See paper Labels reduced belief in RT's election misinformation only when viewers noticed them; labels over the video worked better than below it Single outlet
Betzer et al. (2025), HKS Misinformation Review Pre-registered US experiment, May 2022 2,555 Twitter-style state-media tags had no effect on perceived accuracy of false state-media claims, apparently because few noticed them; fact-check labels did reduce belief Mechanical Turk sample
China Media Project (2021) 33 Chinese official accounts, 50 days before and after Twitter's labels 33 accounts Most accounts got significantly fewer shares and likes; CGTN, Xinhua and People's Daily fell more than 20 percent per tweet Labels coincided with de-amplification, so effects cannot be separated
Bradshaw, Elswah and Perini (2024), American Behavioral Scientist 8,071 YouTube comments before and after labels on AJE, CGTN, RT, TRT World and VOA 5 channels Likes unchanged except RT, which fell; critical comments became less likely after labelling Comments are a proxy
Chuai, Tian, Pröllochs and Lenzini (2024), PACM HCI Difference-in-differences and regression discontinuity on the Community Notes roll-out All notes and source tweets Roll-out did not measurably cut overall engagement with misinformation; notes arrive too late for the early, most viral stage Platform-wide average
Chuai et al. (2026), Nature Communications Difference-in-differences on noted posts 237,180 cascades, 431 million reposts Once a note was shown, further spread fell 61.2 percent on average and deletion odds rose 94.3 percent; weaker for influential accounts and political posts Timing still a limit
Twitter (2020), company figures Read-first prompt test Not published Users opened articles 40 percent more often after the prompt; opening before retweeting rose 33 percent Company data, not peer reviewed
Twitter (2020), company figures Quote-tweet default Not published Quote tweets rose, retweets fell, total sharing fell about 20 percent Company data; reversed after the election
McCabe, Ferrari, Green, Lazer and Esterling (2024), Nature Panel of more than 500,000 active users; natural experiment on the 70,000-account ban 500,000+ Misinformation sharing fell among the banned users' followers as well as by the banned users Authors could not fix the size of the effect because the ban coincided with 6 January events

Null results, decay and critiques

  • Labels have side effects. The implied truth effect means partial labelling can raise trust in the unlabelled falsehoods. State-media tags did nothing when users did not notice them (Betzer et al.), and helped only when noticed (Nassetta and Gross).
  • Accuracy nudges are small. The first replication stage failed; the pooled replication and the authors' meta-analysis find a real but modest effect, measured mostly as stated intentions to share.
  • Notes are slow and selective. The Center for Countering Digital Hate (October 2024) found that 209 of 283 misleading election posts it sampled (74 percent) had accurate notes that were not shown to all users; those posts had 2.2 billion views. The bridging rule, which requires agreement across sides, is least likely to succeed on the most divisive topics.
  • Deplatforming moves people. Ali, Saeed and colleagues (2021) found that users who moved to Gab after a Twitter suspension became more active and more toxic, but reached a much smaller audience. A PNAS Nexus study found that deplatforming did not decrease Parler users' activity on fringe platforms. The analytics firm Zignal Labs reported a 73 percent drop in election-fraud mentions (2.5 million to 688,000) the week after the January 2021 bans; this is a commercial count, not a causal study.
  • Covert networks often reach few people. The takedown reports show what was removed, not what it achieved. In "Unheard Voice," the Stanford Internet Observatory and Graphika found that most posts in the pro-Western network got no more than a handful of likes or retweets and only 19 percent of the accounts had more than 1,000 followers.
  • What was not measured. Most label and nudge results are survey experiments. Company figures on friction and labels are not independently checked. No study measures whether any of these tools changed votes or long-term beliefs.

Entrance sign and driveway at Facebook's headquarters, 1 Hacker Way, Menlo Park, California

Entrance to Facebook's headquarters at 1 Hacker Way, Menlo Park, California, 2 March 2014. Facebook's fact-check labels, state-media labels and takedown reports are described above. Photo: LPS.1, CC0 1.0 (public domain dedication), via Wikimedia Commons.

Twitter's headquarters at Market Square, 1355 Market Street, San Francisco, with the company's sign

Twitter's headquarters at Market Square, 1355 Market Street, San Francisco, 28 August 2016. Twitter ran the labels, prompts and crowd notes described above. Photo: Tobias Kleinlercher, CC BY-SA 3.0, via Wikimedia Commons.

Who uses it and who criticises it

  • Platforms' stated position. Twitter presented state-media labels as context for users deciding what to trust. Facebook said state-controlled publishers "combine the influence of a media organization with the strategic backing of a state." Twitter said in November 2020 that the quote-tweet change slowed misleading information by reducing sharing overall.
  • Takedowns cover many countries, the US included. Meta says networks it disrupted originated in 81 countries and targeted at least 135, with Russia the source it has reported most often. Its reports also name domestic and allied actors: Rally Forge, a US marketing firm working for Turning Point Action (October 2020; 202 Facebook accounts, 54 pages and 76 Instagram accounts); STOIC, a Tel Aviv political marketing firm whose network targeted the US and Canada (May 2024; 510 Facebook accounts, 11 pages and 32 Instagram accounts); and in December 2018 five accounts, including that of New Knowledge chief executive Jonathon Morgan, over tactics in Alabama's 2017 Senate race, run by a project called "Project Birmingham" in support of the Democratic candidate, Doug Jones, who said he had not known about it.
  • The US military case. In July and August 2022 Twitter and Meta removed overlapping networks that "Unheard Voice" (SIO and Graphika, August 2022) described as the most extensive covert pro-Western operation open-source researchers had reviewed: accounts on Twitter, Facebook, Instagram and five other platforms promoting US and allied interests in the Middle East and Central Asia and opposing Russia, China and Iran. In November 2022 Meta said it found "links to individuals associated with the U.S. military." The Washington Post reported in September 2022 that Colin Kahl, under secretary of defense for policy, ordered a full accounting of online psychological operations.
  • Censorship concerns. On the Trump ban, ACLU senior legislative counsel Kate Ruane said "it should concern everyone when companies like Facebook and Twitter wield the unchecked power to remove people from platforms that have become indispensable for the speech of billions." Angela Merkel's spokesperson Steffen Seibert said she considered the permanent block "problematic," since freedom of opinion should be limited only by law. In Murthy v. Missouri two states and five individuals argued that federal officials had pressed platforms to remove speech; on 26 June 2024 the Supreme Court ruled 6 to 3 that they lacked standing, without deciding the merits. Justice Samuel Alito, dissenting with Justices Thomas and Gorsuch, called the pressure "no less coercive" for being subtle.
  • Targeted governments. China's CGTN ran the headline "Twitter labels China and Russia 'state-affiliated media' accounts but not BBC, NPR & VOA." NPR objected when X applied a similar label to it in April 2023, and stopped posting there, saying the label undermined its credibility and independence.
  • "Not enough" critics. The CCDH argues that Community Notes fail on the divisive, high-reach claims where they matter most; in its sample, posts that did show a note got 13 times more views than the note. The Stanford researchers behind "Unheard Voice" and the Pentagon review cited above show that Western-aligned operations used the same methods as Russian and Iranian ones.
  • Opacity. On 28 October 2025 the European Commission preliminarily found that Meta and TikTok had put "burdensome procedures and tools" in the way of researchers seeking public data under the Digital Services Act. Meta has shut its CrowdTangle monitoring tool. Outsiders can rarely check a takedown attribution or a platform's effect figure.

What the evidence supports and doesn't

  • Supported: visible labels and shown notes reduce belief in, or sharing of, the specific post they sit on.
  • Supported: accuracy prompts have a small, replicable effect on stated sharing intentions.
  • Supported: removing large numbers of accounts reduces what circulates on that platform, while some users move to smaller platforms and become more extreme there.
  • Not shown: that any tool changes population-level beliefs or elections, or that companies' own effect figures hold up independently.
  • Unresolved: who should decide what is labelled or removed, and whether the tools are applied evenly across countries and parties. Takedown reports show both foreign and US-linked networks being removed, but outsiders cannot audit what was not.

See also

See also in Learn

Images

Image Source Licence
Platform actions, December 2016 to April 2023 Drawn by FRAME from this page's Background section (make_counter_platform_interventions_diagrams.py) CC BY 4.0 (FRAME)
Community Notes display rule (schematic) Drawn by FRAME from this page and the X Community Notes documentation (same script) CC BY 4.0 (FRAME)
Facebook headquarters, 1 Hacker Way, 2014 Wikimedia Commons, LPS.1 CC0 1.0
Twitter headquarters, San Francisco, 2016 Wikimedia Commons, Tobias Kleinlercher CC BY-SA 3.0

Image gaps

  • Platform labels, Community Notes and prompts: screenshots of X, Facebook and YouTube interfaces are copyrighted designs and none were found on Wikimedia Commons under a free licence.
  • Takedown report figures: Meta, Graphika and Stanford report graphics are copyrighted.

Sources

Techniques

Related