Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor
David J. Grüning, Matthew Katsaros
We investigate the effectiveness of interventions that reduce the visibility of offensive content on the local social platform Nextdoor. Content filtering -- hiding or downranking offensive content that brushes against a platform's rules without clearly breaking them -- is deployed across virtually every major platform, yet almost no field evidence exists on whether it changes user behavior. We report two large-scale randomized controlled trials, each involving 100,000 users. Study 1 (2022) tested a report-triggered filter applied to comments in post threads and produced a modest 12% reduction in views of offensive comments; across eleven further measures of platform behavior we found no significant effects. Study 2 (2023-2024) remedied Study 1's central limitation -- a weak manipulation driven by slow, report-based eligibility -- by proactively scoring posts and comments at creation with Google Jigsaw's Perspective API and filtering them from the newsfeed. This produced a near-complete (95%) reduction in views of offensive posts, yet across thirteen further measures we again found no significant effects. Across two independent trials spanning different content types, filtering mechanisms, classifiers, and countries -- and despite manipulation strength rising from 12% to 95% -- filtering reliably reduced the visibility of offensive content without altering platform visitation, content consumption, or content production. These convergent null results provide rare field evidence on a ubiquitous intervention and underscore the complexity of effectively moderating online platforms.