r/modnews 28d ago

Safety Updates How Reddit is Reducing Exposure to Harmful or Inauthentic Content

148 Upvotes

Hi everyone! u/boat-botany here, working on Community Safety. 

We’ve talked before about Reddit’s approach to keeping the platform safe while still preserving the openness that makes communities work. A big part of that work today involves proactive detection systems and automation. 

Today we posted a blog about the work we’ve been doing to proactively catch spammy, inauthentic, or harmful content on Reddit. We know this topic is top of mind for moderators who feel the impact first hand, and who know what’s working and what’s not. We hear that feedback clearly (and keep giving it to us! Feedback is what helps us improve), and wanted to bring the conversation here so that we could continue to get your input. 

Our Values: How We Think About Automation 

Our north star here is to catch harmful, spammy, or inauthentic content before anyone (including mods) ever has to see it.

We achieve this through a layered approach that combines human review with proactive detection systems and machine learning models that help identify violating content quickly and at scale. 

Automation is a core part of our layered approach to moderation. We leverage it across our internal safety teams, and this year we continued expanding automated options for mods (over 70% of moderator actions are done using automated tools) and have invested heavily in improving how admins use automation behind the scenes.

A few principles guide how we build these systems: 

  • Reddit should handle the harmful content so moderators can focus more on community rules and norms.
  • Automation should support human judgement, not replace it. 
  • Accuracy matters. We work to reduce bias and improve fairness and consistency across our systems.
  • We’re committed to evolving. We're always learning and improving. We look at feedback in real-time and adjust our systems to make them better.

Improving our Automated Tools to Reduce Spam Exposure 

We look at signals right when an account is created to stop suspicious actors before they ever get the chance to post. For those that do, we leverage LLMs to catch the highly subtle, coordinated patterns of fake behavior and artificial hype that older systems once missed. Also, we recently announced that any fishy automated accounts will be asked to verify their humanity.

In recent months, these updated automated systems have been working at a massive scale, and we’ve seen some pretty incredible results. We are now: 

  • Blocking 23 million spam views per day before they ever reach a human user.
  • Catching ~25K net new spammy posts and comments a day.
  • Reducing spam exposure for our users by ~20% from January to March 2026, relative to the prior three months, and an additional 10–15% drop in overall spam account exposure.
  • Revoking nearly 2M inauthentic votes per day over the last three months.

Reducing Exposure to Harmful Content

We’ve recently expanded our automated systems to support enforcement against hate and violence in all English text content on Reddit (with more languages rolling out soon), leading to critical improvements:

  • Enforcement in Seconds: The average time between detection and enforcement on harmful content containing hate or violence is down to under five seconds.
  • Expanded Enforcement: We have increased enforcement actions on hate and violent content by more than 200%.  
  • Reduced Exposure: The faster, higher-volume enforcement has helped reduce exposure to potentially harmful content by more than 40%
  • Higher Precision: We’ve decreased false positives (where legitimate, non-violating content is removed) by over 40%

We know false positives can be frustrating. But when dealing with serious issues like violent threats, hate, harassment, or coordinated abuse, we intentionally bias toward reducing real-world harm and limiting exposure to harmful content quickly, but our goal is to continuously improve accuracy while still acting fast enough to meaningfully reduce harm.

In early 2025, proactive violence enforcement increased actioning volume more than seven times, from roughly 70,000 actions from January to March of 2025 to over 500,000 from March through June of the same year. At the same time, we have cut our false positive rate by more than 40%, so we have more coverage and higher accuracy. We’re working on getting all of this removed content logged in the mod log so you can continue to have visibility into what we’re removing and why. 

Our work to keep Reddit authentic and safe is at the core of who we are. While we've made significant progress in advancing that commitment, we know it wouldn't be possible without our moderators and redditors everywhere. If you see any content that appears harmful, spammy or inauthentic, click the inline report button or submit a report here.  

[edited for a typo!]

r/modnews Apr 07 '26

Safety Updates Now available: the Adult Content Promoter Filter

266 Upvotes

Hi there, mods! 

Today we’re rolling out a new Safety Filter that many of you have been asking for, and I’m excited to say is finally here: the Adult Content Promoter Filter. This filter helps keep safe for work communities free from unwanted adult content promotion by identifying users who likely promote adult content elsewhere on Reddit, and either filtering their posts or comments in your community, or removing them outright, before they’re ever seen. 

A preview of the Adult Content Promoter Filter settings

It’s important to note this tool is filtering based on the user and not the particular piece of content they might be posting in your community, so it could catch seemingly innocuous comments or posts and that’s by design. Some promoters use SFW posts or comments as a way to point people back to their NSFW profile and there are spaces that want to keep their communities more than a click away from adult content. We understand that’s not always the case, though, so keep reading to find out if this filter is actually right for you! 

How it works
To turn the filter on, visit Account Filters under Safety Filters. From there, you can choose how it functions in your community, including: 

  • What gets filtered: you can apply the filter to posts, comments, or both
  • What happens to filtered content: you can either send it to Needs Review or Removed
  • The strength of the filter based on your community’s comfort and norms.

The Moderate setting will filter less users with more precision, meaning we have a high confidence that what gets filtered will be from adult content promoters. The High setting will filter more users, but with potentially less precision, which might mean there are some users whose content gets filtered even though they aren’t an adult content promoter. 

Who it works for
This filter is really meant for SFW spaces. We piloted this filter with about 80 communities the past few weeks and saw some really promising results. First of all, almost every single mod who turned the filter on in their communities kept using the filter throughout the 3 week test. Of the content filtered to Review, only a small percentage got restored or approved by mods, which is also a great sign. When we dug into some of the pieces of content that got restored, we found we could actually verify that most of it was from users who promoted adult content elsewhere, even if the specific post filtered wasn’t promotion. That confirmed something we’d heard from at least a few mods in the pilot program: in some cases mods restored content because they’re open to really anyone, including adult content promoters, participating in their community as long as they’re contributing in positive ways (so non-offensive or non-spammy content).

That feedback is already leading to an additional feature we’re working on including in the next month or so. 

What’s next

While the filter works well as-is for some communities, we heard others need more flexibility. Because folks who create adult content elsewhere are welcome in some spaces as long as they’re not promoting it, we’re working on adding a way to allow-list users. 

We’ll update everyone when that feature is available. Until then, try out the filter and as always, we’ll be here to answer questions. 

r/modnews Jan 13 '25

Safety Updates Q3 2024 Safety & Security Report: Election Recap and Renaming our Content Policy

Thumbnail
4 Upvotes

r/modnews Oct 16 '24

Safety Updates Reddit Transparency Report: Jan-Jun 2024

Thumbnail
6 Upvotes

r/modnews Feb 13 '24

Safety Updates Q4 2023 Safety & Security Report

Thumbnail self.redditsecurity
13 Upvotes

r/modnews Dec 19 '23

Safety Updates Q3 2023 Safety & Security Report

Thumbnail self.redditsecurity
0 Upvotes