BlogsRedditPlatform Safety and Authenticity

Platform Safety and Authenticity

Platform Safety and Authenticity

2
posts
2026

This feature thread tracks Reddit's ongoing efforts to maintain the authenticity and safety of its platform, particularly in the face of evolving threats like AI-generated spam, bots, and harmful content. It began with the deployment of advanced AI tools to enhance automated defenses, significantly reducing user exposure to spam and inauthentic content, and drastically decreasing enforcement time for hateful or violent material. This post details significant upgrades to these AI tools, leading to the introduction of Rules Hub, which uses LLMs to evaluate rule intent rather than just keyword matching, and enhancements to Post & Comment Guidance and Safety Filters, aiming to eventually replace Automod's enforcement workflows. The thread also covers efforts to protect the platform from data scraping and spamming by modernizing legacy systems like Old Reddit and migrating critical bots and moderator workflows to the modern stack, alongside continued support for Reddit's Developer Platform with expanded capabilities and a migration program for third-party apps.

2026

Modernizing Reddit’s Infrastructure and Moderation Tools

8/5/2026

This post details the modernization of Reddit's infrastructure and moderation tools. Key technical contributions include the introduction of Rules Hub, which leverages LLMs for more nuanced rule enforcement compared to Automod's keyword matching. It also mentions enhancements to Post & Comment Guidance and Safety Filters. Furthermore, it outlines plans to protect against data scraping and spamming by making changes to Old Reddit and migrating bots/workflows to the modern stack. The post also discusses the ongoing development of Reddit's Developer Platform, including new plugins and APIs, and the $1 Million App Migration Program to support third-party app transitions.

How We’re Keeping Reddit Real and Safe in the AI Era

7/6/2026

This post details significant upgrades to Reddit's AI-powered automated defense systems. New capabilities include using LLMs to detect subtle, coordinated patterns of fake behavior and artificial hype, and signals at account creation to stop suspicious actors early. The post quantifies improvements such as blocking 23 million spam views daily, catching 25K spammy posts/comments daily, reducing user spam exposure by ~20%, and revoking nearly 2 million inauthentic votes daily. For harmful content, enforcement time is now under five seconds, with a >200% increase in enforcement actions and a >40% reduction in exposure and false positives.