Definition
Content moderation is the set of policies and technical controls platforms use to detect, block, or remove prohibited user content and model outputs — including prompt filtering and output scanning for generative AI Spaces/apps.
Key Points
-
Policy bans without platform-enforced filters create enforcement gaps (developer-opt-in only)
-
Aggregate statistics and policy recommendations are the journalistic framing for sensitive misuse research
-
Distinct from cybersecurity incidents (intrusion) though both affect platform trust
-
2026-07-29: pangram funding/product push for ai-content-detection amid ai-slop (2026-07-29-pangram-9m-ai-detection)