Every plugin listing in the official Em Dash plugin registry is now screened by Cloudflare's Clef decision model before it can appear in the catalog. Clef reviews the listing metadata, outbound links, icon, and screenshots for phishing, impersonation, offensive content, scams, and attempts to manipulate the moderation system. The registry is built on AT Protocol, so plugin authors publish signed releases from their own accounts. Automated moderation lets authors keep that open publishing model without making the catalog an open door for abuse.
Clef is a decision model: instead of generating prose, it answers bounded questions with typed probabilities that application code can act on. When a plugin is published, the Em Dash labeler asks Clef nine questions about its text and links, covering explicit sexual content, hateful or dehumanizing content, graphic violence, phishing, impersonation, scams, spam, deceptive links, misleading media, and attempts to manipulate moderation. Eight corresponding questions are asked about each icon and screenshot. A result of 0.45 or higher sends the listing to an operator for review.
The previous moderation pipeline needed two general-purpose models to agree on text, plus another pass for images. According to Matt Kane, Em Dash lead maintainer, "Previously we had to use a mix of models to get reliable results, but Clef beat them all, while being much faster." Across three repeats of 21 public text fixtures, Clef made the expected decision in all 63 runs, with end-to-end text moderation latency of 1.64 seconds at p95. The moderation change doesn't add a new step for plugin authors.
Source: Hacker News · Summarized by HeadlinesBriefing