Legal · DSA
Transparency Report
Published in accordance with the EU Digital Services Act (Articles 15 and 24) and the UK Online Safety Act 2023.
1. Content moderation practices
NaughtyBot employs a multi-layered moderation system:
- Heuristic keyword detection — a fast pre-check that scans user messages for patterns associated with child exploitation, coercion, and other prohibited categories. Real-time, before messages are processed.
- AI-based classification — a dedicated moderation model evaluates message content against safety categories including CSAM, non-consensual content, and self-harm. Returns a classification and confidence score.
- Stateful context tracking — per-conversation state tracks cumulative safety signals. Repeated violations result in progressively stricter moderation thresholds.
2. Automated decision-making
- Soft block — the AI redirects the conversation away from potentially harmful content (e.g., non-consensual scenarios). The user can continue using the service.
- Hard block — the message is rejected and the user is notified. Applies to content involving minors (absolute zero tolerance) and content that fails moderation due to system errors (fail-safe behaviour).
- Account suspension — triggered by detection of child exploitation content. Account is locked, evidence is preserved, and the incident is queued for human review.
All account suspension decisions triggered by automated moderation are subject to human review. No automated decision has legal or similarly significant effects on users beyond service access.
3. Self-harm safety protocol
NaughtyBot publishes its self-harm safety protocol in line with emerging companion-chatbot safety law (e.g. California SB 243 and comparable US state regimes).
- Detection — the moderation classifier includes a dedicated self-harm category, evaluated on user messages alongside the other safety categories described above.
- Response — a positive self-harm signal triggers a soft block: the AI declines to continue that theme and redirects the conversation, and the interface displays on-screen crisis resources (in the US, call or text 988 — the Suicide & Crisis Lifeline; internationally, findahelpline.com). The notice makes clear that NaughtyBot is an AI and cannot provide crisis support.
- No engagement optimisation on crisis content — we do not apply engagement-, retention-, or upsell-oriented behaviour to conversations exhibiting self-harm signals. The response is to de-escalate and surface help, not to prolong the session.
- AI-disclosure cadence — the chat interface carries a persistent "AI" badge identifying the companion as an AI system, and shows a periodic reminder (approximately every three hours of active use) that the persona is an AI companion, not a human.
4. Reporting illegal content
- The "Report" button available in the chat interface
- Our public report page, which routes legal notices and complaints to abuse@naughtybot.me (general/urgent) and dsa-contact@naughtybot.me (formal notices)
- Email to dsa-contact@naughtybot.me
Reports are acknowledged within 24 hours and reviewed within 7 days. Reporters are informed of the outcome and reasoning. Illegal content is removed and reported to the relevant authorities.
5. Appeal mechanism
Users whose accounts have been suspended or whose content has been moderated may appeal by emailing appeals@naughtybot.me. Appeals are reviewed by a human within 14 days. Human review may reverse an automated decision.
6. Point of contact
For the EU Digital Services Act, our single point of contact for EU Member State authorities, the European Commission, and the European Board for Digital Services is:
Email: dsa-contact@naughtybot.me
Language: English
Operator: Arnhem Labs
7. Moderation statistics
Updated periodically with aggregate statistics on moderation actions, user reports received, and outcomes. The first report covers the period from service launch.
Last updated · April 2026