Docs / Use-case guides
Content moderation
Filter abusive language in comments, chat, usernames and reviews before it reaches other users.
The problem#
Open text fields invite spam and abuse. Manual moderation does not scale, and keyword lists miss misspellings and context. You want a fast first pass that handles the clear cases and queues the uncertain ones for people.
Signals to combine#
- Profanity detection: a
riskScoreand anisSafeverdict for a piece of text, with the option to list the matched words. - Email scoring and IP reputation for repeat offenders who register throwaway accounts.
- Blacklists for users or networks you have already banned.
Integration flow#
- When a user submits text, call profanity detection from your server. Pass
userIDso abusive authors are traceable in events. - If
isSafeis true, publish immediately. - If it is unsafe with a high
riskScore, reject and explain which rule was broken. - For borderline results, publish as pending and queue for a moderator.
- Count repeated violations per user and escalate to a temporary mute.
<?php
$text = $_POST['comment'];
$url = 'https://gurdx.cretip.com/api/scoring/profanity?' . http_build_query([
'text' => $text,
'userID' => $user->id,
]);
$res = json_decode(file_get_contents($url, false, stream_context_create(['http' => [
'header' => 'Authorization: Bearer ' . getenv('GURDX_KEY'),
'timeout' => 4,
]])), true);
if (($res['status'] ?? '') !== 'success') {
$state = 'pending'; // fail safe: queue for a human
} elseif ($res['data']['isSafe']) {
$state = 'published';
} elseif ($res['data']['riskScore'] >= 0.8) {
$state = 'rejected';
} else {
$state = 'pending';
}
Suggested thresholds#
The score for this endpoint is on a 0 to 1 scale, and Gurdx raises an event from 0.5 upwards.
| Result | Action |
|---|---|
isSafe true |
Publish |
| Risk 0.5 to under 0.8 | Hold for moderator review |
| Risk 0.8 or more | Reject, or auto-hide and notify |
Use stricter limits for children's audiences and looser ones for adult forums. Send long texts in chunks if you hit request limits.
Warning: Classifiers make mistakes, particularly with quotes, reclaimed words and mixed languages. Let users appeal and review a sample of auto-rejected content regularly.
Send a chat alert for profanity events to your moderators' channel. Related: Fake account prevention.
Found a mistake? Tell us on the contact page. Contact