Hate speech detection

A hate speech detection API for Turkish and English.

Send a comment, a message or the address of an image. Detectivision answers which kind of hate it found, if any, and how sure it is: hate, racism or sexism in text, hate in images.

What it finds

Three kinds of hate in text, and hateful images

Each is a flag of its own, so your rules can treat them apart.

  • Hate

    Hate against a group of people, such as a race, ethnicity, religion, gender or sexual orientation.

  • Racism

    Racist or ethnic hate speech.

  • Sexism

    Sexist statements.

  • Insults

    An attack on one person rather than on a group is a flag of its own: insult.

  • Hateful images

    Images that spread hate against a group of people: the image flag hate.

In code

Check a comment

Request
curl https://api.detectivision.ai/api/v1/moderation \
  -H "Api-Key: $DETECTIVISION_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "type": "text_moderation",
    "input": "People like you should go back to where you came from.",
    "mode": "detect",
    "flags": ["hate", "racism", "sexism"]
  }'
Response
{
  "id": 2071,
  "type": "text_moderation",
  "mode": "detect",
  "status": "completed",
  "flagged": true,
  "flags": [
    { "flag": "racism", "confidence": 0.88 },
    { "flag": "hate", "confidence": 0.81 }
  ],
  "checked": ["hate", "racism", "sexism"],
  "unchecked": [],
  "tookMs": 640,
  "usage": { "charged": 2, "remaining": 9998 }
}

The flags raised come strongest first. Only the flags you name are checked, so a narrow request answers faster.

How it works
  1. Pick your flags

    Check hate, racism and sexism together, or one of them. GET /api/v1/flags lists the keys your account may send.

  2. Draw your line

    Act on flagged first, then tune with the confidence: for example, send answers between 0.5 and 0.7 to a moderator.

  3. Keep it traceable

    Put your user and comment ids in metadata. They come back in the answer, in Logs and in every webhook, so an action reaches the right account.

Where it helps
Questions

Frequently asked

Which languages does it understand?

Turkish and English. Text in other languages is checked, but the results are not reliable.

What is the difference between hate and an insult?

Hate is aimed at a group of people for who they are; an insult attacks one person. They are separate flags, hate and insult, so a platform can remove hate at once and only warn for insults.

How long can a text be?

1 to 10,000 characters. A text costs 1 credit per started 2,500 characters when the answer comes by webhook, and twice that when it comes at once.

Can I check usernames and bios?

Yes. Any text from 1 to 10,000 characters, as it was written.

Is there a free trial?

Yes. The plans, and the free trial, are on the pricing page.