Developer Dashboard

Moderation Guide

Use the AvalAI Moderation API to check whether text or image inputs are potentially harmful according to defined content policies. This helps ensure safety and compliance in your applications.

If harmful content is identified, you can take corrective action, like filtering content or flagging user accounts. Access to the moderation endpoint via AvalAI might be free or subject to specific pricing; please check AvalAI's Pricing page.

AvalAI provides access to moderation models, potentially including:

  • omni-moderation-latest (Recommended): Supports more categories and multi-modal (text + image) inputs.
  • text-moderation-latest (Legacy): Supports only text inputs and fewer categories.

Check AvalAI's Models Overview for currently available moderation models.

Moderation Decision Workflow

OpenAI's moderation guidance is most useful when it becomes a product workflow rather than a single API call. For AvalAI applications:

  1. Classify input before generation when the user can submit open text, images, URLs, files, or retrieved content.
  2. Generate only after the request passes policy or after you route it to a constrained safe-completion/refusal path.
  3. Classify generated output before display for public, social, marketplace, education, or under-18 surfaces.
  4. Tune thresholds with evals: treat flagged as a strong default signal, but calibrate custom category_scores thresholds with examples from your product and human review labels.
  5. Keep a review trail: log avalai-request-id, model, route, hashed user or safety_identifier, moderation categories, and the final product action without storing unnecessary personal data.
  6. Offer escalation: define what happens for blocked, borderline, and appealed decisions instead of silently dropping user requests.

Quickstart

Moderate Text Inputs

Get classification information for a text input:

python
# Python Example using AvalAI Moderation API
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",  # AvalAI API endpoint
)

try:
    response = client.moderations.create(
        model="omni-moderation-latest",  # Or another model available via AvalAI
        input="Sample text that might violate content policy.",
    )
    print(response)
except Exception as e:
    print(f"An error occurred: {e}")
javascript
// JavaScript Example using AvalAI Moderation API
import { OpenAI } from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY, // Ensure AVALAI_API_KEY is set

  baseURL: "https://api.avalai.ir/v1", // Use AvalAI base URL
});

async function main() {
  try {
    const moderation = await client.moderations.create({
      model: "omni-moderation-latest", // Or another model available via AvalAI

      input: "Sample text that might violate content policy.",
    });
    console.log(moderation);
  } catch (error) {
    console.error("Error calling moderation API: ", error);
  }
}
main();
bash
# cURL Example using AvalAI Moderation API
curl https://api.avalai.ir/v1/moderations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d '{
  "model": "omni-moderation-latest",
  "input": "Sample text that might violate content policy."
}'
php
<?php
// PHP Example using AvalAI Moderation API
require_once 'vendor/autoload.php';

// Using OpenAI PHP client library (https://github.com/openai-php/client)
$apiKey = getenv('AVALAI_API_KEY'); // Or replace with your actual key: 'aa-YOUR_API_KEY'

if (!$apiKey) {
    die("AvalAI API key not found. Please set the AVALAI_API_KEY environment variable.");
}

// Your custom base URL
$customBaseUrl = 'https://api.avalai.ir/v1';

// Create a custom client instance using the factory
$client = OpenAI::factory()
    ->withApiKey($apiKey)
    ->withBaseUri($customBaseUrl)
    ->make();

try {
  // Make the moderation request
  $response = $client->moderations()->create([
  'model' => 'omni-moderation-latest',
  'input' => 'Sample text that might violate content policy.'
  ]);

  // Output the response
  print_r($response->toArray());
} catch (\Exception $e) {
  echo "Error: " . $e->getMessage() . "\n";
}
go
package main

import (
	"context"
	"fmt"
	openai "github.com/openai/openai-go"
)

func main() {
	client := openai.NewClient("AVALAI_API_KEY")
	client.BaseURL = "https://api.avalai.ir/v1"

	resp, err := client.Moderations(
		context.Background(),
		openai.ModerationRequest{

			Input: "Sample text that might violate content policy.",

			Model: openai.ModerationLatest,
		},
	)

	if err != nil {
		fmt.Printf("Moderation error: %v\n", err)
		return
	}

	// Check if the text is flagged
	if resp.Results[0].Flagged {
		fmt.Println("This content was flagged!")
	}

	// Check specific categories
	for category, score := range resp.Results[0].CategoryScores {
		if score > 0.5 {
			fmt.Printf("Content flagged for %s with score %.2f\n", category, score)
		}
	}
}

Moderate Image and Text Inputs (Multi-modal)

Requires a multi-modal moderation model like omni-moderation-latest.

Get classification information for combined image and text input:

python
# Python Example using AvalAI Multi-modal Moderation
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",  # AvalAI API endpoint
)

try:
    response = client.moderations.create(
        model="omni-moderation-latest",  # Ensure model supports multi-modal
        input=[
            {"type": "text", "text": "Description accompanying the image."},
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://example.com/image_to_moderate.png"
                    # Or Base64: "url": "data:image/png;base64,abcdefg..."
                },
            },
        ],
    )
    print(response)
except Exception as e:
    print(f"An error occurred: {e}")
javascript
// JavaScript Example using AvalAI Multi-modal Moderation
import { OpenAI } from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

async function main() {
  try {
    const moderation = await client.moderations.create({
      model: "omni-moderation-latest", // Ensure model supports multi-modal
      input: [
        { type: "text", text: "Description accompanying the image." },
        {
          type: "image_url",
          image_url: {
            url: "https://example.com/image_to_moderate.png",
            // Or Base64: url: "data:image/png;base64,abcdefg..."
          },
        },
      ],
    });
    console.log(moderation);
  } catch (error) {
    console.error("Error calling moderation API: ", error);
  }
}
main();
bash
# cURL Example using AvalAI Multi-modal Moderation
curl https://api.avalai.ir/v1/moderations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d '{
  "model": "omni-moderation-latest",
  "input": [
  { "type": "text", "text": "Description accompanying the image." },
  {
    "type": "image_url",
    "image_url": {
      "url": "https://example.com/image_to_moderate.png"
    }
  }
  ]
}'
php
<?php
// PHP Example using AvalAI Multi-modal Moderation
require_once 'vendor/autoload.php';

$apiKey = getenv('AVALAI_API_KEY'); // Or replace with your actual key: 'aa-YOUR_API_KEY'

if (!$apiKey) {
    die("AvalAI API key not found. Please set the AVALAI_API_KEY environment variable.");
}

// Your custom base URL
$customBaseUrl = 'https://api.avalai.ir/v1';

// Create a custom client instance using the factory
$client = OpenAI::factory()
    ->withApiKey($apiKey)
    ->withBaseUri($customBaseUrl)
    ->make();

try {
  $response = $client->moderations()->create([
  'model' => 'omni-moderation-latest',
  'input' => [
  [
  'type' => 'text',
  'text' => 'Description accompanying the image.'
  ],
  [
  'type' => 'image_url',
  'image_url' => [
  'url' => 'https://example.com/image_to_moderate.png'
  // Or Base64: 'url' => 'data:image/png;base64,abcdefg...'
  ]
  ]
  ]
  ]);

  print_r($response->toArray());
} catch (\Exception $e) {
  echo "Error: " . $e->getMessage() . "\n";
}
go
package main

import (
	"context"
	"fmt"
	openai "github.com/openai/openai-go"
)

func main() {
	client := openai.NewClient("AVALAI_API_KEY")
	client.BaseURL = "https://api.avalai.ir/v1"

	// Create the input structure for multi-modal moderation
	input := []openai.ModerationInput{
		{
			Type: "text",
			Text: "Description accompanying the image.",
		},
		{
			Type: "image_url",
			ImageURL: &openai.ImageURL{
				URL: "https://example.com/image_to_moderate.png",
			},
		},
	}

	resp, err := client.Moderations(
		context.Background(),
		openai.ModerationRequest{
			Input: input,
			Model: "omni-moderation-latest",
		},
	)

	if err != nil {
		fmt.Printf("Moderation error: %v\n", err)
		return
	}

	// Process the response
	if resp.Results[0].Flagged {
		fmt.Println("This content was flagged!")
	}

	// Check specific categories
	for category, score := range resp.Results[0].CategoryScores {
		if score > 0.5 {
			fmt.Printf("Content flagged for %s with score %.2f\n", category, score)
		}
	}
}

Moderate Generated Content Inline

When the selected AvalAI route supports OpenAI-compatible inline moderation, you can request moderation scores in the same /v1/responses or /v1/chat/completions call that generates the answer. This is useful when you need both the generated response and safety signals for the input/output pair.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.5",
    input="Draft a brief refusal and safe alternative for a harmful request.",
    moderation={"model": "omni-moderation-latest"},
)

if response.moderation.input.flagged or response.moderation.output.flagged:
    print("Queue for review before showing the response.")
else:
    print(response.output_text)

Treat inline moderation as a policy signal, not an automatic final decision. A safe refusal may still receive a high score because it discusses harmful content. For streaming responses, moderation scores arrive only after the full generated output is available; they are not included with partial deltas. If inline moderation is not available for the selected model or route, call POST /v1/moderations separately before display or downstream actions.

For tool workflows, moderation can cover tool-call arguments and tool outputs when they appear in conversation content. It does not moderate tool names, tool descriptions, tool schemas, or structured-output schemas, so validate those surfaces separately.

Understanding the Response

The API response provides details about potential policy violations:

json
{
  "id": "modr-...", // Moderation request ID

  "model": "omni-moderation-latest", // Model used

  "results": [
    {
      "flagged": true, // True if any category is flagged above threshold

      "categories": {
        // Boolean flags for each category

        "sexual": false,
        "hate": false,
        "harassment": false,
        "self-harm": false,
        "sexual/minors": false,
        "hate/threatening": false,
        "violence/graphic": false,
        "self-harm/intent": false,
        "self-harm/instructions": false,
        "harassment/threatening": false,
        "violence": true, // Example: Flagged for violence

        // Omni-specific categories:

        "illicit": false,
        "illicit/violent": false
      },
      "category_scores": {
        // Confidence scores (0-1) for each category

        "sexual": 0.0001,
        "hate": 0.0002,
        // ... other scores

        "violence": 0.987, // Example: High confidence for violence

        "violence/graphic": 0.123
        // ... omni scores

      },
      // Only present for omni models:

      "category_applied_input_types": {
        "sexual": ["text", "image"], // Which input type triggered flag

        "hate": ["text"],
        // ... other categories

        "violence": ["image"] // Example: Image triggered violence flag

      }
    }
  ]
}
  • flagged: Overall flag (true if any category score exceeds internal thresholds).
  • categories: Boolean flags indicating if a category is violated.
  • category_scores: Model's confidence score (0 to 1) for each category violation. Use these scores for custom policies, but note they might need recalibration if the underlying model is updated by the provider.
  • category_applied_input_types (Omni models only): Shows whether text or image input (or both) contributed to a category being flagged.

Content Classifications

The moderation endpoint checks for content across several categories. Availability and input type support (text/image) depend on the model used (omni models generally support more categories and image input).

CategoryDescriptionModelsInputs Supported
harassmentExpresses, incites, or promotes harassing language towards any target.AllText only
harassment/threateningHarassment that also includes violence or serious harm threats.AllText only
hateExpresses, incites, or promotes hate based on protected characteristics (race, gender, religion, etc.).AllText only
hate/threateningHateful content that also includes violence or serious harm threats towards the targeted group.AllText only
illicitAdvice or instructions on committing illicit acts (e.g., how to shoplift).Omni onlyText only
illicit/violentIllicit content that also references violence or procuring weapons.Omni onlyText only
self-harmPromotes, encourages, or depicts acts of self-harm (suicide, cutting, eating disorders).AllText & Image
self-harm/intentSpeaker expresses intent to engage in self-harm.AllText & Image
self-harm/instructionsEncourages or gives instructions for self-harm.AllText & Image
sexualContent meant to arouse sexual excitement or promote sexual services (excluding education/wellness).AllText & Image
sexual/minorsSexual content involving individuals under 18.AllText only
violenceDepicts death, violence, or physical injury.AllText & Image
violence/graphicDepicts death, violence, or physical injury in graphic detail.AllText & Image

(Note: "All" typically refers to both omni-moderation-latest and text-moderation-latest. "Omni only" refers to categories added with omni-moderation-latest and its snapshots).