Free · pip install · OpenAI Moderations compatible

Free AI content moderation that plugs right in

Screen text with Ollama — cloud or your own machine — through the same /v1/moderations API your tools already speak. Lightweight, multi-key ready, and especially easy with sub2api.

中文:免费的 Ollama AI 内容审核。兼容 OpenAI Moderations,可直接对接 sub2api。 轻量、多 Key 轮换、即插即用。质量不等同于 OpenAI 官方审核。 完整中文页面 →

Free Ollama audit Multi-key rotation Helm / Kubernetes sub2api ready

Why people use it

Skip OpenAI Moderations fees. Keep the client you already have. Add free Ollama-backed content audit in minutes.

Free AI content audit

Run moderation on Ollama Cloud or a self-hosted model — no OpenAI Moderations bill for everyday screening.

OpenAI Moderations compatible

Same POST /v1/moderations shape. Point the OpenAI SDK, sub2api, or any Moderations client at this gateway.

Multi-key rotation

Drop several Ollama keys in. The gateway round-robins and cools down keys that hit 401 / 403 / 429 / quota — traffic keeps moving.

Lightweight plug-in

One small service. Docker Compose, Helm on Kubernetes, or a quick local start. No heavy platform, no custom client to rewrite.

Cloud or your own box

Defaults to Ollama Cloud. Prefer private? Point it at your local Ollama and keep data on your network.

Doesn’t fake “safe”

If the model is down or the response is unusable, you get an error — never a made-up “all clear.”

Plug into sub2api in minutes

The sweet spot: sub2api content audit (风控中心 · 内容审计) already expects an OpenAI-style Moderations backend. This gateway is that backend — free Ollama scoring behind the UI you already use.

One change. Keep sub2api’s risk-control screens. Swap the Moderations base URL to this gateway. No custom client, no new protocol — just free Ollama AI moderation on the back end.

Quick setup

1

Base URL — no trailing /v1

Use your gateway: http://localhost:8000 (local) or https://moderation.example.com (your deploy) — without trailing /v1. sub2api appends /v1/moderations; adding /v1 yourself breaks the path.

2

Model

moderation-fast (or the omni-moderation-latest alias if your UI prefers that name).

3

API Key = gateway key

Use the gateway’s MODERATION_API_KEY, not an Ollama Cloud key. Use the key you set on your own deploy.

4

Set timeout_ms to 30000

Default 3000 is too short for cloud LLMs.

sub2api: github.com/Wei-Shaw/sub2api

Deploy it for real

Primary path: pip install, Docker Compose, or Helm. An optional Render demo may sleep / cold-start — do not rely on it.

Helm on Kubernetes

Primary production path. Chart at charts/ollama-moderation-gateway (v0.1.0). Create the Secret out-of-band, then helm upgrade --install.

Docker Compose

Copy .env.example to .env, set keys, docker compose up --build. Port 8000. Fine for one machine.

Optional demo

onrender.com may exist for poking the API but often sleeps / times out. Prefer local or Helm (example host moderation.example.com).

Kubernetes with Helm

Prefer an existing Secret — the chart does not create one by default. Image: ghcr.io/aaronyang0628/ollama-moderation-gateway:v0.1.0 (GHCR; private packages may need imagePullSecrets). Proxy env vars default empty — set only if your cluster needs egress proxy (keep real endpoints out of git). Ingress defaults to nginx + cert-manager with placeholder host moderation.example.com (override to yours).

kubectl create namespace moderation
kubectl -n moderation create secret generic ollama-moderation-gateway \
  --from-literal=OLLAMA_API_KEYS='key1,key2' \
  --from-literal=MODERATION_API_KEYS='gw-key1,gw-key2'

helm upgrade --install ollama-moderation-gateway ./charts/ollama-moderation-gateway \
  -n moderation --create-namespace \
  --set secrets.create=false \
  --set secrets.existingSecret=ollama-moderation-gateway \
  --set ingress.hosts[0].host=moderation.example.com \
  --set ingress.tls[0].hosts[0]=moderation.example.com \
  --set ingress.tls[0].secretName=moderation.example.com-tls

Smoke: curl https://moderation.example.com/health then POST /v1/moderations with a gateway Bearer key. Chart notes: charts/ollama-moderation-gateway/README.md.

Docker Compose

cp .env.example .env
# Set MODERATION_API_KEY, OLLAMA_API_KEYS (for Cloud), APP_ENV=production
docker compose up --build

How it fits together

Three steps. No architecture lecture.

1

Your app calls Moderations

OpenAI SDK, sub2api, curl — anything that posts to /v1/moderations.

2

Gateway asks Ollama

Keys rotate. You choose cloud or self-hosted. Models like moderation-fast map to Ollama chat models.

3

You get a familiar result

OpenAI-style flagged / categories / scores — ready for your existing risk flow.

Model aliases

You send Ollama runs Notes
moderation-fast gpt-oss:20b Good default
moderation-standard gemma4:31b Heavier option
moderation-accurate gpt-oss:120b Largest; needs serious hardware self-hosted
omni-moderation-latest gpt-oss:20b Name alias only — not OpenAI quality

Get running in minutes

pip install ollama-moderation-gateway first. Placeholder keys are for local smoke only. Production: Helm.

pip (PyPI)

pip install ollama-moderation-gateway
export APP_ENV=development
export MODERATION_API_KEY=local-moderation-key
# Cloud: export OLLAMA_API_KEYS=...
# Self-host: export OLLAMA_BASE_URL=http://127.0.0.1:11434
ollama-moderation-gateway

From source: pip install -e ".[dev]" then uvicorn app.main:app --host 0.0.0.0 --port 8000.

curl

curl http://localhost:8000/v1/moderations \
  -H 'Authorization: Bearer local-moderation-key' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "moderation-fast",
    "input": "Text to screen"
  }'

Primary: local curl above. Optional demo onrender.com may sleep / cold-start. Production example: moderation.example.com — see Deploy.

One honest note

This is not the same quality as OpenAI’s official moderation models (omni-moderation-latest). It uses general Ollama chat models for free / self-hosted content audit — great for cost and control, different from dedicated OpenAI safety classifiers. Try it on your own traffic before you rely on it.

Free Ollama moderation, ready to plug in

pip install, Compose, or Helm — then point sub2api or your SDK at it.