Free AI content audit
Run moderation on Ollama Cloud or a self-hosted model — no OpenAI Moderations bill for everyday screening.
Screen text with Ollama — cloud or your own machine — through the same
/v1/moderations API your tools already speak.
Lightweight, multi-key ready, and especially easy with
sub2api.
中文:免费的 Ollama AI 内容审核。兼容 OpenAI Moderations,可直接对接 sub2api。 轻量、多 Key 轮换、即插即用。质量不等同于 OpenAI 官方审核。 完整中文页面 →
Skip OpenAI Moderations fees. Keep the client you already have. Add free Ollama-backed content audit in minutes.
Run moderation on Ollama Cloud or a self-hosted model — no OpenAI Moderations bill for everyday screening.
Same POST /v1/moderations shape. Point the OpenAI SDK, sub2api, or any Moderations client at this gateway.
Drop several Ollama keys in. The gateway round-robins and cools down keys that hit 401 / 403 / 429 / quota — traffic keeps moving.
One small service. Docker Compose, Helm on Kubernetes, or a quick local start. No heavy platform, no custom client to rewrite.
Defaults to Ollama Cloud. Prefer private? Point it at your local Ollama and keep data on your network.
If the model is down or the response is unusable, you get an error — never a made-up “all clear.”
The sweet spot: sub2api content audit (风控中心 · 内容审计) already expects an OpenAI-style Moderations backend. This gateway is that backend — free Ollama scoring behind the UI you already use.
/v1
Use your gateway:
http://localhost:8000 (local) or
https://moderation.example.com (your deploy) —
without trailing /v1.
sub2api appends /v1/moderations;
adding /v1 yourself breaks the path.
moderation-fast
(or the omni-moderation-latest alias if your UI prefers that name).
Use the gateway’s MODERATION_API_KEY,
not an Ollama Cloud key.
Use the key you set on your own deploy.
timeout_ms to 30000
Default 3000 is too short for cloud LLMs.
sub2api: github.com/Wei-Shaw/sub2api
Primary path: pip install, Docker Compose, or Helm.
An optional Render demo may sleep / cold-start — do not rely on it.
Primary production path. Chart at charts/ollama-moderation-gateway (v0.1.0). Create the Secret out-of-band, then helm upgrade --install.
Copy .env.example to .env, set keys, docker compose up --build. Port 8000. Fine for one machine.
onrender.com may exist for poking the API but often sleeps / times out. Prefer local or Helm (example host moderation.example.com).
Prefer an existing Secret — the chart does not create one by default.
Image: ghcr.io/aaronyang0628/ollama-moderation-gateway:v0.1.0 (GHCR; private packages may need imagePullSecrets).
Proxy env vars default empty — set only if your cluster needs egress proxy (keep real endpoints out of git).
Ingress defaults to nginx + cert-manager with placeholder host moderation.example.com (override to yours).
kubectl create namespace moderation kubectl -n moderation create secret generic ollama-moderation-gateway \ --from-literal=OLLAMA_API_KEYS='key1,key2' \ --from-literal=MODERATION_API_KEYS='gw-key1,gw-key2' helm upgrade --install ollama-moderation-gateway ./charts/ollama-moderation-gateway \ -n moderation --create-namespace \ --set secrets.create=false \ --set secrets.existingSecret=ollama-moderation-gateway \ --set ingress.hosts[0].host=moderation.example.com \ --set ingress.tls[0].hosts[0]=moderation.example.com \ --set ingress.tls[0].secretName=moderation.example.com-tls
Smoke: curl https://moderation.example.com/health
then POST /v1/moderations with a gateway Bearer key.
Chart notes:
charts/ollama-moderation-gateway/README.md.
cp .env.example .env # Set MODERATION_API_KEY, OLLAMA_API_KEYS (for Cloud), APP_ENV=production docker compose up --build
Three steps. No architecture lecture.
OpenAI SDK, sub2api, curl — anything that posts to /v1/moderations.
Keys rotate. You choose cloud or self-hosted. Models like moderation-fast map to Ollama chat models.
OpenAI-style flagged / categories / scores — ready for your existing risk flow.
| You send | Ollama runs | Notes |
|---|---|---|
moderation-fast |
gpt-oss:20b |
Good default |
moderation-standard |
gemma4:31b |
Heavier option |
moderation-accurate |
gpt-oss:120b |
Largest; needs serious hardware self-hosted |
omni-moderation-latest |
gpt-oss:20b |
Name alias only — not OpenAI quality |
pip install ollama-moderation-gateway first. Placeholder keys are for local smoke only. Production: Helm.
pip install ollama-moderation-gateway export APP_ENV=development export MODERATION_API_KEY=local-moderation-key # Cloud: export OLLAMA_API_KEYS=... # Self-host: export OLLAMA_BASE_URL=http://127.0.0.1:11434 ollama-moderation-gateway
From source: pip install -e ".[dev]" then
uvicorn app.main:app --host 0.0.0.0 --port 8000.
curl http://localhost:8000/v1/moderations \
-H 'Authorization: Bearer local-moderation-key' \
-H 'Content-Type: application/json' \
-d '{
"model": "moderation-fast",
"input": "Text to screen"
}'
Primary: local curl above. Optional demo
onrender.com
may sleep / cold-start. Production example:
moderation.example.com — see Deploy.
This is not the same quality as OpenAI’s official moderation models
(omni-moderation-latest). It uses general Ollama chat models for free / self-hosted content audit —
great for cost and control, different from dedicated OpenAI safety classifiers. Try it on your own traffic before you rely on it.
pip install, Compose, or Helm — then point sub2api or your SDK at it.