Chat apps

Open WebUI with Spicy API: uncensored role-play models

Open WebUI is the self-hosted chat interface for local and API models, usually run with Docker. This guide connects it to Spicy API (spicyapi.com), so Open WebUI can use spicy-companion-1, a role-play model that holds a persona and writes explicit adult scenes, or spicy-companion-1-flash, the same tuning for about a ninth of the cost of a long chat. You pay per token from a prepaid balance, with no subscription.

What you need

  • A Spicy API account: sign in at /auth and accept the terms.
  • A separate key for Open WebUI: /dashboard/api-keys, then Create key. You can revoke it without breaking anything else. Open WebUI runs on your own machine and calls the API from there.
  • A monthly spend limit on the account (same page, Monthly spend limit). Past it, requests fail with a 402 and nothing is charged, so a leaked key cannot drain your balance.
  • A balance (top up from $50 by card or crypto at /dashboard/account), or a sandbox key to try the setup first: tick Sandbox when you create the key. It answers with a fixed reply and bills nothing.

Connect Open WebUI

  1. Open Settings, then Admin: AI, Connections, and click the gear on the OpenAI API connection row (or + to add one).
  2. Set URL to https://api.spicyapi.com/v1, Auth to Bearer with your key, and API Type to Chat Completions. Click the circular arrows: "Server connection verified".
  3. Under Advanced, Model IDs, add spicy-companion-1, spicy-companion-1-flash and spicy-chat-1, then Save. Without this the model menu lists all 29 Spicy API models and new chats start on the first one, an image model.
  4. Go to Admin: AI, Models, Model Defaults, Model Capabilities, untick Builtin Tools and save. Open WebUI otherwise attaches about 35 tools of its own to every request: on spicy-companion-1-flash every message then fails with a 400 ("tools is not available"), and on the other models their definitions are billed as prompt tokens on every message.
  5. Under Admin: Experience, Interface, Tasks, set External Task Model to spicy-companion-1-flash (see Recommended settings for why).
  6. Start a chat and pick spicy-companion-1.
Open WebUI Edit Connection dialog: URL https://api.spicyapi.com/v1, Bearer auth with the key hidden, Chat Completions, and the three Spicy API chat models under Model IDsAn Open WebUI chat answered by spicy-companion-1 (a sandbox key returns a fixed reply)

Or set it all when you first create the container (tested on a clean install; Open WebUI reads these only on its first start):

docker run -d -p 3000:8080 --name open-webui -v open-webui:/app/backend/data \
  -e OPENAI_API_BASE_URL=https://api.spicyapi.com/v1 \
  -e OPENAI_API_KEY=sk-spicy-... \
  -e 'OPENAI_API_CONFIGS={"0":{"model_ids":["spicy-companion-1","spicy-companion-1-flash","spicy-chat-1"]}}' \
  -e DEFAULT_MODELS=spicy-companion-1 \
  -e 'DEFAULT_MODEL_METADATA={"capabilities":{"builtin_tools":false}}' \
  -e TASK_MODEL_EXTERNAL=spicy-companion-1-flash \
  -e ENABLE_OLLAMA_API=False \
  ghcr.io/open-webui/open-webui:main

Which model to pick

ModelBest forContext (tokens)Per 1M tokens (prompt / completion)100 messages at 4k / 16k context
spicy-companion-1Role-play and companions: holds a persona, writes explicit adult scenes, group scenes.131,072$1.00 / $2.80$0.47 / $1.71
spicy-companion-1-flashThe same role-play tuning for about a ninth of the cost of a long chat. Fast.43,000 tested$0.10 / $0.80$0.06 / $0.19
spicy-chat-1General assistant. Not tuned for role-play; use a companion model for characters.200,000$0.80 / $2.40$0.38 / $1.38

Chat apps resend the whole conversation with every message, so the context size you set drives the cost far more than the length of the replies. The last column assumes a 250 token reply at 4k context and 400 at 16k. Every response carries its exact cost_usd.

SettingValue
Builtin Tools (Model Capabilities)Off. With it on, every chat fails with a 400 on spicy-companion-1-flash and costs more on the other models.
Model IDsThe three chat models only, and spicy-companion-1 as the default model.
External Task Modelspicy-companion-1-flash. Open WebUI makes extra calls for titles, follow-up suggestions and tags: four calls on a chat's first message and two on each later one. Or switch off Title, Follow Up and Tags Generation to make it one call per message.

If something goes wrong

StatusMeaningFix
400Wrong model id, or an option the chat models do not take (tools on spicy-companion-1-flash, images, more than 8,192 max tokens)Pick spicy-companion-1, spicy-companion-1-flash or spicy-chat-1; the error names the chat models
401Key missing, mistyped or revokedPaste the key again, or make a new one
402Balance empty, or the monthly spend limit reachedTop up, or raise the limit on the API keys page. Nothing is charged
422Blocked by moderationNothing is charged. See what is allowed below
429Too many requests in a minuteWait a minute; the app can retry
  • Errors appear word for word in place of the reply, for example "spicy-animate-1 is a video-edit model, not a chat model. Chat models: ...".
  • Open WebUI sends only the model, the messages and stream: true, so the model's own defaults apply unless you set parameters in the chat controls.
  • On a Mac with an Apple M5 chip the current image crashed at start (exit code 132) until -e OPENSSL_armcap=0 was added. That is an Open WebUI build issue, not a Spicy API one.

What is allowed

Explicit role-play between adults is allowed, including non-consent between fictional adult characters (force, sleep, intoxication, framed as consensual non-consent or not) and other dark themes. Every message is screened before it reaches the model: minors in any form, real people (with or without consent), incest, real-world harm and the rest of the prohibited list are refused with a 422, before any charge. The screen reads the whole conversation, character card included, so a card that describes a minor or a real person is refused however the latest message is worded. Refusals in the most serious categories count as strikes. Send each user's id as user: the strikes then land on that user, who is refused after ten, and your account is only flagged. Without it they land on your account, flagged after three and suspended after ten. Drop a refused message from the history you send back, or the next message is refused too. A wrong refusal can be reported: POST its moderation_id to /v1/moderations/appeals. Full policy: /acceptable-use

Tested

Open WebUI 0.11.4 (ghcr.io/open-webui/open-webui:main of 21 Sep 2026) in Docker on macOS, 30 Sep 2026: connection verify, the Builtin Tools fix, streamed chat with a sandbox key, the task calls and both ways to reduce them, error display, and the environment variables above on a clean install.

Other apps: /docs/integrations. Questions: contact@spicyapi.com.