# Open WebUI with Spicy API: uncensored role-play models

Open WebUI is the self-hosted chat interface for local and API models, usually run with Docker. This guide connects it to Spicy API (spicyapi.com), so Open WebUI can use `spicy-companion-1`, a role-play model that holds a persona and writes explicit adult scenes, or `spicy-companion-1-flash`, the same tuning for about a ninth of the cost of a long chat. You pay per token from a prepaid balance, with no subscription.

## What you need

- A Spicy API account: sign in at https://www.spicyapi.com/auth and accept the terms.
- A separate key for Open WebUI: https://www.spicyapi.com/dashboard/api-keys, then Create key. You can revoke it without breaking anything else. Open WebUI runs on your own machine and calls the API from there.
- A monthly spend limit on the account (same page, Monthly spend limit). Past it, requests fail with a 402 and nothing is charged, so a leaked key cannot drain your balance.
- A balance (top up from $50 by card or crypto at https://www.spicyapi.com/dashboard/account), or a sandbox key to try the setup first: tick Sandbox when you create the key. It answers with a fixed reply and bills nothing.

## Connect Open WebUI

1. Open **Settings**, then **Admin: AI**, **Connections**, and click the gear on the OpenAI API connection row (or + to add one).
2. Set **URL** to `https://api.spicyapi.com/v1`, **Auth** to Bearer with your key, and **API Type** to Chat Completions. Click the circular arrows: "Server connection verified".
3. Under **Advanced**, **Model IDs**, add `spicy-companion-1`, `spicy-companion-1-flash` and `spicy-chat-1`, then **Save**. Without this the model menu lists all 29 Spicy API models and new chats start on the first one, an image model.
4. Go to **Admin: AI**, **Models**, **Model Defaults**, **Model Capabilities**, untick **Builtin Tools** and save. Open WebUI otherwise attaches about 35 tools of its own to every request: on `spicy-companion-1-flash` every message then fails with a 400 ("`tools` is not available"), and on the other models their definitions are billed as prompt tokens on every message.
5. Under **Admin: Experience**, **Interface**, **Tasks**, set **External Task Model** to `spicy-companion-1-flash` (see Recommended settings for why).
6. Start a chat and pick `spicy-companion-1`.

![Open WebUI Edit Connection dialog: URL https://api.spicyapi.com/v1, Bearer auth with the key hidden, Chat Completions, and the three Spicy API chat models under Model IDs](https://www.spicyapi.com/docs/integrations/open-webui-connection.webp)

![An Open WebUI chat answered by spicy-companion-1 (a sandbox key returns a fixed reply)](https://www.spicyapi.com/docs/integrations/open-webui-chat.webp)

Or set it all when you first create the container (tested on a clean install; Open WebUI reads these only on its first start):

```bash
docker run -d -p 3000:8080 --name open-webui -v open-webui:/app/backend/data \
  -e OPENAI_API_BASE_URL=https://api.spicyapi.com/v1 \
  -e OPENAI_API_KEY=sk-spicy-... \
  -e 'OPENAI_API_CONFIGS={"0":{"model_ids":["spicy-companion-1","spicy-companion-1-flash","spicy-chat-1"]}}' \
  -e DEFAULT_MODELS=spicy-companion-1 \
  -e 'DEFAULT_MODEL_METADATA={"capabilities":{"builtin_tools":false}}' \
  -e TASK_MODEL_EXTERNAL=spicy-companion-1-flash \
  -e ENABLE_OLLAMA_API=False \
  ghcr.io/open-webui/open-webui:main
```

## Which model to pick

| Model | Best for | Context (tokens) | Per 1M tokens (prompt / completion) | 100 messages at 4k / 16k context |
|---|---|---|---|---|
| `spicy-companion-1` | Role-play and companions: holds a persona, writes explicit adult scenes, group scenes. | 131,072 | $1.00 / $2.80 | $0.47 / $1.71 |
| `spicy-companion-1-flash` | The same role-play tuning for about a ninth of the cost of a long chat. Fast. | 43,000 tested | $0.10 / $0.80 | $0.06 / $0.19 |
| `spicy-chat-1` | General assistant. Not tuned for role-play; use a companion model for characters. | 200,000 | $0.80 / $2.40 | $0.38 / $1.38 |

Chat apps resend the whole conversation with every message, so the context size you set drives the cost far more than the length of the replies. The last column assumes a 250 token reply at 4k context and 400 at 16k. Every response carries its exact `cost_usd`.

## Recommended settings

| Setting | Value |
|---|---|
| Builtin Tools (Model Capabilities) | Off. With it on, every chat fails with a 400 on spicy-companion-1-flash and costs more on the other models. |
| Model IDs | The three chat models only, and `spicy-companion-1` as the default model. |
| External Task Model | `spicy-companion-1-flash`. Open WebUI makes extra calls for titles, follow-up suggestions and tags: four calls on a chat's first message and two on each later one. Or switch off Title, Follow Up and Tags Generation to make it one call per message. |

## If something goes wrong

| Status | Meaning | Fix |
|---|---|---|
| 400 | Wrong model id, or an option the chat models do not take (tools on `spicy-companion-1-flash`, images, more than 8,192 max tokens) | Pick `spicy-companion-1`, `spicy-companion-1-flash` or `spicy-chat-1`; the error names the chat models |
| 401 | Key missing, mistyped or revoked | Paste the key again, or make a new one |
| 402 | Balance empty, or the monthly spend limit reached | Top up, or raise the limit on the API keys page. Nothing is charged |
| 422 | Blocked by moderation | Nothing is charged. See what is allowed below |
| 429 | Too many requests in a minute | Wait a minute; the app can retry |

- Errors appear word for word in place of the reply, for example "spicy-animate-1 is a video-edit model, not a chat model. Chat models: ...".
- Open WebUI sends only the model, the messages and `stream: true`, so the model's own defaults apply unless you set parameters in the chat controls.
- On a Mac with an Apple M5 chip the current image crashed at start (exit code 132) until `-e OPENSSL_armcap=0` was added. That is an Open WebUI build issue, not a Spicy API one.

## What is allowed

Explicit role-play between adults is allowed, including non-consent between fictional adult characters (force, sleep, intoxication, framed as consensual non-consent or not) and other dark themes. Every message is screened before it reaches the model: minors in any form, real people (with or without consent), incest, real-world harm and the rest of the prohibited list are refused with a 422, before any charge. The screen reads the whole conversation, character card included, so a card that describes a minor or a real person is refused however the latest message is worded. Refusals in the most serious categories count as strikes. Send each user's id as `user`: the strikes then land on that user, who is refused after ten, and your account is only flagged. Without it they land on your account, flagged after three and suspended after ten. Drop a refused message from the history you send back, or the next message is refused too. A wrong refusal can be reported: POST its `moderation_id` to /v1/moderations/appeals. Full policy: https://www.spicyapi.com/acceptable-use

## Tested

Open WebUI 0.11.4 (ghcr.io/open-webui/open-webui:main of 21 Sep 2026) in Docker on macOS, 30 Sep 2026: connection verify, the Builtin Tools fix, streamed chat with a sandbox key, the task calls and both ways to reduce them, error display, and the environment variables above on a clean install.

Other apps: https://www.spicyapi.com/docs/integrations. Questions: contact@spicyapi.com.
