KoboldAI Lite with Spicy API: uncensored role-play models
KoboldAI Lite is the free browser chat and story front end from the KoboldAI project, at lite.koboldai.net. This guide connects it to Spicy API (spicyapi.com), so KoboldAI Lite can use spicy-companion-1, a role-play model that holds a persona and writes explicit adult scenes, or spicy-companion-1-flash, the same tuning for about a ninth of the cost of a long chat. You pay per token from a prepaid balance, with no subscription.
What you need
- A Spicy API account: sign in at /auth and accept the terms.
- A separate key for KoboldAI Lite: /dashboard/api-keys, then Create key. You can revoke it without breaking anything else. KoboldAI Lite calls the API from your browser and keeps the key in its settings, so never reuse a key your own apps depend on.
- A monthly spend limit on the account (same page, Monthly spend limit). Past it, requests fail with a 402 and nothing is charged, so a leaked key cannot drain your balance.
- A balance (top up from $50 by card or crypto at /dashboard/account), or a sandbox key to try the setup first: tick Sandbox when you create the key. It answers with a fixed reply and bills nothing.
Connect KoboldAI Lite
- Open https://lite.koboldai.net and click AI at the top left.
- Choose
OpenAI Compatible APIin the provider menu. - Set the URL to
https://api.spicyapi.com/v1and paste your key. Leave Use CORS Proxy off. - Click Fetch. The model list fills and Chat-Completions API ticks itself; keep it ticked (Spicy API has no legacy completions endpoint).
- In Model Choice pick
spicy-companion-1(Lite preselectsspicy-animate-1, a video model) and click Ok. - Type in the box at the bottom and send. The footer reads "Last request served by Custom Endpoint using spicy-companion-1".
Which model to pick
| Model | Best for | Context (tokens) | Per 1M tokens (prompt / completion) | 100 messages at 4k / 16k context |
|---|---|---|---|---|
spicy-companion-1 | Role-play and companions: holds a persona, writes explicit adult scenes, group scenes. | 131,072 | $1.00 / $2.80 | $0.47 / $1.71 |
spicy-companion-1-flash | The same role-play tuning for about a ninth of the cost of a long chat. Fast. | 43,000 tested | $0.10 / $0.80 | $0.06 / $0.19 |
spicy-chat-1 | General assistant. Not tuned for role-play; use a companion model for characters. | 200,000 | $0.80 / $2.40 | $0.38 / $1.38 |
Chat apps resend the whole conversation with every message, so the context size you set drives the cost far more than the length of the replies. The last column assumes a 250 token reply at 4k context and 400 at 16k. Every response carries its exact cost_usd.
Recommended settings
| Setting | Value |
|---|---|
| Token streaming (Settings) | On (SSE), the default. Replies arrive as they are written. |
| Max Output | 300 to 600 tokens; Lite's default is 768. |
| Context Size | Lite's default 6,144 is cheap; raise it toward 16,384 for longer memory. |
If something goes wrong
| Status | Meaning | Fix |
|---|---|---|
| 400 | Wrong model id, or an option the chat models do not take (tools on spicy-companion-1-flash, images, more than 8,192 max tokens) | Pick spicy-companion-1, spicy-companion-1-flash or spicy-chat-1; the error names the chat models |
| 401 | Key missing, mistyped or revoked | Paste the key again, or make a new one |
| 402 | Balance empty, or the monthly spend limit reached | Top up, or raise the limit on the API keys page. Nothing is charged |
| 422 | Blocked by moderation | Nothing is charged. See what is allowed below |
| 429 | Too many requests in a minute | Wait a minute; the app can retry |
- Lite sends the conversation without a system message unless you set a memory or character; the model then introduces itself as Spicy Companion 1.
- The key is kept in this browser's storage for lite.koboldai.net.
What is allowed
Explicit role-play between adults is allowed, including non-consent between fictional adult characters (force, sleep, intoxication, framed as consensual non-consent or not) and other dark themes. Every message is screened before it reaches the model: minors in any form, real people (with or without consent), incest, real-world harm and the rest of the prohibited list are refused with a 422, before any charge. The screen reads the whole conversation, character card included, so a card that describes a minor or a real person is refused however the latest message is worded. Refusals in the most serious categories count as strikes. Send each user's id as user: the strikes then land on that user, who is refused after ten, and your account is only flagged. Without it they land on your account, flagged after three and suspended after ten. Drop a refused message from the history you send back, or the next message is refused too. A wrong refusal can be reported: POST its moderation_id to /v1/moderations/appeals. Full policy: /acceptable-use
Tested
KoboldAI Lite v347 (lite.koboldai.net) on 30 Sep 2026: model fetch from the browser, OpenAI Compatible API in Chat-Completions mode, and three streamed turns with a sandbox key.
Other apps: /docs/integrations. Questions: contact@spicyapi.com.