Fit guides

Role-play chat sites on Spicy API: does it fit?

The brief: A JanitorAI-style role-play site: user-made character cards, lorebooks, long contexts, streaming, tens of requests a second. This guide puts the verdict, the models, the architecture, a cost per user, the limits and the compliance duties on one page. Prices come from the same model registry the API bills from.

Verdict

Fits with caveats. Both companion models are OpenAI-compatible with streaming and up to 4 replies per call (n), long contexts are cheap because a stable prompt prefix bills at a fifth of the input price, and chat may depict non-consent between fictional adult characters. The limits: user cards that involve minors (including ageless or young-looking characters), real people, incest or choking are declined in chat too, the screen reads the whole conversation so one such card declines every turn, and 50 requests a second (3,000 a minute) is above the starting 600 a minute: limits rise automatically for paying accounts that keep reaching them, up to 6,000, or ask before launch.

  • Content the chat screen always declines: minors in any form (a stated adult age does not save a character written as a child), real identifiable people, incest, choking or strangulation, sexual gore and real-world harm. Fictional non-consent between adults (force, sleep, intoxication, framed as consensual non-consent or not) is allowed in chat only, not in image or video prompts.
  • Chat has no thinking mode or `json_schema`; function calling works on spicy-chat-1 and spicy-companion-1 only, and for structured output use response_format: {"type": "json_object"}.
  • Sampling fields outside the forwarded list are dropped silently, for example min_p. Forwarded: temperature, top_p, top_k, presence_penalty, frequency_penalty, repetition_penalty, seed, stop, logit_bias.
  • No auto top-up or top-up API. Top-ups are paid in the browser. Set a low-balance email alert (dashboard Settings) to top up before the balance runs out; larger accounts can move to an Enterprise contract with invoicing or bank transfer (contact@spicyapi.com).

Which models to use

ModelUse it forPriceReference
spicy-companion-1-flashDefault for a free or high-volume tier: the same role-play tuning for a fraction of the price, tested to 144,000 tokens of context. No tools.$0.1 per 1M prompt tokens ($0.02 when served from cache) and $0.8 per 1M completion tokens (minimum $0.0005 per request)Create Chat Completion
spicy-companion-1Paid tier: stronger writing, function calling, 131,072 tokens of context.$1 per 1M prompt tokens ($0.2 when served from cache) and $2.8 per 1M completion tokens (minimum $0.001 per request)Create Chat Completion
spicy-chat-1General assistant with about 200,000 tokens; not tuned for role-play and may decline explicit scenes. Use it for utility calls such as summaries.$0.8 per 1M prompt tokens ($0.16 when served from cache) and $2.4 per 1M completion tokens (minimum $0.001 per request)Create Chat Completion
spicy-embed-1Lorebook and memory retrieval: embed entries once, pick the relevant ones per turn.$0.14 per 1M tokensCreate Embeddings
spicy-image-1Optional scene images from a prompt you write (never the chat text word for word).$0.06 per imageCreate Image

Architecture

  • Your server makes the call. It assembles the prompt from the card, lorebook and history, calls POST /v1/chat/completions with stream: true and relays the tokens. The key never reaches the browser.
  • Order the prompt for the cache: card and static lorebook first and unchanged between turns, then the history, then triggered lorebook entries and the latest message. Anything that changes near the start breaks the cached prefix and bills the whole prompt at the full input price.
  • Long role-play chats cost less than they look: the provider caches a repeated prompt prefix (1,024 tokens or more) automatically, and cached tokens bill at a fifth of the input price, so keep the character card and lorebook at the start and stable. One turn with 90% of the prompt cached and a 300-token reply: spicy-companion-1-flash $0.0007 at 16k tokens of context, $0.0019 at 60k, $0.0036 at 120k (tested to 144,000); spicy-companion-1 $0.0053, $0.0176 and $0.0344.
  • Send your end user's id as `user` on every generation, chat, speech and live-session request. Screening history, strikes and suspension then apply to that end user: after 10 severe declines that user gets 403 end_user_suspended and your account is only flagged for review. Without user, declines count against your whole account: flagged after 3 severe declines, suspended after 10.
  • Drop a declined chat message from the history you resend. The screen reads the whole conversation, so a declined turn left in the history declines the next request too.
  • Screen cards before they go public. Check a Prompt runs a card or greeting through the same screen without generating, so a card that will be declined on every turn is caught at upload. It never records a strike.
  • Swipes: n up to 4 returns alternative replies in one call.
  • Wrong declines: every 422 from screening carries error.moderation_id; send it to POST /v1/moderations/appeals (free) and we review it.
  • Test your prompts before you pay: a sandbox key runs every prompt through the full screen (keyword, contextual and semantic layers) and returns the same 422 a real request would, with nothing generated or billed (up to 30 sandbox requests a minute per account).

Cost per active user per month

Assumptions: An active user sends 100 messages a day for 30 days (3,000 turns). Each turn sends about 32,000 tokens (card, lorebook, history) with 90% served from the cache because the start of the prompt stays the same, and gets a 400-token reply. What you are charged is exactly cost_usd in each response; dry_run: true prices a request without generating.

Free tier on Flash

ItemModelPer monthEachCost
Chat turns at 32k context, 400-token repliesspicy-companion-1-flash3,000$0.00122$3.65
Total per active user per month$3.65

Paid tier on Companion 1

ItemModelPer monthEachCost
Chat turns at 32k context, 400-token repliesspicy-companion-13,000$0.0101$30.24
Total per active user per month$30.24

Heavy user on Flash at 120k context

ItemModelPer monthEachCost
Chat turns at 120k context, 400-token repliesspicy-companion-1-flash3,000$0.00368$11.04
Total per active user per month$11.04

Limits

  • Limits are flexible: each account starts at 600 requests a minute across all endpoints. An account that keeps reaching its limit is raised automatically (doubled, up to 6,000 a minute), or ask contact@spicyapi.com for more at once. Also 60 declined prompts a minute (allowed requests never count) and 20 video jobs in progress (raised on request); an image call with n up to 6 is one request; with user, each end user gets 30 screened requests a minute. Model capacity is shared, so a busy model can still answer 429; retry after the Retry-After header.
  • Context: 131,072 tokens on spicy-companion-1, tested to 144,000 on spicy-companion-1-flash. Trim the history on your side so the card plus history plus max_tokens fits.
  • `max_tokens` up to 8,192 (default 4,096). The balance must cover the worst case (prompt plus every completion at max_tokens) before the call, so set it to what you need.
  • Declines are capped per account: at 60 declined prompts in a minute, screened requests answer 429 until the minute is over. Send user (each end user then has 30 screened requests a minute) and slow down users who keep getting declined.

Compliance checklist

Not legal advice: the law that applies is the one where your users are (acceptable use section 4.1), and we do not prescribe a method.

  • Age checks, examples of what the main laws accept. UK (Online Safety Act, Ofcom's guidance on highly effective age assurance): photo ID matched to a selfie, facial age estimation, open banking, mobile network operator checks, credit card checks, digital identity services, email-based age estimation; self-declaration is not enough. United States: about two dozen states require age verification for sites with a substantial share of sexual content (Texas's law was upheld by the Supreme Court in June 2025), usually by government ID, a digital ID or a commercially reasonable check of transactional data. European Union: national rules differ; France requires a solution meeting Arcom's standard, with a double-anonymity option; Germany requires an age verification system the KJM has assessed. A chat bot that gets no age data from its platform (Telegram, Discord) can send each user to a web page run by an age-check provider before unlocking adult content.
  • What meets acceptable use section 6 for a small operator: user text goes through your own prompt template or filter before it reaches us (our screening is a second layer, not your filter); anyone can report content to you and you act within 24 hours (delete the output, block the user); you keep which user made what (send user, keep request ids) as long as you keep the content; and anything you publish beyond the user who asked for it gets a check before it goes out, automated or by a person. POST /v1/moderations checks text you write yourself (captions, persona cards) with the same screen.
  • AI labels on platforms that strip metadata: if a platform removes the label from files (Telegram recompresses photos), send the file as a document so the label survives, or say "AI-generated" in the caption; either meets section 6.3.
  • Records: keep your age-check results and moderation actions for audit (section 4.2); DELETE /v1/end-users/{user} handles a user's erasure request on our side.
  • Copyrighted characters: acceptable use 2.4 bans copyrighted characters without authorization; that applies to user-made cards of existing franchise characters.
  • Companion laws: if your site offers ongoing companions, safety_mode: "companion" adds the AI notice New York and California ask for and crisis lines on turns flagged x-spicyapi-safety: self_harm.
  • Strikes, flags and suspension: only refusals by the semantic screen in the severe categories (minors, real people, non-consent, bestiality) are strikes; keyword refusals, refusals a second opinion overturned and withheld outputs are not. Strikes do not expire. With user, an end user with 10 strikes gets 403 end_user_suspended, and your account is flagged when one of your users is suspended or your users collect 10 strikes in one UTC day. Without user, the account is flagged at 3 and suspended at 10. Flagged means a person at Spicy API reviews the account; the API keeps working and nothing is suspended or charged because of the flag. POST /v1/moderations never records a strike. Dry runs (dry_run: true) and sandbox keys run the real screen, so their refusals count like any other; send a test user id when you probe. Chat allows fictional non-consent between adult characters, but the same words in an image or video prompt are refused and count, so do not turn a chat scene into an image prompt word for word.
  • Acceptable use sections that matter most here: 2.1 (minors, fictional and animated included, whatever age is stated), 2.2 (the chat allowance for fictional non-consent and its limits), 2.3 (incest, choking, gore, also in chat), 2.4 (copyrighted characters, celebrities), 4 (age checks), 6 (your own filtering, reports within 24 hours). Full policy: /acceptable-use. How screening, strikes and flags work: /moderation.

Getting started

  1. Sign in at /auth and accept the terms; the Default key is at /dashboard/api-keys.
  2. Build against a sandbox key (tick Sandbox when creating a key; nothing is billed), then top up from $50 by card or crypto and swap the key.
  3. Set a low-balance email alert and a monthly spend limit in the dashboard.

One streamed turn with the stable prefix first and a declined turn dropped (TypeScript, OpenAI SDK):

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.spicyapi.com/v1", apiKey: process.env.SPICYAPI_KEY });

export async function reply(userId: string, card: string, history: { role: "user" | "assistant"; content: string }[], send: (t: string) => void) {
  try {
    const stream = await client.chat.completions.create({
      model: "spicy-companion-1-flash",
      messages: [{ role: "system", content: card }, ...history], // card first and unchanged: cached
      stream: true,
      max_tokens: 500,
      temperature: 0.9,
      user: userId,
    });
    for await (const chunk of stream) send(chunk.choices[0]?.delta?.content ?? "");
  } catch (err) {
    if (err instanceof OpenAI.APIError && err.status === 422) {
      history.pop(); // declined: never resend it, or the next turn is declined too
      throw err; // err.error.moderation_id can go to /v1/moderations/appeals
    }
    throw err;
  }
}

Other business types: /docs/guides. Everything in one file: /llms-full.txt. Questions: contact@spicyapi.com.