Documentation · v1

SpicyAPI in five minutes

One REST API for uncensored images, video, voice and chat: video edits, expressive speech, live voice calls, transcription, embeddings and role-play companions included. No GPUs to rent, no queues to run: you send JSON, you get media back. Every request is billed per generation from your account balance, with no subscription. The base URL for all endpoints is https://api.spicyapi.com.

Getting started

curl https://api.spicyapi.com/v1/images/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-image-1",
    "prompt": "a woman on a beach at golden hour, photorealistic",
    "size": "832*1216"
  }'

Image requests are synchronous and return CDN URLs directly. Video requests return a task; poll GET /v1/videos/tasks/{id} or register a webhook. Either way you receive the output URL and the exact amount charged.

Runnable curl, Python and Node examples live in the porn-api repository on GitHub.

Models

ModelTypePriceNotes
spicy-image-1-proImage$0.09 / imageBest quality. Highest-quality NSFW text-to-image generation with stronger prompt adherence.
spicy-image-1Image$0.06 / imageCheapest. Fast, realistic NSFW text-to-image generation.
spicy-image-action-1Image$0.15 / imageActions. Puts the woman from your image into a ready-made explicit scene (GET /v1/images/actions): her face, hair and skin tone on the scene's act, pose and room.
spicy-image-edit-1Image edit$0.15 / imageEditing. Prompt-driven image editing with up to 3 reference images (outfit, pose, scene changes).
spicy-motion-3Video (i2v / t2v)from $0.10 / secBest quality. Newest-generation video with native dialogue, music and sound effects.
spicy-motion-3-fastVideo (i2v / t2v)from $0.14 / secFastest. Spicy Motion 3 on the accelerated backbone: same inputs and quality tier, noticeably faster turnaround.
spicy-cinema-1-imageVideo (i2v)from $0.14 / secCinematic. Spicy Cinema 1 from a first frame: animates your image with native audio, 3 to 15 seconds at 480p to 1080p.
spicy-motion-2Video (i2v)from $0.20 / secLip-sync. Latest-generation image-to-video with native audio, up to 1080p and 15s.
spicy-character-video-1Video (i2v / t2v)from $0.20 / secSame character. Reference-to-video: keeps a saved character's identity across a new clip from a text prompt, no first frame needed (pass character).
spicy-cinema-1-characterVideo (references)from $0.14 / secGroup scenes. Spicy Cinema 1 from references: keeps up to 9 people, outfits or props consistent across a new clip (pass character or reference images and name them Image 1, Image 2 in the prompt).
spicy-cinema-1Video (t2v)from $0.14 / secText to video. A second video engine for text-to-video: cinematic camera work with synced native audio, 3 to 15 seconds at 480p to 1080p and nine aspect ratios, from 480p at the lowest per-second price.
spicy-video-1Video (t2v)from $0.20 / secText only. Text-to-video generation, no source image required.
spicy-motion-draft-1Video (i2v)from $0.05 / secCheap draft. Silent draft tier for image-to-video at a quarter of the full price: check motion, framing and timing cheaply, then render the keeper on Spicy Motion 2 or 3.
spicy-motion-1Video (i2v)from $0.20 / secLegacy. NSFW-tuned image-to-video with audio.
spicy-video-edit-1Video editfrom $0.20 / secEdit clips. Edit a clip you generated with a prompt: swap the outfit from up to 4 reference images, restyle the scene, change lighting or props.
spicy-cinema-1-editVideo editfrom $0.28 / secLong clips. Edit longer clips: 3 to 30 second inputs with up to 5 outfit, prop or style references.
spicy-animate-1Motion transferfrom $0.24 / secMotion transfer. Motion transfer: the person in your image performs the moves and expressions of a motion clip, such as a dance.
spicy-companion-1Chat$1.00 in / $2.80 out per 1M tokensBest roleplay. Role-play chat model for AI companions: holds a persona from the system prompt, writes explicit adult scenes, handles group chats and returns up to 4 reply options with n.
spicy-companion-1-flashChat$0.10 in / $0.80 out per 1M tokensCheapest. The fast, low-cost role-play model: same persona and explicit-scene handling as Spicy Companion 1, priced for high-volume companion and NPC traffic.
spicy-chat-1Chat$0.80 in / $2.40 out per 1M tokensGeneral. Uncensored roleplay-capable chat model.
spicy-voice-2Speech$0.40 / 10K charsBest quality. Expressive speech with inline tags: switch emotion with [whispers], [excited] or [crying] and add sounds like [giggles], [gasp] or [sighing] mid-sentence.
spicy-voice-2-flashSpeech$0.30 / 10K charsFastest. The faster, cheaper expressive voice model: the same inline emotion and sound tags as Spicy Voice 2, 9 preset voices including British and American English, streaming with stream: true.
spicy-voice-1-expressiveSpeech$0.23 / 10K charsEmotion. Text-to-speech steered by instructions: describe emotion, pace, pitch and delivery in plain English or Chinese.
spicy-voice-1-customSpeech$0.23 / 10K charsYour voice. Text-to-speech in your own voices: design one from a description, or clone one from a screened clip of a speaker who consented (POST /v1/voices).
spicy-voice-1Speech$0.20 / 10K charsCheapest. Text-to-speech with 45 preset voices in 10 languages plus Chinese dialects.
spicy-live-1Live calls$0.46 to $3.74 per 1M tokensLive calls. Live voice calls with a companion: the caller talks, the persona answers in a natural voice with interruptions handled, over one WebSocket.
spicy-transcribe-1Transcription$0.0042 / minTranscription. Speech-to-text for voice notes and clips: wav, mp3, m4a, ogg and more in 11 languages, with the spoken language detected.
spicy-embed-1Embeddings$0.14 per 1M tokensText. Text embeddings for companion memory, prompt search and recommendations.
spicy-embed-vision-1Embeddings$0.18 text, $0.06 image per 1M tokensImages. Images and text in one 768-dimension space: "more like this" search, tagging and near-duplicate detection over your own generations.

Call GET /v1/models for the live catalog including per-model size, duration and resolution limits.

Authentication

Authenticate with a bearer token. Create keys in the dashboard. A key is shown once at creation and stored only as a hash, so save it somewhere safe.

Authorization: Bearer sk-spicy-xxxxxxxxxxxxxxxxxxxx

Never ship a key in client-side code. Check your balance any time with GET /v1/account. Requests without a valid key return 401.

Image generation

POST /v1/images/generations is synchronous. It returns image URLs directly, typically in 10 to 30 seconds.

curl https://api.spicyapi.com/v1/images/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-image-1",
    "prompt": "your prompt",
    "negative_prompt": "blurry, low quality",
    "size": "832*1216",
    "n": 2
  }'

{
  "id": "sj_2f1c...",
  "object": "image.generation",
  "model": "spicy-image-1",
  "data": [{ "url": "https://cdn.spicyapi.com/outputs/..." }],
  "cost_usd": 0.05
}

On the Spicy Image 1 models n is 1 to 6, billed per image, and size is one of eleven width*height shapes, all in the same price tier: squares (1024*1024, 1280*1280, 1440*1440), portrait and landscape (832*1216, 1024*1536, 1080*1920, 1152*2048 and their mirrors). Prompts are capped at 2000 characters. enhance_prompt: true lets the model add lighting, camera and detail cues before rendering (enhance_mode: "agent" plans the shot first). The same n, size and enhance_prompt work on /v1/images/edits. Images are screened after generation as well as before; anything withheld is refunded and reported as withheld.

Every returned url is permanently hosted on the SpicyAPI CDN and registered to your account. These URLs are what you pass later as inputs to image editing and video generation.

Treat the URL as your master copy and provenance token, not as end-user hosting: download the file once and serve it to your own users from your own storage or CDN. Sustained hotlinking of output URLs into consumer-facing apps falls outside fair use and may be rate-limited.

Serving a copy does not affect reuse: keep the original URL in your database, and it remains valid forever as an input to /v1/images/edits and /v1/videos/generations. Your users view your copy, your backend passes the SpicyAPI URL.

Image editing

POST /v1/images/edits edits or recomposes an existing image from a prompt. Pass up to three reference images. Each must be an output previously generated by your own account: use the url values returned by /v1/images/generations or /v1/images/edits. External image URLs are rejected with a 400; see Input images for why, and for the integration pattern.

curl https://api.spicyapi.com/v1/images/edits \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-image-edit-1",
    "prompt": "change the outfit to a red dress",
    "image_urls": ["https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png"]
  }'

Characters: the same person, againNew

Save a person you generated as a character and pass its id to any generation call to get the same face, hair and body in a new scene, a new outfit, or a video. Characters are built only from images your account generated, so there is nothing to upload and no photograph of a real person can ever become one.

curl https://api.spicyapi.com/v1/characters \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "name": "Mara", "image_urls": ["https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png"] }'
# → { "id": "chr_9f2a…", "sheet_url": "…", "description": "Adult woman in her late 20s, …" }

curl https://api.spicyapi.com/v1/images/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "spicy-image-1", "prompt": "on a hotel balcony at night, red silk robe", "character": "chr_9f2a…" }'

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "spicy-character-video-1", "prompt": "walks to the window and looks back over her shoulder", "character": "chr_9f2a…", "duration": 5 }'

Creation renders one neutral reference sheet (billed as one spicy-image-edit-1 image) and writes a short descriptor. On every later call the sheet goes in as Image 1 and the descriptor is appended to your prompt. Images with a character are billed at the edit rate. For video, use spicy-character-video-1 (reference-to-video, no first frame needed), spicy-motion-3 or spicy-motion-3-fast (which also take a character with an action). Models that need a first frame take an image generated with the character. Up to 3 source images per character and 50 characters per account. See Create Character.

Moving to SpicyAPI with characters you already generated on another platform? Those can be brought over with character import, which we enable per account on request; see Importing existing characters.

Input images: no uploads, by design

SpicyAPI does not accept uploaded or external images anywhere. The only images that can be edited or animated are images that were generated through the API by your own account. This is a safety and legal-compliance requirement, not a technical limitation: because every video frame and edit source traces back to a moderated, machine-generated origin, neither you nor we can be handed real-world photographs of real people, which keeps the API (and every product built on it) compliant with international law on non-consensual and abusive imagery.

Enforcement is server-side: each generated output URL is registered to the account that created it, and input URLs are checked against that registry before anything is charged. A URL that was not minted for your account (another customer's output, a guessed CDN path, or a re-hosted copy) is rejected with a 400.

Integrating this into your product. The pattern is a gallery: your users first generate images, browse what they've made, then pick one to edit or animate.

1. When a generation completes, store the returned urlin your own database against the end-user who made it. Your API account's library is shared across your whole app, so this per-user mapping is yours to keep.

2. To build the gallery, use your stored URLs, or fetch the account-wide library from GET /v1/images, newest first:

curl "https://api.spicyapi.com/v1/images?limit=50" \
  -H "Authorization: Bearer $SPICYAPI_KEY"

{
  "object": "list",
  "data": [
    { "url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
      "source": "image", "created": 1755856800 }
  ]
}

3. When the user picks an image, pass its exact URL straight into /v1/images/edits or /v1/videos/generations:

// e.g. an Express handler behind your own user auth
app.post("/animate", async (req, res) => {
  // look the image up in YOUR db so users can only
  // animate images they generated themselves
  const image = await db.images.findOwn(req.user.id, req.body.imageId);

  const task = await fetch("https://api.spicyapi.com/v1/videos/generations", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.SPICYAPI_KEY}`,
               "Content-Type": "application/json" },
    body: JSON.stringify({
      model: "spicy-motion-2",
      prompt: req.body.prompt,
      image_url: image.url,   // exact URL from /v1/images/generations
      resolution: "720P",
      duration: 5,
    }),
  }).then(r => r.json());

  res.json({ taskId: task.id });
});

Pass the URL byte-for-byte as it was returned. Adding query parameters or changing the encoding will fail the provenance check.

Importing existing characters. You cannot upload your own images, whether to animate them or to edit them. The one exception is for platforms moving to SpicyAPI with AI-generated characters they have already created elsewhere: if you have existing characters you need to bring over, email contact@spicyapi.com from your account email with your site and roughly how many characters you need to import. We can enable character import on your account for a limited window and a set number of characters.

While it is enabled, POST /v1/characters/import takes 1 to 3 images of one character and consent: true, your attestation that they show a fictional, AI-generated adult your platform created, not a real person. Imports are governed by Schedule 1 of the Terms of Service, which your account accepts in the dashboard before the first import, and the feature is for your own library: never expose it to your end users. Every image is screened before anything is charged, files with camera metadata are refused as photographs, and some imports wait for a person to review them. The result is an ordinary chr_... character you use everywhere a character is accepted, and it keeps working after the window closes. The uploaded files themselves never become inputs. See Import Character.

curl https://api.spicyapi.com/v1/characters/import \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -F name=Mara -F consent=true \
  -F file=@mara-front.png -F file=@mara-side.png
# → { "id": "chr_4b7e…", "imported": true, "status": "active", … }

Video generation

POST /v1/videos/generations is asynchronous. It returns a task immediately; poll it or receive a webhook. Video generation typically takes 1 to 5 minutes.

image_url must be an image generated by your own account: first create the frame with /v1/images/generations (or /v1/images/edits), then pass the exact url it returned. Uploaded or external images are rejected with a 400; see Input images.

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-2",
    "prompt": "she turns towards the camera and smiles",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
    "resolution": "720P",
    "duration": 5
  }'

{ "id": "sj_9a2b...", "status": "queued", "cost_usd": 0.125 }

Poll with GET /v1/videos/tasks/{id}. Status moves through queued, processing, then succeeded or failed. Failed generations are refunded automatically.

{
  "id": "sj_9a2b...",
  "status": "succeeded",
  "output": { "video_url": "https://cdn.spicyapi.com/outputs/..." },
  "cost_usd": 0.125
}

As with images, download the finished video and serve it to your end users from your own infrastructure. The CDN URL is your master copy, not a streaming host for your traffic.

Text-to-video models take aspect_ratio: 16:9, 9:16, 1:1, 4:3, 3:4 on spicy-video-1 (default 16:9), those plus 21:9 on spicy-motion-3 and spicy-motion-3-fast, where leaving it out lets the model pick a shape for the prompt. With an input image or clip the output follows the source. Prompt caps are per model (limits.maxPromptChars on GET /v1/models); Spicy Motion 3 has no negative prompt and ignores one.

Last frame and references

last_frame_url fixes where the clip ends: the model travels from image_url to it. Works on spicy-motion-2 and Spicy Motion 3. reference_image_urls (up to 10) guide identity and look without fixing the opening shot, on spicy-character-video-1 and Spicy Motion 3; it cannot be combined with image_url. Both take images your account generated.

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-3",
    "prompt": "she walks from the doorway to the window and turns",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
    "last_frame_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-1.png",
    "duration": 8
  }'

Extend a clip

video_url takes the output.video_url of one of your own finished tasks. On spicy-motion-2 it continues the clip in place of image_url: the input must be 2 to 10 seconds, duration is the whole output including the input, and you are billed for the whole duration. On Spicy Motion 3 it is a reference clip to extend or edit: up to 15 seconds in, and input plus duration stays within 30.

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-2",
    "prompt": "she keeps walking and looks back over her shoulder",
    "video_url": "https://cdn.spicyapi.com/outputs/…/sj_9a2b….mp4",
    "duration": 10
  }'

Smart duration

On spicy-motion-3 and spicy-motion-3-fast, duration: -1lets the model choose the length. You are debited for the maximum (30 seconds, less any input clip) when the task is accepted; when it succeeds the difference is refunded and the task's cost_usd settles to the rendered length, which usage.output_seconds reports.

{
  "id": "sj_9a2b...",
  "status": "succeeded",
  "output": { "video_url": "https://cdn.spicyapi.com/outputs/..." },
  "usage": { "output_seconds": 12, "resolution": "720P" },
  "cost_usd": 2.4
}

Audio inputs

Audio is the one input you can bring from outside. Upload a WAV or MP3 of 2 to 30 seconds (15 MB max) with POST /v1/audio, as a multipart file or a JSON url we fetch. The clip is transcribed and the transcript screened like a prompt; the flat $0.01 covers that. Then pass the returned url as audio_url on any video model except spicy-motion-1 and spicy-character-video-1 for lip-sync and audio-driven motion. Speech from POST /v1/audio/speech is a clip too, usable the same way (see Voices and speech). Spicy Motion 3 takes up to five clips in audio_urls, 15 seconds in total. GET /v1/audio lists your clips and DELETE /v1/audio/{id} forgets one.

curl https://api.spicyapi.com/v1/audio \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -F "file=@line.wav"
# → { "id": "au_3c9f…", "url": "https://cdn.spicyapi.com/audio/…/au_3c9f….wav", "duration_s": 6.4, "transcript": "…", "cost_usd": 0.01 }

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-2",
    "prompt": "she says the line looking straight into the camera",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
    "audio_url": "https://cdn.spicyapi.com/audio/…/au_3c9f….wav",
    "duration": 7
  }'

Driving audio vs reference audio

The same audio_url does two different things. On spicy-motion-2, and spicy-video-1 it is driving audio: the clip becomes the soundtrack and the mouth follows it word for word, alongside a first frame where the model takes one. On spicy-motion-3 and spicy-motion-3-fast it is reference audio: the model generates its own sound and uses the clip for voice, tone and beat while the words come from the prompt; it only goes with a text prompt or references, never with a first frame (there a first frame pairs only with last_frame_url). Send audio_mode (driving or reference) to assert which one you expect; a mismatch is a 400 naming a model that offers it.

What can be combined

Per model, from inputs on GET /v1/models, whose combinations sentences are quoted back in any 400 for an invalid mix. References and a first frame are exclusive everywhere.

ModelFirst frameLast frameReferencesClipAudioAspect ratioSmart duration
spicy-motion-3optionalyesup to 10referencereference, 5 clips, no first frameyesyes
spicy-motion-3-fastoptionalyesup to 10referencereference, 5 clips, no first frameyesyes
spicy-cinema-1-imagerequirednononononenono
spicy-motion-2requiredyesnocontinuedrivingnono
spicy-character-video-1optionalnoup to 10nononenono
spicy-cinema-1-characternonenoup to 9nononeyesno
spicy-cinema-1nonenononononeyesno
spicy-video-1nonenononodriving, no first frameyesno
spicy-motion-draft-1requirednononononenono
spicy-motion-1requirednononononenono

Draft tier

spicy-motion-draft-1 is silent image-to-video at $0.05/s (720P), a quarter of spicy-motion-2's $0.20/s. Use it to check motion, framing and timing cheaply, then render the keeper with the same first frame and prompt on Spicy Motion 2 or 3. 2 to 15 seconds at 720P or 1080P; GET /v1/models marks it silent: true.

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-draft-1",
    "prompt": "she turns towards the camera and smiles",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
    "duration": 5
  }'

The Cinema engine

A second video engine, with cinematic camera work and synced native audio on every clip: spicy-cinema-1 from text alone (nine aspect ratios, including 4:5 and 9:21), spicy-cinema-1-image from a first frame, and spicy-cinema-1-character from a character or up to 9 reference images, named Image 1, Image 2 in the prompt, which it keeps consistent. 3 to 15 seconds at 480P ($0.14/s, the lowest per-second price), 720P ($0.28/s) or 1080P ($0.36/s). No negative prompt, no enhance_prompt and no audio inputs.

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-cinema-1-character",
    "prompt": "Image 1 and Image 2 share a slow dance in a candlelit bar, the camera circling them",
    "reference_image_urls": [
      "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
      "https://cdn.spicyapi.com/outputs/…/sj_7a3d…-0.png"
    ],
    "resolution": "480P",
    "aspect_ratio": "9:16",
    "duration": 8
  }'

60 fps

Every model renders at 30 fps. Send fps: 60 on any video model and the finished clip is frame-interpolated to 60 fps for smoother motion, audio kept. It adds 20% of the video's own price (input and action clip seconds included) and about a minute to the task: a 10 second 720P clip on spicy-motion-3 is $2.00 at 30 fps, $2.40 at 60 fps, and an 8 second 720P action video on spicy-motion-3 (8 output plus 4 input seconds) is $2.40 at 30 fps, $2.88 at 60 fps. If the 60 fps step fails, the 30 fps clip is delivered and the fee refunded; the finished task's usage.fps says 60 or 30.

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "spicy-motion-3", "prompt": "she dances in slow circles", "duration": 10, "fps": 60 }'

Video edits and motion transferNew

POST /v1/videos/edits changes a clip you already generated and returns a task, polled and webhooked like any video. spicy-video-edit-1 ($0.20/s at 720P) swaps an outfit, restyles the scene or changes lighting and props in a 2 to 10 second clip, with up to 4 reference images and an optional aspect_ratio. spicy-cinema-1-edit ($0.28/s at 720P) takes 3 to 30 second clips and up to 5 references and returns the first 15 seconds edited. Both bill the input clip's seconds plus the output's, debited up front and settled when the task finishes. keep_audio: true keeps the original soundtrack.

curl https://api.spicyapi.com/v1/videos/edits \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-video-edit-1",
    "video_url": "https://cdn.spicyapi.com/outputs/…/sj_9a2b….mp4",
    "prompt": "she wears the red dress from Image 1, same room and lighting",
    "reference_image_urls": ["https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png"],
    "keep_audio": true
  }'

Motion transfer

spicy-animate-1 makes the person in one of your images perform the moves and expressions of a motion clip, such as a dance. Pick a clip from the library with motion (mo_slow_sway, mo_hip_dance, mo_hair_turn, mo_catwalk, mo_wave; GET /v1/videos/motions lists them with previews, no key needed), or pass one of your own 2 to 30 second clips as video_url. The output is as long as the clip, billed per output second: $0.24/s standard, $0.36/s with quality: "pro". A full-body image in the clip's 9:16 framing works best. This model refuses explicit or nude images and clips (the task fails with error_code: "blocked" and is refunded), so use a dressed performer.

curl https://api.spicyapi.com/v1/videos/edits \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-animate-1",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-1.png",
    "motion": "mo_hip_dance"
  }'
# → { "id": "sj_7e1d…", "object": "task", "status": "queued", "poll_url": "/v1/videos/tasks/sj_7e1d…" }

Actions: custom sex actions from one imageNew

An action is a ready-made sex act: a short reference clip that fixes the camera angle, framing, pose and motion. Pass its id as action on POST /v1/videos/generations with spicy-motion-3 or spicy-motion-3-fast, plus an image of the woman, and she performs that act second for second, with sound. There are 64 of them; GET /v1/videos/actions lists them with preview clips and posters, no key needed.

POV missionary
POV missionarypov-missionarysex
POV missionary, legs wide
POV missionary, legs widepov-missionary-legs-widesex
Close-up missionary creampie
Close-up missionary creampiecloseup-missionary-creampiesex
Legs up pounding
Legs up poundinglegs-up-poundingsex
Legs up, side view
Legs up, side viewlegs-up-side-viewsex
POV doggystyle
POV doggystylepov-doggystylesex
POV standing doggystyle
POV standing doggystylepov-standing-doggystylesex
Pounding doggystyle, face close-up
Pounding doggystyle, face close-uppounding-doggystylesex
Front view doggystyle, hair pull
Front view doggystyle, hair pullfront-view-doggystyle-hair-pullsex
Rough doggystyle, hair pull
Rough doggystyle, hair pullrough-doggystyle-hair-pullsex
Cowgirl, face close-up
Cowgirl, face close-upcowgirl-face-closeupsex
Cowgirl, leaning back
Cowgirl, leaning backcowgirl-lean-backsex
Cowgirl, undressing
Cowgirl, undressingcowgirl-undresssex
Reverse cowgirl, bouncing
Reverse cowgirl, bouncingreverse-cowgirl-bouncysex
Reverse cowgirl, ass grab
Reverse cowgirl, ass grabreverse-cowgirl-ass-grabsex
Doggystyle, looking back
Doggystyle, looking backdoggystyle-looking-backsex
Reverse cowgirl, facing camera
Reverse cowgirl, facing camerareverse-cowgirl-facing-camerasex
Reverse cowgirl, rubbing herself
Reverse cowgirl, rubbing herselfreverse-cowgirl-rubbingsex
POV missionary creampie
POV missionary creampiepov-missionary-creampiesex
POV titfuck
POV titfuckpov-titfucksex
Anal doggystyle
Anal doggystyleanal-doggystyleanal
Anal full nelson
Anal full nelsonanal-full-nelsonanal
Ass to mouth
Ass to mouthass-to-mouthanal
Anal, legs pulled back
Anal, legs pulled backanal-legs-backanal
Close-up dildo ride
Close-up dildo ridecloseup-dildo-ridesolo
Pussy spread
Pussy spreadpussy-spreadsolo
Turn around and spread
Turn around and spreadturn-and-spreadsolo
POV blowjob and deepthroat
POV blowjob and deepthroatpov-blowjob-deepthroatoral
Eye contact deepthroat
Eye contact deepthroateye-contact-deepthroatoral
Sloppy deepthroat
Sloppy deepthroatsloppy-deepthroatoral
Face fucking
Face fuckingface-fuckingoral
Face fuck, hand on head
Face fuck, hand on headface-fuck-hand-on-headoral
Arched back blowjob
Arched back blowjobarched-back-blowjoboral
Seductive blowjob
Seductive blowjobseductive-blowjoboral
Side view blowjob
Side view blowjobside-view-blowjoboral
Top-down blowjob, ass view
Top-down blowjob, ass viewtop-down-blowjoboral
Two-handed blowjob
Two-handed blowjobtwo-handed-blowjoboral
Cock worship
Cock worshipcock-worshiporal
POV underside licking
POV underside lickingpov-underside-lickingoral
POV licking balls
POV licking ballspov-licking-ballsoral
POV blowjob, looking up
POV blowjob, looking uppov-blowjob-looking-uporal
Kneeling blowjob, eye contact
Kneeling blowjob, eye contactkneeling-blowjob-eye-contactoral
Face fuck, two angles, facial
Face fuck, two angles, facialface-fuck-facialoral
POV handjob cumshot
POV handjob cumshotpov-handjob-cumshothandjob
Handjob while touching herself
Handjob while touching herselfhandjob-while-touchinghandjob
POV handjob, eye contact
POV handjob, eye contactpov-handjob-eye-contacthandjob
Handjob facial
Handjob facialhandjob-facialcumshot
Handjob finish and swallow
Handjob finish and swallowhandjob-finish-swallowcumshot
Cum in mouth
Cum in mouthcum-in-mouthcumshot
Cumshot on tits
Cumshot on titscumshot-on-titscumshot
POV licking cumshot
POV licking cumshotpov-licking-cumshotcumshot
Pussy to mouth cumshot
Pussy to mouth cumshotpussy-to-mouth-cumshotcumshot
Titjob cumshot
Titjob cumshottitjob-cumshotcumshot
Handjob facial, tongue out
Handjob facial, tongue outhandjob-facial-tongue-outcumshot
Facial, tongue out
Facial, tongue outfacial-tongue-outcumshot
Cum gargling
Cum garglingcum-garglingcumshot
Trans cowgirl
Trans cowgirltrans-cowgirltrans
Trans stroking cumshot
Trans stroking cumshottrans-stroking-cumshottrans
Trans edging cumshot, low angle
Trans edging cumshot, low angletrans-edging-cumshottrans
Trans cum dripping onto her balls
Trans cum dripping onto her ballstrans-cum-driptrans
Trans fucking him from behind
Trans fucking him from behindtrans-fucking-himtrans
Trans fucking him, side view
Trans fucking him, side viewtrans-fucking-him-sidetrans
Trans fucking her doggystyle
Trans fucking her doggystyletrans-fucking-hertrans
Trans lingerie selfie
Trans lingerie selfietrans-lingerie-posetrans
curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-3",
    "action": "pov-blowjob-deepthroat",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
    "prompt": "hotel room at night, warm lamp light, black lace lingerie",
    "resolution": "720P"
  }'

Pricing: the reference clip is the first 4 seconds of the act (5 on two actions); its seconds are billed as input seconds on top of the output seconds, at the model's per-second rate. An 8 second video at 720P on spicy-motion-3 bills 12 seconds ($0.20/s, so $2.40). On spicy-motion-3-fast the same clip costs $3.36.

Tips

  • image_url is who she is, not the first frame: a clear, well-lit image of the woman (face and body) from your library. reference_image_urls or a character work too.
  • prompt is optional and sets the scene: setting, outfit, lighting. The action already supplies the motion, so do not describe it. Without a prompt she is in a plain light bedroom.
  • The output takes the action's shape (3:4, 9:16 or 1:1) unless you send aspect_ratio; any ratio the model supports works.
  • duration defaults to 8; longer clips repeat the same act (action seconds plus duration up to 30). video_url and last_frame_url are not accepted with an action.

Step by step, with prices and examples: Custom sex actions.

Image actions: her, in a ready-made explicit sceneNew

An image action is a ready-made explicit still. Pass its id as action on POST /v1/images/generations with spicy-image-action-1, plus an image of the woman, and her face, hair and skin tone are put into that scene. The still keeps the act, pose, framing, room and her body. There are 68; GET /v1/images/actions lists them, no key needed. Most share their id with a video action.

POV missionary
POV missionarypov-missionary
POV missionary, legs wide
POV missionary, legs widepov-missionary-legs-wide
Close-up missionary creampie
Close-up missionary creampiecloseup-missionary-creampie
Legs up pounding
Legs up poundinglegs-up-pounding
Legs up, side view
Legs up, side viewlegs-up-side-view
POV doggystyle
POV doggystylepov-doggystyle
Pounding doggystyle, face close-up
Pounding doggystyle, face close-uppounding-doggystyle
Front view doggystyle, hair pull
Front view doggystyle, hair pullfront-view-doggystyle-hair-pull
Rough doggystyle, hair pull
Rough doggystyle, hair pullrough-doggystyle-hair-pull
Cowgirl, face close-up
Cowgirl, face close-upcowgirl-face-closeup
Cowgirl, leaning back
Cowgirl, leaning backcowgirl-lean-back
Cowgirl, undressing
Cowgirl, undressingcowgirl-undress
Reverse cowgirl, bouncing
Reverse cowgirl, bouncingreverse-cowgirl-bouncy
Reverse cowgirl, ass grab
Reverse cowgirl, ass grabreverse-cowgirl-ass-grab
Anal doggystyle
Anal doggystyleanal-doggystyle
Anal full nelson
Anal full nelsonanal-full-nelson
Close-up dildo ride
Close-up dildo ridecloseup-dildo-ride
POV blowjob and deepthroat
POV blowjob and deepthroatpov-blowjob-deepthroat
Eye contact deepthroat
Eye contact deepthroateye-contact-deepthroat
Sloppy deepthroat
Sloppy deepthroatsloppy-deepthroat
Face fucking
Face fuckingface-fucking
Face fuck, hand on head
Face fuck, hand on headface-fuck-hand-on-head
Arched back blowjob
Arched back blowjobarched-back-blowjob
Seductive blowjob
Seductive blowjobseductive-blowjob
Side view blowjob
Side view blowjobside-view-blowjob
Top-down blowjob, ass view
Top-down blowjob, ass viewtop-down-blowjob
Two-handed blowjob
Two-handed blowjobtwo-handed-blowjob
Cock worship
Cock worshipcock-worship
POV underside licking
POV underside lickingpov-underside-licking
POV licking balls
POV licking ballspov-licking-balls
POV handjob cumshot
POV handjob cumshotpov-handjob-cumshot
Handjob while touching herself
Handjob while touching herselfhandjob-while-touching
Handjob facial
Handjob facialhandjob-facial
Handjob finish and swallow
Handjob finish and swallowhandjob-finish-swallow
Cum in mouth
Cum in mouthcum-in-mouth
Cumshot on tits
Cumshot on titscumshot-on-tits
POV licking cumshot
POV licking cumshotpov-licking-cumshot
Pussy to mouth cumshot
Pussy to mouth cumshotpussy-to-mouth-cumshot
Kneeling nude on the bed
Kneeling nude on the bedkneeling-nude-on-bed
Facial, tongue out
Facial, tongue outfacial-tongue-out
Lying on her front, feet up
Lying on her front, feet uplying-on-front-feet-up
Topless selfie in bed
Topless selfie in bedtopless-bed-selfie
Oiled up on the beach
Oiled up on the beachoiled-on-the-beach
Touching herself in the shower
Touching herself in the showershower-masturbation
Creampie, lying back
Creampie, lying backcreampie-lying-back
Cat-ear cosplay, licking
Cat-ear cosplay, lickingcat-ears-licking
Doggystyle, looking back
Doggystyle, looking backdoggystyle-looking-back
Reverse cowgirl, facing camera
Reverse cowgirl, facing camerareverse-cowgirl-facing-camera
Reverse cowgirl, rubbing herself
Reverse cowgirl, rubbing herselfreverse-cowgirl-rubbing
POV missionary creampie
POV missionary creampiepov-missionary-creampie
POV titfuck
POV titfuckpov-titfuck
Ass to mouth
Ass to mouthass-to-mouth
Anal, legs pulled back
Anal, legs pulled backanal-legs-back
Pussy spread
Pussy spreadpussy-spread
POV blowjob, looking up
POV blowjob, looking uppov-blowjob-looking-up
Kneeling blowjob, eye contact
Kneeling blowjob, eye contactkneeling-blowjob-eye-contact
Face fuck, two angles, facial
Face fuck, two angles, facialface-fuck-facial
POV handjob, eye contact
POV handjob, eye contactpov-handjob-eye-contact
Titjob cumshot
Titjob cumshottitjob-cumshot
Handjob facial, tongue out
Handjob facial, tongue outhandjob-facial-tongue-out
Cum gargling
Cum garglingcum-gargling
Trans cowgirl
Trans cowgirltrans-cowgirl
Trans stroking cumshot
Trans stroking cumshottrans-stroking-cumshot
Trans edging cumshot, low angle
Trans edging cumshot, low angletrans-edging-cumshot
Trans cum dripping onto her balls
Trans cum dripping onto her ballstrans-cum-drip
Trans fucking him from behind
Trans fucking him from behindtrans-fucking-him
Trans fucking him, side view
Trans fucking him, side viewtrans-fucking-him-side
Trans fucking her doggystyle
Trans fucking her doggystyletrans-fucking-her
curl https://api.spicyapi.com/v1/images/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-image-action-1",
    "action": "pov-blowjob-deepthroat",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png"
  }'

Pricing: $0.15 per image, one to four per call. A swap takes a minute or two.

Tips

  • image_url must be one of your own generated images with a clear face. A character works too.
  • Drawn women: pass style (anime, 3d or cartoon) to put her into the drawn version of the scene; without it the scene is a photo. Each action lists the styles it comes in under styles.
  • No prompt or size: the still sets the scene and the shape. Describing a body type does not change it and can make the model dress her, so it is refused rather than ignored.

Styles: one look across image and videoNew

Pass style on POST /v1/images/generations or POST /v1/videos/generations(every video model) to pick an art style. On images it adds a tuned style prompt and matching negatives; on video it tells the model the whole clip's style, matched to the first frame or reference image. Styles are free: same models, same prices. Omit style for each model's natural look. GET /v1/styles lists them, no key needed.

Photorealistic style
PhotorealisticphotorealisticCamera-real photo look
Studio style
StudiostudioGlossy glamour studio look
Anime style
AnimeanimeModern detailed TV anime
3D style
3D3dAnimated-film CGI look
Cartoon style
Cartooncartoon2D adult cartoon look
curl https://api.spicyapi.com/v1/images/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-image-1-pro",
    "style": "anime",
    "prompt": "a woman with long silver hair in a red kimono on a balcony at dusk",
    "size": "832*1216"
  }'

Then keep her in that style through an action. Without style, the photoreal reference clip pulls a stylised woman back toward realistic; with it she stays anime, 3D or cartoon for the whole act.

curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-3",
    "action": "pov-missionary",
    "style": "anime",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png"
  }'

Voices and speechNew

POST /v1/audio/speech turns up to 600 characters of text into speech (5,000 on Spicy Voice 2). The response carries a url to a WAV on the CDN (24 kHz, 16-bit mono; MP3 on request with Spicy Voice 2) that is already one of your audio clips, so it goes straight into video as audio_url. Speech is billed per 10,000 characters of input (CJK ideographs count 2). The text is screened like a prompt first, and a blocked request costs nothing. language defaults to Auto; English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian can be set explicitly.

ModelVoicesPriceNotes
spicy-voice-145 presets$0.20 / 10K charsPreset voices in 10 languages plus Chinese dialect voices
spicy-voice-1-expressive21 presets$0.23 / 10K charsAdds instructions for emotion, pace, pitch and delivery
spicy-voice-1-customYour vc_... voices$0.23 / 10K charsDesigned voice $0.40, cloned voice $0.02, once each
spicy-voice-22 presets$0.40 / 10K charsInline tags, instructions, MP3 and streaming, up to 5,000 characters
spicy-voice-2-flash9 presets$0.30 / 10K charsInline tags, instructions, MP3 and streaming, up to 5,000 characters

Preset voices

GET /v1/voices lists every preset voice with the models that speak it, then your own voices. Pass the name as voice.

curl https://api.spicyapi.com/v1/audio/speech \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-voice-1",
    "input": "Come here. I have been waiting all evening.",
    "voice": "Cherry"
  }'
# → { "id": "au_7d2e…", "object": "audio.speech", "url": "https://cdn.spicyapi.com/audio/…/au_7d2e….wav", "duration_s": 3.4, "voice": "Cherry", "cost_usd": 0.00086 }

Expressive delivery

On spicy-voice-1-expressive, instructions describe how the line is said (emotion, pace, pitch, delivery) in up to 1,600 characters of plain English or Chinese. They are included in the debit, and the charge settles down to the billed count.

curl https://api.spicyapi.com/v1/audio/speech \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-voice-1-expressive",
    "input": "You kept me waiting. Come closer.",
    "voice": "Serena",
    "instructions": "Low and breathy, slow, teasing, with a smile in the voice."
  }'

Expressive voice with inline tags

spicy-voice-2 (voices Lingxin, Lufeng) and spicy-voice-2-flash (voices Fengyue, Yuanfei, Lingxi, Xiaoxin, Huan, Chuanshu, Mary, Eva, John) take up to 5,000 characters in Auto, English, Chinese, plus instructions, and act out tags written in square brackets. Control tags change the delivery until the next tag: [sad] [amazed] [deep and loud shouting] [trembling] [angry] [excited] [sarcastic] [curious] [like dracula] [bored] [tired] [scornful] [shouting] [asmr] [panicked] [mischievously] [empathetic] [whispers] [reluctantly] [crying] [serious] [very slowly] [very fast]. Sound tags insert a sound: [gasp] [sighing] [clears throat] [giggles] [laughing] [cough] [snorts]. Tags are billed as characters. response_format is wav (default) or mp3. Custom voices stay on spicy-voice-1-custom.

curl https://api.spicyapi.com/v1/audio/speech \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-voice-2",
    "input": "[excited]Hey you, I have been waiting all night.[giggles] Come here.[whispers] Closer.",
    "voice": "Lingxin",
    "response_format": "mp3"
  }'

Streaming

With stream: true on a Spicy Voice 2 model the response body is the audio itself, sent as it is synthesized, with the first audio in about 1.5 seconds: good for voice replies in a chat app. The default is raw pcm (24 kHz, 16-bit, mono, little-endian); ask for wav or mp3 with response_format. The headers carry x-spicyapi-audio-id, x-spicyapi-sample-rate and x-spicyapi-cost-usd; the finished clip is stored under that id afterwards, so it also appears in GET /v1/audio and works as audio_url on video.

const res = await fetch("https://api.spicyapi.com/v1/audio/speech", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.SPICYAPI_KEY}`, "Content-Type": "application/json" },
  body: JSON.stringify({ model: "spicy-voice-2-flash", voice: "Eva", input: "[whispers]Stay on the line with me.", stream: true, response_format: "mp3" }),
});
console.log(res.headers.get("x-spicyapi-audio-id"));
for await (const chunk of res.body) player.write(chunk); // mp3 bytes as they arrive

Design a character voice

POST /v1/voices with a name and a description of vocal qualities: gender, age range, pitch, pace, tone, accent. A voice must be invented, not copied: a name or description that names, references or imitates a real person is refused with code: "real_person". Designing costs $0.40 once and the response includes a preview_url. Speak with it on spicy-voice-1-custom, passing the vc_... id as voice.

curl https://api.spicyapi.com/v1/voices \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Mara",
    "description": "Woman in her thirties, low and warm, slightly husky, unhurried, a playful smile in the voice, neutral American accent."
  }'
# → { "id": "vc_4e1b…", "object": "voice", "kind": "designed", "models": ["spicy-voice-1-custom"], "preview_url": "https://cdn.spicyapi.com/audio/…-preview.wav", "cost_usd": 0.4 }

Clone your own voice

Upload a clean 10 to 20 second sample with POST /v1/audio (transcribed and screened for $0.01), then send its url as audio_url with consent: true. That flag is your attestation that the voice is your own, or that you hold the speaker's written consent to clone it and to generate speech with it; without it the request is a 400. The attestation is stored with the voice and returned as consent_attested_at. Cloning anyone without that consent breaks the acceptable use policy. Cloning costs $0.02 once; DELETE /v1/voices/{id} removes a voice.

curl https://api.spicyapi.com/v1/voices \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "My voice",
    "audio_url": "https://cdn.spicyapi.com/audio/…/au_3c9f….wav",
    "consent": true
  }'

Make a character speak in a video

Generate the line, then hand its url to video generation as audio_url. On spicy-motion-2 it is driving audio: the clip becomes the soundtrack and the mouth follows the words, from a first frame. Audio for video must be 2 to 30 seconds long, so split long text across requests.

# 1. The line, in your character's voice
curl https://api.spicyapi.com/v1/audio/speech \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-voice-1-custom",
    "input": "I missed you. Stay a little longer tonight.",
    "voice": "vc_4e1b…"
  }'
# → { "id": "au_7d2e…", "url": "https://cdn.spicyapi.com/audio/…/au_7d2e….wav", "duration_s": 3.9, … }

# 2. The clip: a first frame from your library, the line as driving audio
curl https://api.spicyapi.com/v1/videos/generations \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-motion-2",
    "prompt": "she says the line softly, looking into the camera",
    "image_url": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png",
    "audio_url": "https://cdn.spicyapi.com/audio/…/au_7d2e….wav",
    "duration": 5
  }'

# 3. Poll until succeeded
curl https://api.spicyapi.com/v1/videos/tasks/sj_… \
  -H "Authorization: Bearer $SPICYAPI_KEY"

On spicy-motion-3 and spicy-motion-3-fast the same URL is reference audio: send it in audio_url or audio_urls with reference images or a character (never a first frame) and write the dialogue in the prompt; the model takes voice, tone and beat from the clip and generates the sound itself. Sandbox keys return a fixture clip and a fixture voice, and nothing is billed.

Live voice callsNew

spicy-live-1 holds a spoken conversation: the caller talks, the persona answers in a natural voice, and interruptions are handled. It takes two steps. Your server creates a session with your key and gets back a one-time WebSocket url (valid once, within 60 seconds); the browser connects to that URL, so the key never leaves your server. 28 voices (default Tina), calls up to 13 minutes per session, billed per turn from audio and text tokens, which comes to roughly half a cent per minute of back-and-forth.

curl https://api.spicyapi.com/v1/realtime/sessions \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spicy-live-1",
    "voice": "Serena",
    "instructions": "You are Mara, 29, a playful bartender on a late-night call. Flirt, tease, keep replies short.",
    "turn_detection": { "type": "server_vad" }
  }'
# → { "id": "rt_5c2e…", "url": "wss://api.spicyapi.com/v1/realtime?session=rts_9f1b…", "max_seconds": 780, … }

The socket speaks the OpenAI Realtime event shape. Send microphone audio as input_audio_buffer.append (base64 PCM16 mono 16 kHz, about 100 ms per event) and play response.audio.delta (base64 PCM16 mono 24 kHz). Transcripts of both sides arrive as events, and spicy.usage reports the running cost after each turn.

// Browser, after your server returned session.url
const ws = new WebSocket(url);
const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
const inCtx = new AudioContext({ sampleRate: 16000 });
const proc = inCtx.createScriptProcessor(2048, 1, 1);
inCtx.createMediaStreamSource(mic).connect(proc);
proc.connect(inCtx.destination);
proc.onaudioprocess = (e) => {
  const f32 = e.inputBuffer.getChannelData(0);
  const pcm = Int16Array.from(f32, (s) => Math.max(-1, Math.min(1, s)) * 0x7fff);
  const b64 = btoa(String.fromCharCode(...new Uint8Array(pcm.buffer)));
  if (ws.readyState === 1) ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: b64 }));
};

const outCtx = new AudioContext({ sampleRate: 24000 });
let at = 0;
ws.onmessage = ({ data }) => {
  const ev = JSON.parse(data);
  if (ev.type === "response.audio.delta") {
    const pcm = new Int16Array(Uint8Array.from(atob(ev.delta), (c) => c.charCodeAt(0)).buffer);
    const buf = outCtx.createBuffer(1, pcm.length, 24000);
    buf.getChannelData(0).set(Float32Array.from(pcm, (s) => s / 0x8000));
    const src = outCtx.createBufferSource();
    src.buffer = buf;
    src.connect(outCtx.destination);
    at = Math.max(at, outCtx.currentTime);
    src.start(at);
    at += buf.duration;
  }
  if (ev.type === "spicy.usage") console.log("call so far $" + ev.total_cost_usd);
  if (ev.type === "error") console.warn(ev.error.code, ev.error.message);
};

The persona in instructionsis screened when the session is created, and both sides' transcripts are screened during the call; a violation ends it with a content_blocked error event. Only session.turn_detection can change mid-call (null for push-to-talk). No cloned voices, camera input or tools on calls yet. Full event list: Live Call WebSocket.

TranscriptionNew

POST /v1/audio/transcriptions turns voice notes and clips into text with spicy-transcribe-1, at $0.0042 per minute billed by the second. Send a multipart file (the OpenAI SDK's audio.transcriptions.create works) or JSON with a public url: wav, mp3, m4a, ogg, flac, webm and more, up to 10 MB and 5 minutes. The language is detected; pass language as a hint if you know it. Audio is not stored.

curl https://api.spicyapi.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -F "file=@voice-note.m4a" \
  -F "model=spicy-transcribe-1"
# → { "object": "audio.transcription", "text": "Hey, it's me…", "language": "en", "duration_s": 6, "cost_usd": 0.00042 }

EmbeddingsNew

POST /v1/embeddings is OpenAI-compatible. spicy-embed-1 embeds text for companion memory, prompt search and recommendations (1024 dimensions by default, 64 to 2048 with dimensions). spicy-embed-vision-1puts images and text in one 768-dimension space for "more like this" search, tagging and near-duplicate detection; its images must be outputs of your own account. Up to 10 inputs per request.

from openai import OpenAI

client = OpenAI(api_key="sk-spicy-...", base_url="https://api.spicyapi.com/v1")
memory = client.embeddings.create(
    model="spicy-embed-1",
    input=["She likes rainy evenings and old jazz records.", "Her favourite drink is a dirty martini."],
    dimensions=512,
)

# images and text in one space: search your own library with words
curl https://api.spicyapi.com/v1/embeddings \
  -H "Authorization: Bearer $SPICYAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "spicy-embed-vision-1", "input": [{ "image": "https://cdn.spicyapi.com/outputs/…/sj_2f1c…-0.png" }, { "text": "red silk robe on a balcony at night" }] }'

Chat completions

POST /v1/chat/completions is OpenAI-compatible, including streaming. Point any OpenAI SDK at the SpicyAPI base URL and change the model name.

from openai import OpenAI

client = OpenAI(
    api_key="sk-spicy-...",
    base_url="https://api.spicyapi.com/v1",
)

stream = client.chat.completions.create(
    model="spicy-chat-1",
    messages=[
        {"role": "system", "content": "You are Emma, a flirty girlfriend."},
        {"role": "user", "content": "hey, what are you up to?"},
    ],
    stream=True,
)

What is supported

Text in, text out. Forwarded as you send them: temperature, top_p, top_k, presence_penalty, frequency_penalty, repetition_penalty, seed, stop, logit_bias, stream. Capped: max_tokens default 4096 and at most 8192, n up to 4, about 200,000 tokens of input across all messages. response_format is text or json_object(for JSON mode the word "JSON" must appear in your prompt). Rejected with a 400 before any charge: thinking (enable_thinking, thinking_budget), tools and tool_choice, web search, json_schema, and image or audio content parts.

Role-play companionsNew

spicy-companion-1 ($1.00 in / $2.80 out per 1M tokens) is fine-tuned for role-play and AI companions. Put the persona in the system message and it stays in character, writes explicit adult scenes, handles group scenes with several characters, and with n returns up to 4 alternative replies for your users to pick from. It is the recommended model for explicit role-play, which spicy-chat-1 may decline. spicy-companion-1-flash ($0.10 in / $0.80 out per 1M tokens) is the same behaviour priced for high-volume traffic such as NPCs. Same request body as any chat call.

resp = client.chat.completions.create(
    model="spicy-companion-1",
    messages=[
        {"role": "system", "content": "You are Mara, 29, a confident bartender who flirts with the user. Stay in character, first person, under 120 words."},
        {"role": "user", "content": "Last call. What are you pouring me?"},
    ],
    n=3,
    temperature=0.9,
)
options = [c.message.content for c in resp.choices]  # three replies to choose from

Pair it with Spicy Voice 2 for spoken replies, embeddings for long-term memory, or move the conversation to a live voice call.

Webhooks

Video tasks (generations, edits and motion transfer) are asynchronous. Set a webhook URL in the dashboard under Settings → Webhooks and SpicyAPI POSTs there when a task reaches a terminal state, so you do not have to poll. Image, speech, transcription, embedding and chat requests are synchronous and never produce a webhook.

POST https://yourapp.com/hooks/spicy
Content-Type: application/json
X-SpicyAPI-Event: task.succeeded
X-SpicyAPI-Delivery: sj_9a2b...
X-SpicyAPI-Timestamp: 1758556800
X-SpicyAPI-Signature: <hex hmac-sha256 of the raw body>

{
  "event": "task.succeeded",
  "id": "sj_9a2b...",
  "created": 1758556800,
  "data": {
    "id": "sj_9a2b...",
    "object": "task",
    "type": "video",
    "model": "spicy-motion-3",
    "status": "succeeded",
    "output": { "video_url": "https://cdn.spicyapi.com/outputs/..." },
    "error": null,
    "error_code": null,
    "cost_usd": 3.2
  }
}

Events are task.succeeded, task.failed and webhook.test (sent by the Send test event button). The URL must be public https. Respond with any 2xx within 8 seconds; anything else counts as a failed delivery and is retried up to 6 times, after 1 minute, 5 minutes, 15 minutes, 1 hour and 4 hours. The same task id is used for every attempt, so treat deliveries as idempotent.

Verify the signature: compute HMAC-SHA256 of the raw request body with the signing secret shown in Settings → Webhooks, hex encode it, and compare it to X-SpicyAPI-Signature with a constant-time comparison.

import crypto from "node:crypto";

const expected = crypto
  .createHmac("sha256", process.env.SPICYAPI_WEBHOOK_SECRET)
  .update(rawBody)
  .digest("hex");
const ok = crypto.timingSafeEqual(
  Buffer.from(expected), Buffer.from(req.headers["x-spicyapi-signature"] ?? ""));

Errors & limits

Errors return a standard shape with an HTTP status: 401 bad key, 402 insufficient balance, 400 invalid parameters, 422 blocked by content screening, 429 rate limited or model at capacity, 503 content screening unavailable.

Limits per account: 600 requests a minute across all endpoints (task polling included), and 120 a minute that go through content screening (generations, chat, speech, moderations, dry runs). When you send a user id, each of your users also gets 30 screened requests a minute, so one of them cannot use up your account's allowance. At most 20 video tasks per account can be in progress at once; the next gets 429 with type concurrency_limit_reached until one finishes. Every 429 carries a Retry-After header in seconds. Need higher limits? Email contact@spicyapi.com.

{ "error": { "message": "...", "type": "insufficient_balance" } }

Content screening errors carry extra fields so your UI can tell the user what to change: code is the category (minor, real_person, non_consent, bestiality, gore, injection, keyword or a contextual rule id), outcome is block or review (uncertain: reword to make ages and consent explicit), and categories holds per-category probabilities when the semantic layer decided. stage: "output" means a generated image was withheld after generation; those are refunded. Finished videos get the same check: a withheld video fails its task with error_code: "blocked" and is refunded.

{
  "error": {
    "message": "Prompt was declined by content screening (real person): ...",
    "type": "moderation_blocked",
    "code": "real_person",
    "outcome": "block",
    "categories": { "minor": 0.04, "real_person": 0.99, "non_consent": 0.10, "bestiality": 0.0, "gore": 0.0, "injection": 0.03 },
    "docs_url": "https://www.spicyapi.com/acceptable-use?utm_source=api"
  }
}

Every prompt is screened before generation and every image is screened after; nothing blocked is ever charged. A 503 means the screen itself was unreachable: we fail closed rather than generate unscreened, so retry after a few seconds. Check prompts ahead of time with POST /v1/moderations. How the screen works is described on the moderation page; the rules are in the acceptable use policy.

Are you an AI agent?

Connect Claude Code, Cursor or Codex to the SpicyAPI MCP server to quote prices, create keys and generate on a user's behalf. Human sign-up is still required once for billing. Setup guide.

claude mcp add --transport http spicyapi https://www.spicyapi.com/api/mcp --header "Authorization: Bearer sk-spicy-..."