# Create Speech

`POST /v1/audio/speech`

Turn text into speech. Returns a durable audio URL you can pass straight to video generation as `audio_url`, or streams the audio as it is synthesized.

The request follows the OpenAI speech shape (`model`, `input`, `voice`, `instructions`, `response_format`). By default the call is synchronous and the response is JSON with a `url`, not raw audio bytes: the clip is stored on the CDN (WAV 24 kHz 16-bit mono, or MP3 on Spicy Voice 2) and registered as one of your audio clips, so it shows up in `GET /v1/audio` with `source: "speech"`.
Five models. `spicy-voice-1` speaks 45 preset voices. `spicy-voice-1-expressive` speaks 21 of them and also takes `instructions` describing emotion, pace, pitch and delivery. `spicy-voice-1-custom` speaks only your own voices, the `vc_...` ids from `POST /v1/voices`. `spicy-voice-2` (voices `Lingxin`, `Lufeng`) and `spicy-voice-2-flash` (voices `Fengyue`, `Yuanfei`, `Lingxi`, `Xiaoxin`, `Huan`, `Chuanshu`, `Mary`, `Eva`, `John`) are the expressive generation: inline tags, `instructions`, up to 5,000 characters, MP3 output and streaming, in `Auto`, `English`, `Chinese`; custom voices stay on `spicy-voice-1-custom`. A custom voice on a preset model, a preset voice on the custom model, or `instructions` on a model without them is a 400 that names the right model. `GET /v1/voices` lists every voice with the models that speak it.
Inline tags (Spicy Voice 2 models): write them in square brackets inside `input`. Control tags change the delivery until the next tag: `[sad]`, `[amazed]`, `[deep and loud shouting]`, `[trembling]`, `[angry]`, `[excited]`, `[sarcastic]`, `[curious]`, `[like dracula]`, `[bored]`, `[tired]`, `[scornful]`, `[shouting]`, `[asmr]`, `[panicked]`, `[mischievously]`, `[empathetic]`, `[whispers]`, `[reluctantly]`, `[crying]`, `[serious]`, `[very slowly]`, `[very fast]`. Sound tags insert a vocal sound at that point: `[gasp]`, `[sighing]`, `[clears throat]`, `[giggles]`, `[laughing]`, `[cough]`, `[snorts]`. Tags count as billed characters. Example: `[excited]Hey you, I've been waiting all night.[giggles] Come here.[whispers] Closer.`
Streaming (Spicy Voice 2 models): with `stream: true` the response body is the raw audio as it is synthesized, the first audio arriving in about 1.5 seconds. `response_format` picks the container: `pcm` (default: 24 kHz, 16-bit, mono, little-endian, no header), `wav` or `mp3`. Headers: `x-spicyapi-audio-id` (the clip is stored afterwards as an audio asset under that id, so it appears in `GET /v1/audio`), `x-spicyapi-sample-rate: 24000` and `x-spicyapi-cost-usd` (the amount debited before the stream starts; settled to the billed count afterwards). If the stream fails before any audio arrives, the charge is refunded.
Billed per 10,000 characters of `input`: $0.20 on `spicy-voice-1`, $0.23 on `spicy-voice-1-expressive`, $0.23 on `spicy-voice-1-custom`, $0.40 on `spicy-voice-2`, $0.30 on `spicy-voice-2-flash`. CJK ideographs count as 2 characters. `instructions` are included in the debit; when the billed count comes back lower the difference is refunded, and `cost_usd` is the settled amount. Failed generations are refunded.
The `input` and the `instructions` are screened like a prompt before anything is charged; a blocked request returns the same 422 as a blocked prompt. `dry_run: true` validates, screens and returns the price without generating. Sandbox keys get a fixture clip flagged `sandbox: true` with the `cost_usd` a live key would pay (one second of silence when streaming); nothing is billed.
Make a character speak: pass the returned `url` byte-for-byte as `audio_url` on `POST /v1/videos/generations`. On `spicy-motion-2` and `spicy-video-1` it is driving audio: the clip becomes the soundtrack and the mouth follows the words, alongside a first frame where the model takes one. On `spicy-motion-3` and `spicy-motion-3-fast` it is reference audio: the model uses it for voice, tone and beat while the words come from the prompt, with a text prompt or references and never with a first frame (up to 5 clips, 15 seconds in total). Audio for video must be 2 to 30 seconds long, so split long text across requests.

Base URL: `https://api.spicyapi.com`

## Authorizations

- `Authorization` (string, header, required): Bearer authentication header of the form `Bearer <token>`, where `<token>` is your SpicyAPI key (`sk-spicy-…`). Create one in the dashboard under API Keys.

## Body (application/json)

- `model` (string, required): `spicy-voice-1`, `spicy-voice-1-expressive`, `spicy-voice-1-custom`, `spicy-voice-2` or `spicy-voice-2-flash`.
- `input` (string, required): The text to speak: at most 600 characters on Spicy Voice 1 models, 5,000 on Spicy Voice 2 models (inline tags included). Longer text is a 400: split it into several requests.
- `voice` (string, required): A preset voice name the model speaks (case-insensitive, e.g. `Cherry` on Spicy Voice 1, `Lingxin` on `spicy-voice-2`, `Eva` on `spicy-voice-2-flash`), or on `spicy-voice-1-custom` one of your `vc_...` voice ids. See `GET /v1/voices`.
- `instructions` (string, optional): `spicy-voice-1-expressive`, `spicy-voice-2` and `spicy-voice-2-flash`: how to say it, in plain English or Chinese (emotion, pace, pitch, delivery), at most 1,600 characters. Counted in the debit, settled to the billed count.
- `language` (string, optional): One of `Auto`, `English`, `Chinese`, `German`, `Italian`, `Portuguese`, `Spanish`, `Japanese`, `Korean`, `French`, `Russian` on Spicy Voice 1 models, `Auto`, `English`, `Chinese` on Spicy Voice 2 models (case-insensitive; the model's `speech.languages`). Defaults to `Auto`, which detects the language as it goes.
- `response_format` (string, optional): Spicy Voice 1 models: `wav` only (`mp3` is a 400). Spicy Voice 2 models: `wav` (default) or `mp3`; with `stream: true`, `pcm` (default), `wav` or `mp3`.
- `stream` (boolean, optional): Spicy Voice 2 models only: `true` returns the audio bytes as they are synthesized instead of JSON (see Streaming above). A 400 on Spicy Voice 1 models.
- `dry_run` (boolean, optional): When `true`, validates and screens the request and returns `{ object: "dry_run", cost_usd, moderation }` without generating or charging. A blocked prompt still returns the normal 422.
- `user` (string, optional): Your own id for the end user making this request (up to 128 characters, hashed at rest). Send it if your product serves many people: declined-prompt history, strikes and suspensions are then kept per end user, so one person's behaviour never affects another's requests or your account. After 10 severe violations that user gets `403 end_user_suspended`; the id is echoed back as `user` in every screening error so you can act on it.

## Request

```bash
curl --request POST \
  --url https://api.spicyapi.com/v1/audio/speech \
  --header 'Authorization: Bearer $SPICYAPI_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "spicy-voice-1",
    "input": "Come here. I have been waiting all evening.",
    "voice": "Cherry"
  }'

# expressive: inline tags on Spicy Voice 2, as MP3
curl --request POST \
  --url https://api.spicyapi.com/v1/audio/speech \
  --header 'Authorization: Bearer $SPICYAPI_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "spicy-voice-2",
    "input": "[excited]Hey you, I have been waiting all night.[giggles] Come here.[whispers] Closer.",
    "voice": "Lingxin",
    "response_format": "mp3"
  }'

# streaming: raw 24 kHz PCM as it is synthesized
curl --request POST \
  --url https://api.spicyapi.com/v1/audio/speech \
  --header 'Authorization: Bearer $SPICYAPI_KEY' \
  --header 'Content-Type: application/json' \
  --data '{ "model": "spicy-voice-2-flash", "input": "[whispers]Stay on the line with me.", "voice": "Eva", "stream": true }' \
  --dump-header headers.txt --output line.pcm
```

## Response: 200 application/json

Speech

The JSON response. With `stream: true` the body is the audio itself (`audio/pcm; rate=24000`, `audio/wav` or `audio/mpeg`) and the id and cost are in the `x-spicyapi-audio-id` and `x-spicyapi-cost-usd` headers.

- `id` (string, required): Clip id, `au_...`; the same id appears in `GET /v1/audio`.
- `object` (string, required): Always `audio.speech`.
- `model` (string, required): The model that spoke.
- `url` (string, required): CDN URL of the clip. Pass it byte-for-byte as `audio_url` (or in `audio_urls`) on `POST /v1/videos/generations`.
- `duration_s` (number, required): Length in seconds, to one decimal.
- `voice` (string, required): The preset name or `vc_...` id that was used.
- `format` (string, required): `wav` or `mp3`.
- `cost_usd` (number, required): What this request cost in US dollars, settled to the billed character count.

```json
{
  "id": "au_7d2e9c1a4b6f8d0e2a",
  "object": "audio.speech",
  "model": "spicy-voice-1",
  "url": "https://cdn.spicyapi.com/audio/a1b2c3d4e5f6a7b8/au_7d2e9c1a4b6f8d0e2a.wav",
  "duration_s": 3.4,
  "voice": "Cherry",
  "format": "wav",
  "cost_usd": 0.00086
}
```
