How to make realistic POV sex videos
This is the exact recipe behind the POV position clips on our examples page: kneeling blowjob, missionary, doggystyle and cowgirl, 8 seconds, 1080P, portrait, with native sound. It takes three calls. The first makes the opening frame, the second gets the motion right with a fine-tuned POV model, and the third re-renders that motion at the newest video quality with sound.
Why three calls: a general video model on its own tends to get sexual motion wrong (oral that stays at the tip, slow-motion thrusting, bodies that drift out of position). The spicy-pov-* fine-tunes were trained on exactly these motions, and spicy-motion-3 can follow a reference clip closely while adding its own rendering and audio. Together they give the motion of the first and the look and sound of the second.
Every input image and clip must be one of your own outputs from this API, so run the steps in order on one key.
Step 1: the opening frame
Generate a portrait still with spicy-image-1-pro at 1152*2048. This frame decides everything that follows, so make three or four with separate calls and different seeds, and pick the best. (Asking for n: 4 in one call returns near-duplicates; separate seeded calls give real variety.)
Rules that made the difference in our testing:
- Write one calm photographic paragraph, not a list and not instructions.
- Describe the act by what disappears into what ("his penis is inside her", "the shaft deep in her mouth, lips stretched around it"). Saying something is "near" or "at" another body part puts it in the wrong place.
- Shoot it as first-person POV and say his face is never visible. Place him with one short clause ("he kneels upright behind her, the tops of his thighs in the lower frame"). Saying nothing about him cuts his body off; listing his body parts makes them the focus.
- Anchor her pose to surfaces ("her forearms flat on the sheets, her head resting on the pillow"), not to where she looks. "Looking up at the camera" from above makes her prop herself up.
- Keep a short negative prompt for the failures you actually see.
Example: doggystyle frame
{
"model": "spicy-image-1-pro",
"size": "1152*2048",
"seed": 11,
"prompt": "A first-person view from above in a softly lit bedroom, a man's face never visible, only his bare hips, lower abdomen and hands at the bottom of the frame. He kneels upright behind her on the bed, the tops of his thighs in the lower frame. A beautiful woman in her late twenties with a toned nude figure kneels on all fours on the mattress, her ass pressed against his hips, a visible line of skin where they meet, his hands gripping her hips with fingers visible, her back arched away from him, her forearms flat on the sheets, her hair falling forward. Her face is turned over her shoulder, eyes on the camera. His penis is inside her neat natural vulva, her anus visible just above the point of insertion. The woman has copper red hair, green eyes, light freckles and a heart-shaped face, light natural makeup. She is clearly enjoying herself: a blissful, seductive expression, eyes heavy-lidded and locked on the camera, lips parted, cheeks flushed.",
"negative_prompt": "floating penis, external penis, uninserted penis, deformed genitals, sitting up, lying flat, merged bodies, extra limbs, missing anus, legless torso, smiling, showing teeth"
}POST /v1/images/generations returns the image URL in data[0].url. Keep it: it is the input for steps 2 and 3.
Step 2: the motion, from a POV fine-tune
Animate the frame with the matching fine-tune: spicy-pov-blowjob-1 (kneeling oral), spicy-pov-missionary-1 or spicy-pov-doggystyle-1. The frame goes in as image_url, the first frame. The fine-tune's activation token is added for you, so a short prompt is enough.
{
"model": "spicy-pov-doggystyle-1",
"image_url": "<data[0].url from step 1>",
"prompt": "POV doggystyle motion. Forward thrusts drive deeper with accelerating rhythm, glutes rippling on impact, shaft disappearing inside on each stroke. Camera stays fixed behind without panning.",
"resolution": "1080P",
"duration": 8
}POST /v1/videos/generations returns a task. Poll GET /v1/videos/tasks/{id} (or set a webhook) until it succeeds, then keep output.video_url.
Step 3: the re-render with native sound
Now give spicy-motion-3 the frame as a reference image and the step 2 clip as a reference video. A first frame and a reference clip cannot go in the same request, so the frame goes in reference_image_urls, where the prompt calls it "Image 1", and the clip goes in video_url, where it is "Video 1".
{
"model": "spicy-motion-3",
"reference_image_urls": ["<data[0].url from step 1>"],
"video_url": "<output.video_url from step 2>",
"aspect_ratio": "9:16",
"resolution": "1080P",
"duration": 8,
"prompt": "Image 1 is the opening shot: reproduce the woman exactly as in Image 1, her face, hair, body and skin, the same bedroom, lighting and the same POV framing and composition as Image 1. Video 1 is the motion reference: follow the motion, depth, pacing, rhythm and camera of Video 1 exactly, second for second, including where the bodies connect. Real-time speed, never slow motion.\n\nThe camera is completely locked, the exact same framing, angle, distance and composition as the input image. No camera movement, no zoom, no pan, no flash and no lighting change.\n\nCRITICAL, DEPTH AND TONE: short deep strokes in real time, about two thrusts per second. He never pulls out more than halfway, so the tip ALWAYS STAYS INSIDE her. The shaft keeps the same warm natural flesh tone along its WHOLE LENGTH at all times, NEVER PALE and NEVER WHITE. On each drive his pelvis meets her ass and her cheeks ripple from the impact.\n\nShe looks DIRECTLY INTO THE CAMERA throughout, lips parted, clearly enjoying it.\n\nAudio: breathy moans in time with the rhythm, plus rhythmic skin-on-skin sounds at each impact. No music, no speech."
}In our clips the result followed the reference motion at 0.9 to 1.0 similarity every second and still opened on the chosen frame.
Writing the motion prompt
Video prompts work differently from image prompts. What worked for us is four blocks, in this order:
- Camera: completely locked, the exact framing of the input image, no zoom, pan or lighting change.
- Critical motion: the act stated precisely, with capitals on the one or two things that must not go wrong (the tip stays inside, her lips never leave the shaft, her hands stay where they are).
- Face: pretty, seductive, enjoying it, eyes on the camera where the frame has her looking at it.
- Audio: muffled humming and wet sounds for oral, breathy moans in time with the rhythm plus skin sounds for penetration, ending "No music, no speech."
Avoid the words slow, slowly, gentle, gently, smooth, steady, languid and sensual: they produce slow-motion clips. spicy-motion-3 has no negative prompt, so say what you want positively ("the shaft keeps a warm natural flesh tone, never pale").
Positions without a fine-tune
For reverse cowgirl, titjob or solo poses, skip step 2 and animate the frame directly on spicy-motion-3 with image_url as the first frame and a four-block motion prompt. The motion is less reliable than with a fine-tune; generate two or three and keep the best.
Cost
At current prices one finished 8 second 1080P clip costs about $9.07:
- three candidate frames on
spicy-image-1-pro: $0.27 - 8 seconds on a
spicy-pov-*fine-tune at 1080P: $2.40 - 8 seconds on
spicy-motion-3at 1080P, which bills the 8 seconds of reference clip as well as the 8 seconds of output: $6.40
Check live prices on the pricing page or ask the API for a quote first.