Image-to-video API — animate a still frame
Send a reference image and a motion prompt; get a clip that starts from that frame. Used for product turntables, avatar animation and bringing a hero image to life.
Crazyrouter is an AI API gateway that serves 6 image-to-video models — Kling, Seedance, Veo 3.1 and Hailuo — through one asynchronous endpoint that accepts a reference frame, billed per clip from $0.0357.
Clips from these models
Generated through this API, 2026-09-28Image-to-video models, by price per clip
| Model | Vendor | Crazyrouter price (per generation) | Discount | |
|---|---|---|---|---|
| Kling 2.6 aigc-video-kling-2.6 | Kling AI | $0.0357–$0.268 / second | −15% | Try |
| Kling 3.0 aigc-video-kling-3.0 | Kling AI | $0.0714–$0.143 / second | −15% | Try |
| Hailuo H3 hailuo-h3 | Hailuo (MiniMax) | $0.0714–$0.251 / second | — | Try |
| Google Veo 3.1 aigc-video-gv-3.1 | $0.20–$0.60 / second | — | Try | |
| Seedance 1.5 Pro doubao-seedance-1-5-pro | ByteDance | $1.14–$2.29 / 1M tokens | — | Try |
| Seedance 2.0 doubao-seedance-2-0 | ByteDance | $4.00–$6.57 / 1M tokens | — | Try |
Prices come from the live billing table and already include the Crazyrouter discount; list prices are each vendor's public rate, for comparison only.
What is a image-to-video API?
Image-to-video keeps the first frame fixed to your image and generates the motion described in the prompt. It gives far more control over subject identity than text-to-video, which is why e-commerce and avatar products use it.
How to choose
Kling 3.0 / 2.6 hold product geometry best across the rotation.
Seedance 2.0 and Hailuo H3 for natural facial motion.
Veo 3.1 adds synchronized audio to the animated clip.
Seedance 1.5 Pro for cheap iterations on the motion prompt.
Animate a frame
Swap model for any row in the table; the example uses aigc-video-kling-2.6.
curl https://api.crazyrouter.com/v1/video/generations \
-H "Authorization: Bearer $CRAZYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"aigc-video-kling-2.6","prompt":"the camera slowly orbits the product","image":"https://example.com/frame.jpg","duration":5}'FAQ
How do I pass the reference image?
Add an image field (URL or base64) to the same POST /v1/video/generations body; the model treats it as the first frame.
Can I control the end frame too?
Kling supports first- and last-frame inputs; see the model page for the exact fields.
Is it billed differently from text-to-video?
No — per clip at the rate in the table.
Which image-to-video model is cheapest?
Kling 2.6 at $0.0357 per clip as of the latest sync.
What image size should I send?
Match the output aspect ratio (16:9 or 9:16) and at least 1280 px on the long side for best results.
