Seedance 2.0 Image-to-Video

bytedance/seedance-2-0/image-to-video
OfficialImage-to-Video

Seedance 2.0 Image-to-Video offers the highest image quality. It generates 4-15 second videos with text prompts and supports various aspect ratios, audio generation, and enhanced web search capabilities.

Read Me

Seedance 2.0 Image-to-Video API

Overview

Seedance 2.0 Image-to-Video turns a prepared first frame into a directed 4–15 second video while keeping the opening subject, composition, lighting, and visual style as the starting point.

The Prompt should describe what changes after the first frame: subject movement, camera travel, environmental motion, pacing, transition, and sound. Through the API, an optional last frame can also define the intended final composition.

Use this endpoint when the visual direction already exists and the model should animate it rather than build the entire scene from text. Use Reference-to-Video instead when several images, videos, or audio files must guide identity, style, motion, camera, or sound.

The iCreat endpoint supports 480P, 720P, 1080P, and 4K output, optional generated audio, fixed or automatic duration, and fixed or adaptive aspect ratio.

Use the model through:

bytedance/seedance-2-0/image-to-video

Key Features

Visual continuity from a prepared first frame. The starting image establishes the opening subject, composition, lighting, environment, and style before motion begins.

Motion-first Prompt control. The Prompt should focus on action, camera movement, environmental change, pacing, and sound instead of repeating details already visible in the image.

Optional last-frame direction. Through the API, a last frame can define where the shot should end. It works as a destination for the transition, not as a general identity or style reference.

Complex subject and environmental motion. Seedance 2.0 can develop body movement, object interaction, cloth motion, weather, lighting change, and camera combinations from a static visual foundation.

Optional synchronized audio. Set generate_audio to true to add dialogue, ambience, sound effects, or music while the image develops into video.

Adaptive framing and timing. Use adaptive when the source image and intended motion should guide the frame shape. Use duration: -1 when the model can choose a suitable length.

Output up to 4K. Use 480P or 720P for lower-cost motion testing. Use 1080P or 4K for selected outputs that need more visible detail or larger delivery dimensions.

Review handling for protected content. Set need_review: true when an input image contains faces or copyrighted IP. Incorrect review settings can cause failure or increase processing time.

Model Comparison

Seedance 2.0 Text-to-Video vs Image-to-Video vs Reference-to-Video

Route Primary Input Best Fit
Text-to-Video Text Prompt Full scene from text
Image-to-Video Prompt + Starting Image Animate an existing visual
Reference-to-Video Prompt + Reference Assets Reference-guided generation

Seedance 2.0 Image-to-Video vs Seedance 2.0 Fast Image-to-Video

Metric Seedance 2.0 I2V Seedance 2.0 Fast I2V
Latency Higher Lower
Cost Higher Lower
Input Mode Image-to-Video Image-to-Video
Output Resolution 480P, 720P, 1080P, 4K 480P, 720P
Best Fit Final output + complex motion Testing + variations

Inputs

The Playground accepts one starting image. The API also supports an optional last frame for a controlled destination.

Input Count Format Notes
Text Prompt 1 Required
Images 1–2 JPEG, PNG, WebP, BMP, TIFF, GIF, HEIC, HEIF First Frame + Optional Last Frame
Videos Not supported No video input
Audio Not supported No audio input; generated audio is optional

Each image must have an aspect ratio from 0.4 to 2.5, dimensions from 300 to 6,000 pixels, and a file size under 30 MB. The total request body must remain under 64 MB.

Chinese Prompts can contain up to 2,000 characters, and English Prompts can contain up to 2,000 words. When the image already defines the subject and scene, use the Prompt to describe change rather than restating the full composition.

Set need_review: true when a submitted image contains faces or copyrighted IP.

Parameters

Parameter Supported Values What It Controls
generate_audio true, false Generated audio
ratio 16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive Output aspect ratio
resolution 480p, 720p, 1080p, 4k Output detail and unit price
duration 415, -1 Output duration and total cost
watermark true, false Output watermark
tools web_search Web-assisted context

Pricing

Resolution Unit Price 5 Seconds 10 Seconds 15 Seconds
480P $0.077/sec $0.385 $0.770 $1.155
720P $0.164/sec $0.820 $1.640 $2.460
1080P $0.409/sec $2.045 $4.090 $6.135
4K $0.830/sec $4.150 $8.300 $12.450
Total Cost = Unit Price × Output Video Duration

Quick Start

This request starts from one product frame and uses a second frame to define the final composition.

curl --fail-with-body --connect-timeout 10 --max-time 60 \
  -X POST https://api.icreat.ai/v1/task/submit/bytedance/seedance-2-0/image-to-video \
  -H "Authorization: Bearer ${ICREAT_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "content": [
      {
        "type": "text",
        "text": "Begin with the product resting on the dark platform. The camera slowly pushes forward as thin beams of light sweep across the surface and small dust particles drift through the air. The product rotates slightly while the background lighting changes from cool blue to warm gold. Keep the product shape, materials, logo placement, and proportions consistent. End smoothly in the supplied last-frame composition. Add a low mechanical hum, a soft rising sound effect, and restrained cinematic music."
      },
      {
        "type": "image_url",
        "image_url": {
          "url": "https://example.com/first-frame.jpg"
        },
        "role": "first_frame",
        "need_review": false
      },
      {
        "type": "image_url",
        "image_url": {
          "url": "https://example.com/last-frame.jpg"
        },
        "role": "last_frame",
        "need_review": false
      }
    ],
    "generate_audio": true,
    "ratio": "16:9",
    "resolution": "1080p",
    "duration": 8,
    "watermark": false
  }'

The request returns a task_id.

Submit request → Receive task_id → Poll status → Retrieve result

Query the task status until it reaches SUCCEEDED, then use the same task_id to retrieve the result.

Output

Seedance 2.0 Image-to-Video uses an asynchronous task flow.

The submission request returns a task_id. Task status is checked through the query-status endpoint, and the completed result is retrieved through the get-result endpoint after the task succeeds.

Detailed result fields are not listed until they are confirmed from an actual iCreat response.

Use Cases

Product-state transition. Start from a prepared product image and guide the shot toward a different position, lighting state, material reveal, or final arrangement.

Character performance from key art. Animate expression, body movement, clothing, hair, camera behavior, and sound from an established character visual.

Scene evolution from a concept frame. Turn a still environment into a short sequence through weather, lighting, object movement, or character interaction.

First-to-last campaign transformation. Use two prepared frames to control a reveal, time change, product transformation, or narrative destination.

Image-led audio-visual shot. Add dialogue, ambience, sound effects, or music to a scene whose visual identity already exists.

Limitations

The source image guides the result but is not preserved pixel for pixel. Faces, anatomy, clothing, product details, logos, text, background objects, and lighting can change as motion develops.

Large body movement, rapid camera travel, major perspective changes, and several competing actions increase the risk of drift, flicker, or unwanted scene changes.

The selected output ratio may differ from the source image ratio. The model may crop, extend, or reposition parts of the image to fit the requested frame shape.

First and last frames that differ greatly in subject identity, composition, perspective, lighting, or environment can produce unstable transitions or unexpected intermediate frames.

Readable text is not guaranteed. Logos, labels, signs, and interface elements may distort as the image moves.

Detail stability, hyper-realism, dynamic vitality, and audio quality can still weaken in demanding scenes. Occasional audio distortion or timing mismatch may occur.

Incorrect need_review settings can cause failure or longer processing time. The Playground currently exposes only one image upload, so optional last-frame control is available through the API.

The provided endpoint does not expose a seed, so exact repeatability should not be promised.

FAQ

What should the Prompt describe when the image already contains the subject and scene?

Describe change rather than repeating the full image. Focus on subject action, camera movement, environmental motion, pacing, transition, and sound. Repeat a visual detail only when it must remain stable.

When should a Last Frame be added, and how similar should it be to the First Frame?

Add a Last Frame when the shot must reach a specific pose, product position, lighting state, environment, or composition. The two frames should share a recognizable subject and a believable transition path. Large differences in identity, viewpoint, scale, lighting, or background make the transition harder to control.

How can face, clothing, and product details be preserved during motion?

Use a clear source image, keep movement within a believable range, avoid extreme rotations, and state which details must remain unchanged. Slower motion and smaller camera-angle changes usually reduce drift.

Why does the output crop, reframe, or introduce an unwanted scene change?

The requested ratio may differ from the source image ratio, and the intended movement may require additional frame space. Conflicting camera instructions, several environments, unrelated actions, or a very different Last Frame can also cause reframing or scene changes.

Why is need_review required for some images?

Images containing faces or copyrighted IP may require official review before generation. Submitting them without review can cause failure, while enabling review unnecessarily may increase processing time.