Seedance 2.0 Fast

bytedance/seedance-2-0-fast
OfficialText-to-VideoImage-to-VideoVideo-to-VideoAudio-to-Video

Seedance 2.0-fast offers rapid video generation. It generates 4-15 second videos with text prompts, supports various aspect ratios, audio generation, and enhanced web search capabilities.

Read Me

Seedance 2.0 Fast API

Overview

Seedance 2.0 Fast is the faster, lower-cost tier of ByteDance’s multimodal video model. It generates short videos from text, images, video, and audio while keeping scene direction, camera movement, reference assets, editing instructions, and synchronized sound inside one workflow.

The model supports more than one way to begin a generation. A prompt can create the full scene from scratch. A first-frame image can define the opening composition. Images, video, and audio can guide the subject, style, movement, environment, or sound. Existing footage can also be used as the basis for a focused edit.

On iCreat, these workflows are available through one model ID:

bytedance/seedance-2-0-fast

Seedance 2.0 Fast produces 4–15 second videos at 480p or 720p. It is a practical choice for prompt testing, creative variations, reference-driven shots, and other workloads where speed and generation cost matter more than the highest available render quality.

Key Features

Multimodal references in one generation. Seedance 2.0 Fast can combine text with images, video, and audio. Reference images can guide a character, product, scene, composition, or visual style. Reference video can provide motion, camera, or editing context, while reference audio can influence rhythm, dialogue, ambience, or music.

Stronger shot direction from the prompt. Prompts can describe the subject and action together with the environment, camera movement, sequence of events, pacing, and sound. This makes the model suitable for shots that need deliberate direction rather than simple image animation.

First-frame and last-frame control through the API. A first_frame image defines how the output begins. Adding a last_frame also guides the final composition and the transition between the two. These controls are currently available through the API rather than the iCreat Playground.

Native audio generation. Set generate_audio to true to create dialogue, ambience, music, or sound effects with the video. Generated audio is included in iCreat’s listed per-second price and does not require a separate audio-generation request.

Reference-based editing. Existing video can be combined with an instruction and supporting references to change a subject, object, action, setting, or visual treatment. Focused changes are generally easier to control than requests that replace every part of the original clip.

A faster tier for repeated iteration. Fast retains the main Seedance 2.0 generation modes while prioritizing lower latency and cost. It is better suited to drafts, variations, and repeated generations; Seedance 2.0 is the stronger choice when maximum output quality is the main priority.

Model Comparison

Seedance 2.0 Fast vs Seedance 2.0

Metric Seedance 2.0 Fast Seedance 2.0
Latency Lower Higher
Cost Lower Higher
Input Modes T2V; first/last frames; image/video/audio references; native audio; editing Same core modes
Output Quality Speed/cost-first; 480p–720p Quality-first; 480p–4K
Best Fit Drafts, variations, faster iteration Final assets, high-detail output, 4K delivery

Seedance 2.0 Fast vs Kling VIDEO 3.0 vs Google Veo 3.1 vs Runway Gen-4.5

Metric Seedance 2.0 Fast Kling VIDEO 3.0 Google Veo 3.1 Runway Gen-4.5
Output Resolution 480p–720p 720p–4K 720p–4K 720p
Native Clip Length 4–15 sec 3–15 sec 4, 6, or 8 sec 2–10 sec
Reference Support 9 images; 3 videos; 3 audio files; first/last frames Start/end frames; reusable elements First/last frames; up to 3 images Text or one starting image
Native Audio Yes Yes Yes Yes
Shot & Editing Control Multi-shot prompts; frame guidance; video editing Structured multi-shot; per-shot timing Frame guidance; video extension Prompt-led motion; separate editing tools

Inputs

Every request must contain a text prompt. Media files are optional and tell the model what to preserve, reference, guide, or change.

Input Count Format Notes
Text Prompt 1 Required. Explain how each uploaded asset should influence the result. Refer to assets consistently, such as Image 1, Video 1, and Audio 1.
Images Up to 9 JPEG, PNG, WebP, BMP, TIFF, GIF, HEIC, HEIF Each file must be under 30 MB, between 300 and 6,000 px, and within a 0.4–2.5 aspect ratio. Supported roles are first_frame, last_frame, and reference_image.
Videos Up to 3 MP4, MOV Each video can be 2–15 seconds; combined duration cannot exceed 15 seconds. Input resolution can range from 480p to 4K at 24–60 FPS, with each file under 200 MB.
Audio Up to 3 MP3, WAV Each file can be 2–15 seconds, with a combined duration of up to 15 seconds and a maximum size of 15 MB per file. The current documented workflow pairs reference audio with an image or video.

A 4K reference video does not produce a 4K result. Seedance 2.0 Fast output remains limited to 480p or 720p.

Set need_review to true when a reference contains faces or copyrighted IP. Leaving it disabled when review is required can cause the generation to fail; enabling it unnecessarily may increase processing time.

Parameters

Parameter Supported values What changes
content Required multimodal array Provides the prompt, frames, references, and source media
generate_audio true, false Creates synchronized sound with the output
ratio 16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive Sets the output frame shape
resolution 480p, 720p Changes visible detail and price
duration 415, or -1 Sets the output length or allows automatic selection
watermark true, false Adds or removes the AI-generated watermark
tools web_search Enables optional web-assisted generation

Use 480p for lower-cost testing and 720p when the selected result needs more visible detail. Shorter durations are usually easier to control when the shot contains one main action or camera movement. Longer clips provide more room for sequences and transitions but increase cost.

Use a fixed aspect ratio when the output format is already known, such as 16:9 for landscape video or 9:16 for vertical content. Use adaptive when the model should choose a suitable format from the input.

Pricing

iCreat charges for Seedance 2.0 Fast by output video duration.

Resolution Reference video Price per output second 5 seconds 10 seconds
480p No $0.062 $0.31 $0.62
480p Yes $0.055 $0.275 $0.55
720p No $0.132 $0.66 $1.32
720p Yes $0.118 $0.59 $1.18

Only the generated output duration is billed. The duration of an uploaded reference video is not added to the charge. Generated audio is included at no extra cost.

Quick Start

This request uses a reference image and reference audio to show how multimodal inputs work in one generation.

curl --fail-with-body --connect-timeout 10 --max-time 60 \
  -X POST https://api.icreat.ai/v1/task/submit/bytedance/seedance-2-0-fast \
  -H "Authorization: Bearer ${ICREAT_API_KEY}" \
  -H "X-Icreat-Ai-Group: default" \
  -H "Content-Type: application/json" \
  -d '{
    "content": [
      {
        "type": "text",
        "text": "Keep the product shape and materials consistent with Image 1. Create a slow camera orbit as the lighting becomes warmer. Use Audio 1 to guide the pace."
      },
      {
        "type": "image_url",
        "image_url": {
          "url": "https://example.com/product.png"
        },
        "role": "reference_image"
      },
      {
        "type": "audio_url",
        "audio_url": {
          "url": "https://example.com/music.mp3"
        },
        "role": "reference_audio"
      }
    ],
    "generate_audio": true,
    "ratio": "16:9",
    "resolution": "720p",
    "duration": 8,
    "watermark": false
  }'

The request returns a task_id. Use that ID to query the task status and retrieve the result when the status reaches SUCCEEDED.

Submit request → Receive task_id → Poll status → Retrieve result

A Text-to-Video request uses the same endpoint with only the text item in content. For API-based frame control, use first_frame or combine first_frame with last_frame.

Output

Seedance 2.0 Fast uses an asynchronous task flow. The submission call returns a task_id, followed by separate status and result requests.

The detailed result-field schema is not included until it has been confirmed from a successful iCreat response. This prevents fields from another provider or an outdated response format from being presented as iCreat output.

Use Cases

Product and advertising video. Reference images can preserve product shape, material, color, packaging, or brand styling while the prompt changes the camera movement, environment, action, or sound.

Character and narrative scenes. Character references, scene images, and written dialogue can be combined to create short interactions, multi-shot moments, and story-led clips with generated audio.

Storyboard and concept visualization. Text prompts and rough visual references can turn an early idea into a short moving concept for evaluating composition, camera direction, timing, and audiovisual tone.

Music- and audio-led video. Reference audio can guide pacing, movement, transitions, or mood, while native audio generation can add dialogue, ambience, music, and effects in the same output.

Video transformation and editing. Source footage can be used with an instruction and additional references to change a selected subject, action, prop, environment, or visual treatment.

Limitations

Seedance 2.0 Fast produces 4–15 second videos at 480p or 720p. The current iCreat implementation does not support 1080p output. Every request requires a text prompt, and explicit first-frame and last-frame controls are currently API-only.

Complex scenes may still lose fine detail, realistic motion, dynamic energy, or consistency across multiple subjects. Audio can occasionally distort or miss the intended timing. Results are generally easier to control when each shot has one primary action, one clear camera instruction, and references that agree on identity, styling, and environment.

FAQ

How should a multi-shot prompt be structured?

Separate the prompt into Shot 1, Shot 2, and Shot 3. For each shot, describe who or what is present, the action, the setting, and the camera movement. Use one main camera instruction per shot and avoid forcing exact second-by-second timing unless it is essential.

How should multiple reference files be assigned?

Give every reference one clear purpose and name it directly in the prompt. One image might define the character, another the product, a video the camera movement, and an audio file the rhythm. Remove references that conflict in identity, style, environment, or direction.

Can the same character or product remain identical across separate generations?

Consistent references and repeated descriptions can reduce drift, but exact continuity across separate tasks is not guaranteed. Reuse the same main images and keep defining details such as clothing, materials, colors, lighting, and visual style unchanged.

What can reduce jumps between the first and last frame?

Use source images with compatible composition and aspect ratios. Large differences in framing, subject position, lighting, or scene structure can create abrupt movement. Matching the output ratio to the source images, or using adaptive, can also reduce stretching and cropping.

Is Seedance 2.0 Fast a dedicated lip-sync model?

It can generate spoken dialogue and synchronized audiovisual content, but it is not described as a specialist lip-sync or talking-avatar model. Test close-up dialogue carefully. A dedicated lip-sync or avatar model may be more suitable when exact mouth movement must follow a fixed recording.