Seedance 2.0

bytedance/seedance-2-0
OfficialVideo-to-VideoImage-to-VideoText-to-VideoAudio-to-Video

Seedance 2.0 offers the highest image quality. It generates 4-15 second videos with text prompts and supports various aspect ratios, audio generation, and enhanced web search capabilities.

Read Me

Seedance 2.0 API

Overview

Seedance 2.0 is ByteDance’s quality-first multimodal video model for generating and editing 4–15 second videos from text, images, video, and audio. It combines reference-guided creation, complex motion, camera direction, native sound, and output up to 4K in one model workflow.

A generation can begin from a prompt, a first frame, a defined first and last frame, or a collection of image, video, and audio references. Existing footage can also be used as source material for focused edits. This gives Seedance 2.0 a wider role than a single Text-to-Video or Image-to-Video endpoint.

Use the model through one iCreat model ID:

bytedance/seedance-2-0

Seedance 2.0 is best suited to selected outputs where higher resolution and maximum visual quality matter more than the lowest latency or cost. It supports 480p, 720p, 1080p, and 4K output, while Seedance 2.0 Fast is the lower-cost option for drafts, previews, and repeated iteration.

Key Features

Quality-first output up to 4K. Seedance 2.0 adds 1080p and 4K output tiers to the core Seedance workflow. Lower resolutions remain useful for tests, while the higher tiers fit selected shots that need more visible detail, cleaner product surfaces, or larger delivery formats.

Multimodal reference control. Images can guide a character, product, scene, composition, or style. Video can provide motion, camera behavior, source footage, or editing context. Audio can guide rhythm, dialogue, ambience, music, or sound design. Up to nine images, three videos, and three audio files can be included in one request.

Complex motion and camera direction. Prompts can describe multiple subjects, physical interactions, action sequences, camera movement, pacing, and shot changes. Clear priorities still matter: one primary action and one main camera instruction per shot are easier to control than several competing directions.

Native audio generation. Set generate_audio to true to create dialogue, ambience, music, or sound effects with the video. Generated audio is included in the listed price, so the visual and audio tracks do not require separate generation requests.

First- and last-frame guidance. Use first_frame to define how a video begins. Add last_frame to guide the final composition and the transition between both images. These roles differ from reference_image, which provides visual guidance without fixing the opening or ending frame.

Reference-based video editing. An existing video can be combined with a prompt and supporting references to change a subject, object, action, environment, or visual treatment. Focused edits are generally easier to preserve than instructions that rebuild every part of the source footage.

Model Comparison

Seedance 2.0 vs Seedance 2.0 Fast

Metric Seedance 2.0 Seedance 2.0 Fast
Latency Higher Lower
Cost Higher Lower
Input Modes T2V, first/last frames, image/video/audio references, native audio, editing Same core modes
Output Quality Quality-first; 480p–4K Speed/cost-first; 480p–720p
Best Fit Final assets, high-detail output, 4K delivery Drafts, variations, faster iteration

Seedance 2.0 vs Kling VIDEO 3.0 vs Google Veo 3.1 vs Runway Gen-4.5

Metric Seedance 2.0 Kling VIDEO 3.0 Google Veo 3.1 Runway Gen-4.5
Output Resolution 480p–4K 720p–4K 720p–4K 720p
Native Clip Length 4–15 sec 3–15 sec 4、6或8 sec 2–10 sec
Reference Support 9 images, 3 videos, 3 audio files; first/last frames Start/end frames; reusable elements First/last frames; up to 3 images Text or one starting image
Native Audio Yes Yes Yes Yes
Shot & Editing Control Multi-shot prompts, frame guidance, video editing Structured multi-shot, per-shot timing Frame guidance, video extension Prompt-led motion; separate editing tools

Inputs

The content array is required. It can contain text, image, video, and audio items in supported combinations. A text item is optional at the array level; when an item uses type: text, its text value is required.

Input Count Format Notes
Prompt Optional Supports Chinese, English, Japanese, Indonesian, Spanish, and Portuguese. Chinese prompts can contain up to 2,000 characters; English prompts can contain up to 2,000 words.
Image Max 9 JPEG, PNG, WebP, BMP, TIFF, GIF, HEIC, HEIF Each image must be under 30 MB, between 300 and 6,000 px, and within a 0.4–2.5 aspect ratio. Roles include first_frame, last_frame, and reference_image.
Video Max 3 MP4, MOV Each video can be 2–15 seconds; combined duration cannot exceed 15 seconds. Input supports 480p–4K, 24–60 FPS, and files under 200 MB each.
Audio Max 3 MP3, WAV Each file can be 2–15 seconds; combined duration cannot exceed 15 seconds. Each file must be under 15 MB.

Use need_review: true when a reference contains faces or copyrighted IP. Leaving it disabled when review is required can cause the request to fail, while enabling it unnecessarily may increase processing time.

A 4K reference video is an input asset and does not determine the output resolution. Set the output tier separately with the resolution parameter.

Parameters

Parameter Supported Values What It Controls
content Text, image, video, and audio items Prompt instructions, frame roles, references, and source media
generate_audio true, false Whether the model generates audio with the video
ratio 16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive Output aspect ratio
resolution 480p, 720p, 1080p, 4k Output detail and unit price
duration 415, or -1 Output length or automatic duration selection
watermark true, false Whether to add an “AI Generated” watermark
tools web_search Optional web-assisted generation

Use 480p or 720p for lower-cost tests and early iterations. Use 1080p or 4K for selected outputs where detail and delivery size justify the higher price.

Shorter clips are easier to control when one action or camera move matters. Longer clips allow more development, shot changes, or transitions, but increase the output charge. Requests with reference video also include the reference-video duration in the final calculation.

Use a fixed aspect ratio when the delivery format is known. Use adaptive when the source media should guide the frame shape.

Pricing

Seedance 2.0 has separate rates for requests with and without reference video.

Resolution Reference Video Unit Price 5-Second Output Subtotal
480p No $0.077/sec $0.385
480p Yes $0.068/sec $0.340
720p No $0.164/sec $0.820
720p Yes $0.147/sec $0.735
1080p No $0.409/sec $2.045
1080p Yes $0.366/sec $1.830
4K No $0.830/sec $4.150
4K Yes $0.751/sec $3.755

Without reference video:

Total = no-reference unit price × output duration

With reference video:

Total = reference-video unit price ×
(output duration + reference-video duration)

For rows marked Yes, the five-second value is the output subtotal before reference-video duration is added.

A 480p request with a five-second output and a five-second reference video costs:

$0.068 × (5 + 5) = $0.68

A 1080p request with a five-second output and a five-second reference video costs:

$0.366 × (5 + 5) = $3.66

A five-second 4K output without reference video costs:

$0.830 × 5 = $4.15

Generated audio is included in these rates.

Quick Start

The following request combines a text instruction with a product reference image, generated audio, and 1080p output.

curl --fail-with-body --connect-timeout 10 --max-time 60 \
  -X POST https://api.icreat.ai/v1/task/submit/bytedance/seedance-2-0 \
  -H "Authorization: Bearer ${ICREAT_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "content": [
      {
        "type": "text",
        "text": "Keep the product shape, materials, and branding consistent with Image 1. Create a slow camera orbit while the lighting shifts from cool studio light to a warmer premium advertising look. Add subtle mechanical sound and room ambience."
      },
      {
        "type": "image_url",
        "image_url": {
          "url": "https://example.com/product.png"
        },
        "role": "reference_image"
      }
    ],
    "generate_audio": true,
    "ratio": "16:9",
    "resolution": "1080p",
    "duration": 8,
    "watermark": false
  }'

The submission returns a task_id.

Submit request → Receive task_id → Poll status → Retrieve result

Query the task until its status reaches SUCCEEDED, then send the same task_id to the result endpoint.

Output

Seedance 2.0 uses an asynchronous task flow. The submission call returns a task_id; task status and the completed result are retrieved through separate requests.

The detailed result fields are not listed until they have been confirmed from a successful iCreat response. This avoids presenting fields copied from another provider or an outdated response format.

Use Cases

Commercial product and advertising video. Product images can preserve shape, materials, packaging, and brand styling while prompts direct camera movement, lighting, environment, action, and sound. Higher output tiers suit selected campaign assets that need more detail.

Film and television previsualization. Multi-shot prompts, complex action, camera direction, native audio, and longer 15-second outputs can be used to explore scene staging before full production.

Character and narrative scenes. Character images, scene references, dialogue, and generated audio can be combined for short interactions and story-driven sequences. Consistent references help reduce identity and wardrobe drift.

High-resolution campaign assets. The 1080p and 4K tiers fit approved shots that need larger delivery formats or cleaner fine detail than a draft-tier output.

Reference-based transformations and editing. Source footage and supporting references can guide focused changes to a subject, prop, action, setting, or visual treatment without recreating every part of the clip.

Limitations

Seedance 2.0 generates 4–15 second videos. Image, video, and audio references are subject to file-count, duration, dimension, and size limits. When a request contains reference video, its duration is included in the final bill.

Explicit first-frame and last-frame roles are supported by the API, but dedicated controls are not visible in the current Playground screenshot.

Complex scenes can still lose fine detail, physical realism, dynamic energy, or consistency across several subjects. Audio may occasionally distort or miss the intended timing. Results are easier to control when each shot has one primary action, one main camera instruction, and references that agree on identity, style, lighting, and environment.

References containing faces or copyrighted IP may require need_review.

Seedance 2.0 Mini: Generate lower-cost video drafts and social content with faster output for frequent, lightweight production workflows.

Seedance 2.0 Fast: Create videos with shorter processing times for rapid testing, creative iteration, and time-sensitive content production.