Skip to main content
POST
Submit Video Task

Introduction

The submit video task API is used to create a new video generation task. Upon successful submission, it returns a task ID that you can use to query the task status. Important Note: Video generation is an asynchronous task. You need to first submit a task to get a task ID, then poll the task status until it succeeds.

Authentication

string
required
Bearer Token, e.g. Bearer sk-xxxxxxxxxx

Request Parameters

string
required
Model identifier, supported models and features:Sora 2 Series:
  • sora-2 - Supports text-to-video, image-to-video, video-to-video (Remix mode). Also POST /v1/videos — see Sora 2
Google Veo Series:
  • veo-3.0-fast-generate-001 - Text-to-video (first frame mode)
  • veo-3.1-fast-generate-preview - Text-to-video (first frame mode, first/last frame mode)
Ali Wanxiang Series:
  • wan2.7-t2v-2026-04-25 - Text-to-video (multi-shot, custom audio). See Wan 2.7
  • wan2.7-i2v-2026-04-25 - Image-to-video (first frame, first+last, continuation). See Wan 2.7
  • wan2.7-r2v - Reference-to-video (multi-modal refs, voice clone). See Wan 2.7
  • wan2.5-t2v-preview - Text-to-video. See Wan 2.5
  • wan2.5-i2v-preview - Image-to-video (first frame mode). See Wan 2.5
Seedance 2.0 Series:
  • seedance-2-0 - T2V, I2V, multi-modal refs, asset library (asset://). See Seedance 2.0
  • seedance-2-0-fast - Fast (no 1080p). See Seedance 2.0
Doubao Seedance Series:
  • doubao-seedance-1-0-lite-t2v-250428 - Text-to-video
  • doubao-seedance-1-0-lite-i2v-250428 - Image-to-video (first frame mode, first/last frame mode, reference image mode)
  • doubao-seedance-1-0-pro-250528 - Text-to-video (first frame mode)
  • doubao-seedance-1-5-pro-251215 - Text-to-video, image-to-video (first frame mode, first/last frame mode), supports audio generation
  • doubao-seedance-1-5-pro-251215-noAudio - Text-to-video, image-to-video (first frame mode, first/last frame mode), no audio generation
string
Video generation prompt, describing scene actions and settings. Note: Doubao Seedance series models do not require this field, the prompt should be written directly in the text field of the metadata.content array
string
Reference image for image-to-video (supports Base64 or URL format)
integer
default:"5"
Video duration (seconds), different models support different durations
string
default:"720p"
Video resolution: 480p, 720p, 1080p, 4k
string
default:"16:9"
Aspect ratio: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, adaptive (adaptive, only supported by some models)

Model-Specific Parameters

Different models support different specific parameters. Below are detailed descriptions by model series:
string|integer
default:"4"
Video duration (seconds), supports: 4, 8, 12
string
default:"720x1280"
Video resolution, supports: 720x1280 (portrait), 1280x720 (landscape)
string
Reference image (supports URL or Base64 format), for image-to-video
string
Remix mode: Regenerate based on existing video ID (must start with video_)

Usage Examples

1. Text-to-Video (Basic Example)
2. Text-to-Video (Landscape, 8 seconds)
3. Image-to-Video (First Frame Mode)
4. Remix Mode (Video-to-Video)
Important Notes (Doubao Seedance Series):
  • The content array must be placed in the metadata object
  • All parameters are passed through special markers in the prompt text (e.g. --ratio 16:9)
  • Images must be placed in the content array using the image_url type
  • First/last frame mode requires two images, marked with role: "first_frame" and role: "last_frame" respectively
  • Reference image mode requires using markers like [图1], [图2] in the prompt to reference images, with images marked role: "reference_image"
  • doubao-seedance-1-0-lite-t2v-250428 does not support image input and adaptive aspect ratio
  • doubao-seedance-1-0-pro-250528 only supports first frame mode
  • doubao-seedance-1-5-pro-251215 automatically generates audio, suitable for scenes requiring background music
  • doubao-seedance-1-5-pro-251215-noAudio does not generate audio, faster rendering, suitable for scenes requiring post-production audio
  • 1.5 pro series supports text-to-video and image-to-video (first frame mode, first/last frame mode), does not support reference image mode
  • 1.5 pro series resolution limit: Only supports 480p and 720p (does not support 1080p)
  • 1.5 pro series duration range: Supports any integer between 4-12 seconds
  • First/last frame mode: Requires providing two images, marked with role: "first_frame" and role: "last_frame" respectively

Response Example

Query Video Task

Query video generation task status and results

Download Video

Download completed video files (Sora 2 only)