Pangram verdict · v3.3
We believe that this text is a mix of AI, AI-assisted, and human-written content.
AI likelihood · overall
MixedArticle text · 1,485 words · 12 segments analyzed
Today, we are officially launching Seedance 2.5, the new-generation video creation model. Since the release of Seedance 2.0, we have noticed a shift in what users expect from video creation models: from merely generating a clip to completing a creative work.
Building on the unified multimodal audio-video joint-generation architecture of Seedance 2.0, Seedance 2.5 centers on foundational generation and reference-based generation, delivering major breakthroughs in long-form storytelling, multimodal reference, and editing.
Grounded in real-world use cases, it opens up greater creative imagination and control, and further unlocks productivity.Key highlights include:Up to 30 seconds per generation, with multi-round extensions: Seedance 2.5 can generate high-quality, 30-second audio-video clips in a single pass and supports multiple rounds of extension.
It also improves shot transitions and scene changes for stronger continuity in longer videos, and delivers notable gains in image, audio, and motion quality, resulting in a more natural, polished visual quality than commonly seen in AI-generated video.
As a result, users can produce high-quality multi-minute content with a consistent audiovisual language, bringing a complete story to life in one take.Fully upgraded multimodal referencing: Users can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. The model also strengthens a range of reference capabilities, including clay render, motion, and creative references, enabling it to better grasp the creator's intent and realize complex ideas that span multiple subjects, scenes, and shot changes.More precise and stable editing capabilities: Seedance 2.5 offers timestamp-level control for targeted editing of audio and video content, notably improving efficiency and controllability. The model also enhances advanced editing features, such as green screen, camera perspective, and reference-based editing, to meet the rigorous demands of professional, complex fields like film and advertising.With advancements in long-form storytelling, multimodal reference, and editing, Seedance 2.5 goes beyond longer single-pass video generation. The model better understands creative intent and delivers the journey from idea to finished video with greater control. Now, we'd like to invite you to watch a short creative film, produced end-to-end by Seedance 2.5.您的浏览器不支持视频播放。Today, Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro, and other platforms, with API access coming soon via BytePlus ModelArk. We invite you to give it a try and share your feedback.Project homepage:https://seed.bytedance.com/seedance2_5Access:Jimeng Web -> Video Generation -> Select Seedance 2.5Doubao Pro -> Video Generation -> Select Seedance 2.530-second long-form storytelling with multi-round extensions: presenting complete stories in a single passSeedance 2.5 extends single-pass video generation from 15 to 30 seconds and further strengthens its storytelling in longer videos. Within 30 seconds, the model can organize multiple logically connected shots so that a story unfolds through setup, development, turning points, and resolution, rather than simply extending a single moment. For example, in a one-take clip of a singer's stage performance, the model portrays the full story of the singer interacting with staff in the dressing room, then walking through the backstage corridor, meeting the dancers, and stepping onto the stage with them for the performance, instead of only the moment of walking on stage.您的浏览器不支持视频播放。T2V prompt: One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor.
The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.Thanks to the model's multi-round extension capability, users can smoothly append subsequent shots to existing video outputs.
Throughout the extension process, it maintains the consistency of main characters, environments, and narrative pacing. This allows users to output videos lasting several minutes at once, reducing the effort required to split clips, repeatedly splice footage, and fix transitions.您的浏览器不支持视频播放。R2V prompt: Extend the video. Continue from the visuals and subjects in @Video 1 and generate another 30-second clip, keeping the character subjects, scene, visual style, and sound effects consistent. The little boy runs along the train carriage holding a soccer ball. When the subway stops, the side door opens and he immediately dashes out, with the male lead chasing after him. The two run across the platform and out onto the street, startling passersby and vehicles along the way. The male lead finally catches up and grabs him. The boy looks up, aggrieved. The male lead's anger slowly fades; he pats the boy's head and shows a helpless smile.In terms of visual presentation, the model achieves smoother transitions between camera movements.
The main subject remains stable across multiple cuts, and the audio and visuals remain in sync, resulting in highly coherent long-form videos. For example, in a Peking Opera scene, the camera executes a graceful circular pan following the lead actor's flowing sleeves, while the main subject and background remain entirely consistent.
The swinging of the sleeves forms natural arcs in the air, closely adhering to real-world physics.您的浏览器不支持视频播放。R2V prompt: 16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot from @Image 1, following her water sleeves through the arc.
She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side in a combat exchange.
Consort Yu stands slightly behind and to the side of the Overlord, weaving in water-sleeve movements to set softness against strength. The camera slowly pulls back from a medium-close shot of the warrior to a full stage view. At the end, all three face the audience and strike a synchronized Peking opera finale pose.Additionally, to address the overly artificial look often seen in AI-generated videos, Seedance 2.5 systematically optimizes details such as object textures, skin and eye features, lighting, and color saturation. The model also minimizes uncontrolled occurrences in subtitles and background music, delivering final products that closely resemble the cinematic quality of live-action footage.Comprehensive upgrades to multimodal reference, bringing greater control to complex creative tasksSeedance 2.5 further strengthens its multimodal reference generation capabilities. It allows users to input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. A larger volume and wider variety of references can better capture the user's intent, producing complex videos with more subjects, richer scenes, and more flexible camera work.The model comprehensively understands elements such as visual composition, scenes, styles, characters, and props across all materials, applying them to the video generation process as instructed. Even in complex scenarios like multi-character shots or group storytelling, it can preserve the appearances and voices of multiple characters while keeping each subject's characteristics stable.您的浏览器不支持视频播放。R2V prompt: A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting and shadows, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly follow @Image 5. Reference @Images 6 to 10 for the rest of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the audience seating.
The lead vocalist walks from center stage toward the front edge. The pianist is positioned by the piano. The orchestra is arranged on both sides and toward the rear. The choir stands at the back of the stage. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera naturally moves across the violin, cello, and orchestra as they perform together, with the violin feeling bright and the cello warm. In the latter part, the choir joins in. The lead vocalist briefly makes eye contact with front-row audience members, who respond with a smile and a slight nod.