Seedance 2.0 Mini vs Fast vs Standard: What's the Difference? We Compared Them With the Exact Same Prompt
A New York chase scene and a monster showdown at the Arc de Triomphe — we ran two identical prompts through all three models to compare their action staging side by side, plus the prompt formula that gets the most out of Seedance 2.0.

One of the questions we hear most often from Toonkit creators is: "There are several models — which one should I pick?" Seedance 2.0 comes in three flavors — Mini, Fast, and Standard — and the names alone don't tell you much about how differently they perform. So for this post, we fed the exact same prompt into all three models and placed the results side by side. We're also publishing the full prompts we used, so feel free to try them yourself. At the end, you'll find a prompt-writing formula that raises the quality of your output no matter which model you choose.
The Three Seedance 2.0 Models at a Glance
All three models belong to the same Seedance 2.0 family, but they're built for different priorities.
- Mini — The fastest and lightest model. It reliably captures the core elements of a scene, but its staging is relatively restrained and static.
- Fast — The sweet spot between speed and quality. The amount of motion and detail increases noticeably over Mini.
- Standard — The highest-quality model. It delivers a clear step up in what we'd call "directorial density": camera work, crowd animation, and particle effects.
If we had to sum up the difference in one line: as you move from Mini → Fast → Standard, the action becomes more expressive and the staging more spectacular and detailed. But words only go so far — let's look at the actual results.
Comparison ① The New York Chase: How Well Does Each Model Bring a Crowd and a City to Life?
(from left to right: Mini / Fast / Standard)
In this scene, Minjun — clutching a black briefcase — sprints through downtown New York at dusk with Carter, in a brown leather jacket, chasing close behind. Yellow cabs, digital billboards, wet asphalt, packed sidewalks, and a parkour sequence sliding over a taxi hood and grabbing onto a fire escape — this prompt packs in everything a demanding urban chase scene needs. Here is the full prompt we used.
Emphasize Minjun's black briefcase and the sleeve patch on Carter's brown leather jacket.
Late afternoon in a bustling New York downtown. Low backlight spills between the buildings to the west, spreading across glass facades, while teal and orange light from digital billboards pulses through a volumetric dust layer every 0.5 seconds. Wind funneled between buildings whips the advertising flags on both sidewalks as if tearing them apart, pushing paper cups and plastic scraps toward the roadway. Crowds move constantly between the crosswalk and storefront entrances, and yellow taxis and black SUVs inch forward through slow-moving congestion.
0-3s: Minjun cuts through the crowd. Shoulders lowered, he weaves between people in an S-curve, his black briefcase swinging widely at his side in a 40 cm arc. Billboard light washes over his face in sequence, teal to orange, and a sharp highlight grazes the corner of the briefcase. The coat hems and hair of pedestrians he passes sway a beat late, and scraps of paper spiral up in the vortex at his heels.
3-6s: Carter gives chase, shoving through the crowd. He charges forward, prying open narrow gaps with his broad shoulders, the sleeve patch on his brown leather jacket flashing in the billboard light. His breath bursts into short white puffs in the cold air, and each footfall splashes small droplets off the rain-slicked pavement. Pedestrians he collides with stagger backward off balance, and the advertising flags bend a beat late in the direction of his sprint.
6-9s: Minjun leaps into the road. The instant the light changes, he cuts between the cars without slowing, only changing direction, pulling the briefcase tight against his body with both hands. Brake lights ignite in a red chain reaction, and the glass reflections stretching across a taxi's hood cover his face in fractured patterns of light. Tires scrape the wet road, raising thin gray smoke, and passengers inside the vehicles lurch forward, bracing their hands against the windows.
9-12s: Minjun slides over the taxi's hood. He springs up with his feet together, plants both hands on the hood, and throws his body low — the black briefcase swings a half-turn through the air before snapping back against his ribs. His palm print and streaks of reflected light smear across the taxi's yellow paint, and Carter's fingertips, right behind him, graze only the air 30 cm behind Minjun's back. The taxi slams to a stop, its front bumper dipping, and the silhouettes of the two figures tear past the surrounding car bodies to the rhythm of horns and shouting.
12-15s: Minjun kicks off a fire hydrant and grabs onto a fire-escape railing. Rounding the corner, he plants one foot on the red hydrant, twists his body, and seizes the steel railing of the side alley. The briefcase strap snaps taut with a jolt, and a thin afterglow of blue billboard light runs along the steel railing. Below, Carter reaches out and grasps only empty air, coming to a halt as dark shadow settles over his jacket sleeve and jawline. Alley dust settles slowly, and the steel stairs tremble faintly in time with his furious footsteps.
Camera: For 0-3s, start with a handheld tracking shot 1.5 m behind Minjun's right side, squeezing through the crowd's shoulders alongside him → for 3-6s, skim past Minjun's swinging briefcase and pull left to catch Carter's head-on pursuit → for 6-9s, as Minjun leaps into the road, the camera vaults low over the curb and slides between lanes → for 9-12s, glide 30 cm above the taxi hood, connecting Minjun's hands, the briefcase, and Carter's fingertips in a single breath → for 12-15s, wrap around the corner and crane up from the hydrant to the railing, ending with a slow fade to black behind Carter's upturned face. One continuous move throughout.
Audio: Low urban ambience and distant subway rumble → startled breaths and the rustle of clothing in the crowd → the briefcase's metal buckle clicking like a rhythm → traffic-signal chime as horns rise a half-step at a time → tire skids, brake pressure hiss, the dull friction of sliding across the taxi hood → the short Doppler whoosh of Carter's fingertips slicing empty air → the clatter of the steel railing → sudden silence → only the residual tremor of the alley stairs and the faint sound of settling dust.
IMAX, cinematic 3D animation, volumetric light, wet-road reflections, 2.35:1, tense urban chase, epic action tone, do not cut arbitrarily, one shot.
Viewed side by side, the differences are unmistakable. Mini renders the protagonist and the backdrop competently, but the surrounding characters move with noticeable restraint. Move up to Fast and the crowd actually starts to react — heads turn, people break into a run, and the frame gains a real sense of speed. Standard goes further still: a tracking shot that stays glued to the character, a hair's-breadth encounter with a taxi crossing the crosswalk, neon reflections shimmering on the wet pavement — the result feels like a scene from a theatrical animated feature. In the hydrant-to-fire-escape sequence specified late in the prompt, Standard also delivers the most dramatic high-angle composition and the most expressive acting from Carter.
Comparison ② The Monster Showdown at the Arc de Triomphe: Scale and Effects
(from left to right: Mini / Fast / Standard)
The second test is a scene of massive scale: Paris at sunset, where a giant gorilla and a red dragon clash in front of the Arc de Triomphe — pure kaiju-movie territory. The prompt demands a full dramatic arc within 15 seconds: a shockwave-generating ground pound, a vertical climb up the monument, a mid-air collision, and a crash landing on top of the Arc.
The circular plaza in front of the Arc de Triomphe in Paris at sunset. The 50 m tall Arc anchors the center of the avenues radiating in every direction, and vehicles squeeze into the outer lanes to avoid the impact, tracing tiny trajectories like a colony of ants. The low-hanging sun rakes the Arc's sculpted reliefs with orange backlight, and dust and fallen leaves on the plaza floor curl upward in circles from the pressure differentials created by the movements of the colossal creatures. The genre tone is fantasy kaiju action where golden sunset collides with red flame.
0-3s: The giant gorilla stands centered beneath the Arc, feet planted wide, gazing up at the sky. A short burst from his nostrils ripples the fur on his chest; the five fingers of his right hand curl slowly closed as his shadow thickens on the stone ground. In the next instant his fist drops vertically from a height of 1 m and slams into the plaza floor — the shockwave spreads in a ring, rattling the windows of six nearby vehicles. Ground dust leaps 40 cm into the air, and without lowering his head, the gorilla lifts only his eyes to glare at the red dragon in the sky.
3-6s: The red dragon, circling 80 m overhead, spreads its crimson wings to a 12 m span and dives. As the wing membranes press down on the air, light splits along the edges of the clouds, and the sunset reflection flashes like a blade off the tips of its horns. The moment the dragon skims past the top of the Arc with 3 m to spare, compressed wind slams down into the plaza, sweeping the dust and leaves around the gorilla to one side. The fur on the gorilla's shoulders flattens in the headwind and rises again — he doesn't retreat a single step, only planting his left hand on the ground to lower his center of gravity.
6-9s: As the dragon banks and comes down again, the gorilla charges toward the Arc's pillar. His first step onto the stone stairs sends some twenty fragments flying off the edges; with his second step he stamps into the wall face and hauls his enormous body vertically upward. As his fingers grip the gaps between the sculpted reliefs, lime dust streams down between the hairs on the back of his hands. At the Arc's mid-height he arches backward, then hurls himself 15 m through the air at the dragon — the dragon throws its neck back, gathering orange heat deep in its throat, and releases a short burst of flame. The air warps like a lens where the fire passes.
9-12s: In mid-air, the gorilla twists his left shoulder and lets the flame graze past with 30 cm to spare. The tip of the blaze singes the fur on his arm, black ash scattering behind him, and he only presses his lips shut as if swallowing the pain. In the same motion, one arm seizes the scales at the dragon's nape — the wings buckle for an instant, and the dragon's tail strikes the top of the Arc, sending stone dust cascading down like a waterfall. The gorilla's other fist hammers down onto the scales, red sparks and golden fragments erupting at once, and the recoil of the impact sends the two bodies spinning as they plummet onto the Arc.
12-15s: As the two massive bodies crash onto the top of the Arc, the entire structure shudders low, and dust rises slowly from the seams between the statues. The dragon folds one wing and pushes itself upright; the gorilla, one fist planted into the surface, slowly raises his head. Their breath mingles in hot steam and low snorts, and the sun, caught on the horizon, leaves only their silhouettes as enormous black shapes. The small fragments on the surface stop vibrating and settle one by one — the two hold their charging stances, frozen, and in the brief silence the frame fades to black.
Camera: Open on a low-angle wide shot facing the Arc, framing the gorilla's fist and the dragon in the sky together → the camera jolts 20 cm with the impact of the fist hitting the ground, then dollies in through the dust particles → at the dragon's dive, a vertical crane up along the top of the Arc, then an orbital move tracing the dragon's trajectory the instant its wings skim past → handheld tracking in close on the stone fragments and the gorilla's fingers as he scales the wall → at the mid-air collision, a 180-degree orbital wrap around the two creatures, then after the fall, end on a high-angle wide of the standoff atop the Arc's silhouette. One continuous move throughout.
Audio: Mid-range traffic noise from the plaza and distant sirens → the low-frequency impact of the gorilla's fist striking the ground and the high-frequency crack of splitting stone → the Doppler whistle of dragon wings tearing the air and low-end wind pressure → the dry percussive hits of stone fragments as he climbs the wall → the sub-bass pressure of the flame burst and the metallic crack of blows landing on scales → the moment they fall, all sound cuts to 0.4 seconds of total silence → only a low reverb circling inside the Arc and the faint sound of falling dust remain, decaying into the fade-out.
Cinematic 3D animation, sunset backlight, volumetric dust and flame light, 2.35:1, grand and ferocious. Do not cut arbitrarily, do not change the characters' appearance, do not change the Arc de Triomphe setting to another location, one shot.
Mini locks in the core composition — the two monsters and the Arc — with reassuring stability. Fast makes the choreography far more dynamic: the gorilla scales the monument while the dragon wheels overhead. And with Standard, the staging climbs to another level entirely. Sparks fly in the brawl atop the Arc, shattered stone fragments and dust clouds fill the air, and the camera pulls out to an aerial view of the entire Parisian cityscape before diving back into the action. Standard was also the model that most faithfully executed the prompt's camera directions ("vertical crane up → orbital → high-angle wide").
So Which Model Should You Use, and When?
Rather than a strict quality ladder, the three models are best understood as tools for different jobs.
- Exploring ideas or drafting? Use Mini. It's the most efficient way to run many prompts quickly and validate composition and concept.
- Make Fast your default. It produces sufficiently dynamic results at a reasonable speed. We recommend doing most of your iteration on Fast.
- Render your finals on Standard. For a video you'll publish, a cut you'll send to a client, or a final export for YouTube or social media, Standard's directorial firepower is what makes the difference.
How to Write Seedance 2.0 Prompts That Actually Deliver
Whichever model you use, the quality of your prompt determines half of your result. The two prompts published above are good examples — look closely and you'll find a shared structure and set of principles.
1. Structure your prompt as: scene setup → second-by-second timeline → camera → audio → style & constraints. This is the skeleton that works best with Seedance 2.0. First, pin down the key details that must stay consistent (e.g., "Emphasize Minjun's black briefcase and the sleeve patch on Carter's brown leather jacket"), then establish the setting and atmosphere, and design the action in ordered time segments — 0-3s, 3-6s, and so on. Keep the camera, audio, and style keywords in their own separate paragraphs. Finally, state your constraints — the things the model must NOT do — such as "do not cut arbitrarily," "do not change the characters' appearance," and "one shot." This goes a long way toward keeping the scene consistent across the full 15 seconds.
2. Separate camera movement from subject movement. Mixing the two is the single most common mistake. As in the prompts above, keep character action in the timeline paragraphs and give the camera its own chronological paragraph — "handheld tracking → crane up → 180-degree orbital" — and you'll avoid shaky, uncontrollable footage. Specific cinematography terms like dolly in, tracking shot, orbital, and high-angle wide work even better.
3. A single lighting description changes quality more than anything else. If you can add only one element to your prompt, make it lighting. "Low backlight spills across glass facades," "teal and orange billboard light pulses through a volumetric dust layer," "the sunset rakes the sculpted reliefs with orange backlight" — in both videos, lighting descriptions like these formed the backbone of the atmosphere.
4. Direct motion with concrete numbers and cause-and-effect. Seedance 2.0 excels at motion, but if you don't specify what moves and how, you'll get static results. Write with numbers, and with the reactions the action triggers in its surroundings: "the briefcase swings in a 40 cm arc," "the shockwave rattles the windows of six vehicles," "the pedestrians' coat hems sway a beat late." This is how you draw out the full directorial power of the Standard model.
💡 Pro tip: You don't have to write all of this yourself
If you looked at those prompts and thought, "There's no way I can write something that detailed…" — that's a completely fair reaction. Crafting a second-by-second timeline, camera choreography, and sound design by hand every time is a heavy lift even for professionals.
That's why Toonkit ships a dedicated Seedance Prompt Wizard. Type in nothing more than a simple idea — Simply list brief sentences describing each scene of two men chasing each other through New York, and the Prompt Wizard will automatically transform them into a structure Seedance understands best—scene setup, timeline-based actions, camera direction, audio, and style constraints. You can get theatrical-quality prompts like the two examples in this post in a matter of seconds — don't miss it.
Wrapping Up
The same prompt can produce dramatically different levels of directorial density depending on the model. And whichever model you choose, a well-constructed prompt is what unlocks its potential. Toonkit gives you access to all three Seedance 2.0 models — Mini, Fast, and Standard — plus a Prompt Wizard that polishes your idea into a Seedance-optimized prompt. Try comparing the three models yourself using the structure we covered today. You might just watch a single idea turn into a theatrical-quality scene.