Summary
A comparison of Seedance 2.0 and MiniMax H3 video models across five scenarios and three input modes reveals that MiniMax H3, priced one-third as much, better adheres to detailed specifications and storyboards, while Seedance 2.0 tends to reinterpret briefs rather than follow them precisely. The key finding shifted the authors' default choice away from their months-long preference for Seedance 2.0 toward the more cost-effective MiniMax H3 for most production work.
- MiniMax H3 shoots storyboards in specified order; Seedance 2.0 elaborates and reinterprets briefs (e.g., adding a second head to a titan)
- Both models produce competent, coherent video output, but the difference lies in adherence to specifications versus creative interpretation
- MiniMax H3 costs one-third as much as Seedance 2.0, making it significantly more cost-effective for production pipelines
- The comparison tested five realistic scenarios (UGC selfie, car chase, cartoon, anime fight, news opener) across three input modes (text only, opening frame, full storyboard)
- Production usability and brief compliance are more important evaluation criteria than individual frame aesthetics when choosing video models
Two of the pack of strongest video models available today run against each other across five scenarios and three input modes. The one that follows a storyboard costs a third as much as the one that rewrites it.
The storyboard we handed to both models was a six-panel anime sequence: a lone hero on a broken rooftop, a stone titan as tall as the towers around it, and a single white slash passing through the titan’s wrist. MiniMax H3 shot the panels in the order they were drawn. Seedance 2.0 treated the first panel as a point of departure and elaborated on it, and in the clip it returned, the titan had two heads.
Neither outcome is a failure in any strict sense, since both models produced ten seconds of competent, well-lit, coherent video from the same input. But the distinction between a model that shoots your plan and a model that improves on it turned out to be the most consequential thing we learned from a fifteen-pair comparison, and it led somewhere we had not anticipated when we set the test up. The model we have been treating as our default for months is not the one we should be using for most of our work, and it has been costing us three times more than the alternative.
Why we ran the comparison
Seedance 2.0 has been the default video model at Everypixel since this spring. That was a reasonable decision at the time and it remains a defensible one now, because the model handles the majority of what we ask of it without argument. The problem with settling on a default, though, is that it removes the reason to keep looking, and we had not seriously re-examined the choice in several months.
In that period MiniMax released H3, which is positioned against the same class of work and priced considerably lower. We had no basis for judging how the two compared beyond the usual promotional material, and since the question directly affects what we spend on every generation, we decided to answer it ourselves rather than wait for someone else to publish something.
How the comparison was built
We chose five scenarios that between them cover most of what we generate in practice: a handheld UGC selfie on a city street, a night car chase across wet neon-lit roads, a stylized 3D cartoon involving a small robot and a cookie jar, a cel-shaded 2D anime fight, and a 3D broadcast news opener with rendered chrome typography. Each scenario was specified in detail, with the ten seconds broken down beat by beat and every camera position described.
Each scenario was then run three separate ways, because in our experience the result depends nearly as much on how you present the brief as on what the brief contains. In the first mode the model received text alone and had to invent every visual element itself. In the second it received an opening frame that we had generated in advance, which fixes the characters, the palette, and the framing before the animation begins. In the third, it received a six-panel storyboard sheet as a reference, with the panels numbered, and was asked to shoot them in that order.
Five scenarios, three input modes, two models, which comes to fifteen head-to-head pairs. Four of them are discussed below. The remaining eleven produced results that were closer together, and the four we have chosen are the ones where the differences are visible without stepping through the footage frame by frame.
What we were actually looking for
The obvious question to ask about a video model is whether the output looks good, and it is also the least useful question available at this point in the technology’s development. Most current models produce attractive individual frames, which means that anyone choosing between them on the basis of a showreel is being shown precisely the footage that was selected to look attractive.
The criteria that determine whether a model is usable in a production pipeline are less photogenic. We wanted to know whether each model shot what the brief specified rather than what it implied the brief was getting at and whether anything appeared in the footage that had never been requested. We wanted to know whether physical objects held their shape under motion, particularly mechanical ones, which are where generative models still tend to come apart. We wanted to know whether a sequence of shots read as a single continuous event or merely as several plausible shots placed in succession. And we wanted a realistic sense of how many attempts were required before a clip could be sent to a client, since that number governs both cost and schedule more directly than any quality judgment does.
None of these properties are apparent from a highlight reel, and all of them become apparent very quickly when a delivery date is involved.
Four pairs, examined
In each comparison below, MiniMax H3 is on the left and Seedance 2.0 is on the right, and both models received identical input.
A UGC selfie, where the two models are effective
[VIDEO: A-S1 — text-only input, 9:16, 10s] MiniMax H3 and Seedance 2.0 generating the same UGC selfie clip from a text prompt
The brief here was deliberately ordinary: a woman in her late twenties filming herself on a phone held at arm’s length while walking down a sunlit European street, a paper coffee cup in her free hand, the whole thing captured as a single continuous take with no cuts. This is the register most commercial video work actually occupies, and it is the case that matters most in aggregate even though it is the least interesting to watch.
Both models handled it. Differences exist if you go looking for them at the level of individual details, and in ordinary viewing they are not perceptible. We consider this the most important of the four results, because it establishes that the conclusions drawn from the remaining three do not come at the expense of the category that accounts for most of our volume.
A night chase, where both models fail in different ways
[VIDEO: A-S2 — text-only input, 16:9, 10s] Seedance 2.0 losing car geometry in a night chase scene next to MiniMax H3
The chase brief is considerably more demanding, asking for a black sedan pursued through wet neon streets across six distinct camera setups in ten seconds: a high wide establishing shot, an extreme close-up on a wheel at road level, an interior shot over the driver’s shoulder, a macro on his eyes as he spots a gap between buildings, a tracking shot as the car turns into it, and a crane upward to resolve the outcome. This was run from text alone, with no opening frame to anchor the geometry.
Both models struggled, and in our runs roughly three attempts were needed before either produced something usable. The failures, however, are not the same failure. On the Seedance side the car itself does not survive: the vehicle loses its structure partway through and becomes a shape that the model evidently cannot resolve into a coherent object, which is the kind of defect no amount of grading or trimming will rescue. MiniMax keeps the car intact throughout but loses the connective logic between shots, so the six setups never quite assemble into one continuous chase.
Of the two, Seedance comes closer to a usable result on this category, and it is the clearest advantage it holds anywhere in the test.
A 3D cartoon specified second by second
[VIDEO: A-S3 — text-only input, 16:9, 10s] MiniMax H3 and Seedance 2.0 animating a 3D robot reaching for a cookie jar
This scenario describes a small robot attempting to reach a cookie jar on a high shelf, stacking a book and a mug into an unstable tower, climbing it, and being defeated when the tower tips and a cookie lands on its head. The prompt is written as a timed breakdown with six beats, each carrying its own camera instruction, and it leaves very little room for interpretation.
MiniMax performs at its best here. It respects the timing, and the camera moves it produces read as deliberate choices rather than as drift, which is not something these models reliably manage in stylized animation.
Seedance executes most of the sequence correctly and then introduces material that was never requested, including keys, disembodied hands, and other objects that arrive in frame from outside the described scene and remain there for the rest of the clip. Each of these requires either a regeneration or manual removal, and because they tend to sit in the periphery of the frame rather than at its center, they are the sort of thing that gets noticed after a clip has already been shown to a client.
A six-panel storyboard, and the extra head
[VIDEO: C-S4 — six-panel storyboard sheet as input, 16:9, 10s] Seedance 2.0 adding an extra head to the titan while MiniMax H3 follows the storyboard

Both models received the same numbered storyboard for the anime sequence described at the start of this article, along with an instruction to use it as the reference for shot order, framing and style.
MiniMax worked through the sheet and returned the sequence as drawn. Seedance took the opening panel as a basis and developed it, adding the second head to the titan along with other elements that do not appear anywhere in the reference. This was consistent rather than incidental: across the storyboard runs, Seedance treated the panels as artistic direction rather than as a specification, reordering beats and introducing shots wherever it apparently judged that the sequence would benefit.
Improvisation against execution
A single behavioral difference accounts for all four results, and it is worth stating carefully because it is easy to mistake for a quality gap when it is nothing of the kind. Seedance fills in whatever the brief leaves unspecified, drawing on its own judgment about what the scene requires. MiniMax stays much closer to the literal content of the instruction it was given.
When the brief is thin, Seedance’s tendency is a genuine advantage, and it explains why the model gets further into the chase sequence than MiniMax does. A ten-second action scene contains far more information than any prompt realistically encodes, so the model must invent most of what appears on screen, and Seedance’s inventions are generally plausible enough to hold together.
The same disposition becomes a liability the moment the brief stops being thin. A storyboard exists precisely because the creative decisions have already been made and someone now needs them executed, so a model that continues to improve on those decisions converts every generation into an inspection task, in which the operator must identify what changed before anything can be approved. Whether that trade favors one model or the other depends entirely on which kind of work dominates a given pipeline. In ours, the specified case dominates, which is why this result matters more to us than the chase result does.
The cost difference, and what it does not account for
Every ten seconds of Seedance output costs us three times what the equivalent ten seconds costs on MiniMax, at the settings of APIs we run in production.
That figure requires an important qualification, without which it would be misleading. We do not call Seedance through its direct endpoint, because the route we use handles work involving human faces more reliably for our purposes, and reaching the model that way carries a markup. The threefold difference, therefore, combines two separate things: an actual difference in what the two models charge and a premium we pay for reasons unrelated to model quality. We have not attempted to separate the two, and we are not going to publish a breakdown we did not measure. Anyone running a similar comparison should price the endpoint they can realistically use rather than the one on the published rate card, since those are frequently not the same endpoint.
What we can state without qualification is the effect on the invoice. On the route available to us, moving this work to MiniMax removes roughly two thirds of the cost, and across the categories that account for most of our generation volume we cannot identify what we would be giving up in exchange.
The case for staying with Seedance
It would be a misreading of these results to conclude that MiniMax is simply the better model, and the honest summary is more qualified than that.
Seedance wins the hardest category in the set. Physical complexity and multi-plane action remain the point at which generative video most visibly breaks down, and it breaks down less on Seedance than on MiniMax. The improvisational habit that undermines its storyboard work is the same capability that rescues an underspecified prompt, and in practice a great many prompts are underspecified. A team working primarily on exploratory or loosely defined material would reasonably reach the opposite conclusion from ours.
We also excluded Seedance 2.5 from this comparison deliberately. It accepts up to thirty reference images and generates up to thirty seconds in a single pass, and MiniMax H3 does neither, so running them against each other would have demonstrated nothing beyond the fact that one model has capabilities the other lacks. For long-form work, or for anything requiring extensive visual references, 2.5 remains in a category of its own, and this article does not speak to it.
The conclusion is narrower than a verdict on which model is better. MiniMax is close enough to Seedance across most of what we generate, meaningfully more disciplined where discipline matters to us, and substantially cheaper on the endpoint we are able to use.
What self-hosting might change
MiniMax H3 is open enough to deploy on your own infrastructure, which is the part of this that interests us most and the part we can say least about. A model that already wins on API economics while performing comparably to a stronger competitor becomes a different proposition once the hardware is owned and the weights can be fine-tuned on a team’s own material.
Nobody here has a credible estimate for how far the cost per generation would fall in that scenario, and until the deployment itself is working properly, that estimate is not worth attempting.
Which model to use
For work that is storyboarded or otherwise specified in advance, MiniMax H3 is the better production choice as of August 2026. It shoots the plan as drawn, it matches Seedance on everyday UGC content, it is stronger on 3D character animation, and on the endpoint we can reach, it costs roughly a third as much.
For work that is loosely defined or heavy on motion and physical interaction, Seedance 2.0 remains the better option and is worth the difference in price. It handles complexity more gracefully, and it will supply what the brief omitted.
What we find genuinely notable about this outcome is how little of it turned on output quality. Seedance is the stronger model on the dimension that most public comparison of video models is concerned with, and we are moving a substantial part of our pipeline off it regardless, on the basis of instruction-following and price.