Summary
A comparison of six AI video editing models (Runway Aleph 2, FLUX 3, Gemini Omni Flash, Kling 3.0 Pro, Seedance 2.5, and WAN 3.0) tested on five editing tasks reveals significant differences in how they interpret instructions—some treat footage as source material to edit, while others remake scenes entirely. Object removal is now reliably solved across all models, but more complex tasks like adding objects or replacing backgrounds show varying levels of control, with some models introducing unwanted creative additions.
- Object removal is commodity-level competency; FLUX 3 at 27 cents per 10 seconds performs as well as models costing 15 times more for this task
- Models differ fundamentally in interpretation: some adjust existing elements (shadows move with relighting) while others reconstruct scenes from scratch (camera and character movement changes)
- WAN 3.0 and Kling 3.0 Pro add unapproved elements (cars, hair) that improve plausibility but reduce control, creating brand/continuity risks
- Four models (Seedance 2.5, WAN 3.0, Kling 3.0 Pro, FLUX 3) successfully held character performance while replacing backgrounds; Gemini Omni Flash reshoot scenes instead of editing them
- Testing methodology was deliberate: one generation per model with no re-rolls, making this a first-look assessment rather than a definitive benchmark
What Are Your Ideas Today?
Ask a video model to change the lighting in a shot and you find out very quickly whether it is editing your clip or quietly remaking it. We asked six models to turn an evening scene into daylight. The cheapest one in the test, at about 27 cents per ten seconds, moved the shadows in the right direction. A model costing five times more moved the light source and left every shadow exactly where it was.
That is not a quality gap. Both clips look good on their own. It is a difference in what the model thinks you asked for: one of them treated your footage as a source, the other treated it as a suggestion.
Text-driven video editing is the request we hear most from Workroom users after image editing. Remove the car. Make it night. Turn this into a cartoon. Six models on Wavespeed say they can do it, so we gave all six the same clips, the same instructions, and exactly one take each. Here is what came back, task by task.
What we asked for
Six models: Runway Aleph 2, FLUX 3 (video edit), Gemini Omni Flash, Kling 3.0 Pro, Seedance 2.5 and WAN 3.0. Every other video-edit endpoint we tried on Wavespeed fell out on the first task with output nobody would ship.
Five edits, ordered from easy to unfair:
- Remove an object from the frame.
- Add an object that was never there.
- Replace most of the location behind the subject.
- Change the lighting so the shadows have to move with it.
- Stylize the whole clip into an animated look and keep the character recognizable.
Same source clips, same prompt text, one generation per model per task, no re-rolls and no picking the best of three. That makes this a first look rather than a scored benchmark, and we are saying so up front because the results below are interesting enough without pretending otherwise.
Removing an object: the boring task everyone passes
All six models removed the object cleanly and left the rest of the frame alone.
This is the least dramatic result in the article and the most useful one. Object removal is solved well enough that paying for it is a decision, not a necessity. FLUX 3 does it for roughly 27 cents per ten seconds, and for small objects on a moving background you will not see the difference against a model costing fifteen times more.
If removal is all you need, buy the cheapest option in the list and put the savings into the shot that actually needs help.
Adding an object: the question is what else shows up
All six added what we asked for. The differences live in what arrived alongside it.
WAN 3.0 added a car driving past a window in the background. Nobody asked for a car. It is beautifully rendered, it moves correctly, and it is still a change you did not approve. Kling 3.0 Pro has a milder version of the same habit and likes to add hair.
There is a real argument for this behavior. A model that only does the literal thing often produces a shot that feels pasted together, and WAN’s inventions usually make the frame more plausible, not less. The problem is control. If your clip has to match a plate, a brand guideline or a shot from the same scene, watch the whole frame before you approve the take, not just the part you edited.
Replacing the location: who keeps your take
Four models held the performance while swapping most of the background: Seedance 2.5, WAN 3.0, Kling 3.0 Pro and FLUX 3. Runway Aleph 2 made one small mistake.
Gemini Omni Flash did something more interesting and more disqualifying. The composite looked fine, but the character moved differently and the camera drifted off its original path. It had re-shot the scene instead of relighting the set.
Kling deserves a specific mention here, because it holds camera motion more faithfully than anything else in the group. For handheld footage and moving shots, that is the property you notice first when it is missing.
Changing the light: the task that splits the group
This is the one that separates editors from generators, and shadows are the tell.
FLUX 3, WAN 3.0, Runway Aleph 2 and Seedance 2.5 changed the light and moved the shadows with it. Kling 3.0 Pro did not react at all; the relight simply failed. Gemini Omni Flash moved the light source and left the shadows pointing at the old one, which is the visual equivalent of a continuity error your viewer will feel without being able to name.
A model that moves shadows correctly is reasoning about the scene as a space. A model that repaints colors is reasoning about the picture as a surface. Both are legitimate pieces of technology. Only one of them can be trusted with a shot that has to cut next to the original.
And note which model got this right for 27 cents. Price told you almost nothing here.
Stylizing the whole clip: only one model kept the face
The hardest task in the set, and the only one where the field genuinely separated.
WAN 3.0 was the only model that came out of a full stylization with the same person on screen. Seedance 2.5 got there too, but it needed extra prompting to hold the identity. Runway Aleph 2 produced a good-looking animated clip starring someone else entirely. FLUX 3 went simple and flat. Kling 3.0 Pro stayed weak.
If your project is a brand character, a presenter or anyone the audience is supposed to recognize, this task is the only one on the list that matters, and the ranking is short: WAN 3.0, then Seedance 2.5 with a longer prompt.
Which model is best for what
| Model | Price per 10s | Best at | Watch out for |
| FLUX 3 | ~$0.27 | Cheap iteration, follows instructions, relights correctly | Slow-drifting “boiling” noise in every frame |
| Gemini Omni Flash | ~$1.35 | Good-looking output on its own terms | Regenerates motion and camera instead of editing |
| Kling 3.0 Pro | ~$1.68 | Holding camera motion through an edit | Relighting does not work at all |
| WAN 3.0 | ~$2.70 | Stylization that keeps the character | Invents objects that were not in the source |
| Runway Aleph 2 | ~$2.97 | Reliable removal, addition, relighting | Loses the character in stylization |
| Seedance 2.5 | ~$3.96 | Every task in this test | Price, by a wide margin |
Prices are what one run costs on our test clip, eight seconds at 720p, at each model’s September 2026 list price. They move often.
The numbers behind that table come from a wider shortlist and a slightly longer task list than this article walks through. Sixteen engines got the first task and six survived it. Every run used the same eight-second 720p clip and the instruction text alone, with no reference images, masks or keyframes even where the model accepts them, which is why the reference-driven strengths of Runway Aleph 2 and Seedance 2.5 never got a chance to show. A sixth task, swapping a knitted cardigan for a denim jacket, went in alongside the five above. Measured against the untouched areas of the frame, background replacement turned out to be the most destructive request in the set, changing the part of the shot nobody asked about by ten to sixty times the baseline noise, with Gemini Omni Flash and Runway Aleph 2 the worst offenders. Three models also quietly rewrite your format: FLUX 3 returns 1248×704 instead of 1280×720, Kling ignores the requested resolution and hands back 1920×1080 every time, and WAN 3.0 stretches the clip to nine seconds at 30fps.
The part that argues with the table
The cheapest model in the test has the most irritating defect, and it still might be the right choice for most of your work.
FLUX 3 puts a slow, blotchy noise over the entire frame on every task. On a replaced wall or a clear sky it is obvious once you see it. On a three-second social cut with a moving subject and grain already in the footage, several people will not see it at all. Colors stay stable between frames, which is the harder problem, and the model obeys instructions that models costing ten times more ignored.
So the honest version of the verdict is not “Seedance wins.” It is that Seedance 2.5 wins the tasks where being wrong is expensive, and FLUX 3 wins everything else on price by an order of magnitude.
Gemini Omni Flash deserves the same fairness. It failed this test because we were testing editing, and it is not an editor. As a model that takes a clip and a description and makes something new and attractive out of both, it is good and it is cheap. Just do not put it anywhere near a shot that has to match what was already filmed.
What we would actually ship
For a text-based video editing tool, one model is the wrong answer. Draft edits, standard edits and final-quality edits have different price ceilings, and this group splits along them cleanly:
Draft tier: FLUX 3 at about $0.27 per ten seconds. Iterate fast, accept the noise or clean it in post.
Standard tier: Kling 3.0 Pro at a mid-range price. Good on everything except relighting, and the best in the set at holding camera motion.
Final tier: Seedance 2.5 at about $4 per ten seconds. Any task, with the right prompt.
Routing each request to the cheapest model that can handle it costs a fraction of running everything through the universal one. Specialization is cheaper than generalism, and in this test it is also more accurate, because the cheap model was better at relighting than two models above it.
What this test does not tell you
One take per model per task means a model that invents details might invent different ones tomorrow. Five tasks on a small set of clips leave out long takes, dialogue scenes, and edits that touch several objects at once. None of this went through a blind panel yet, which is the next step for whichever models survive a product decision. Treat the table as a shortlist, not a scoreboard.
The question worth asking next
The gap between the 27-cent model and the four-dollar one is no longer about who understands the instruction. Both of them moved the shadows. It is about noise, and noise is the kind of problem that gets fixed in a point release.
So how long does a four-dollar tier stay worth four dollars? That is the number we will be watching in the next round, and it is a better question than which model is winning this month.