On July 23, Black Forest Labs announced FLUX 3. The headline capability is easy to understand: one generation can produce up to 20 seconds of video with native audio.
It was quickly described in some corners as an open-source answer to Veo and Sora. I think that claim is premature.
At launch, FLUX 3 is still in Early Access. Black Forest Labs has not published API pricing and has not promised when, or whether, FLUX 3 video weights will be released. The company has a real history of open-weight releases, but a company having open-weight products does not make every new model open source.
So there is no honest answer yet to “How much will a 20-second FLUX 3 video cost?” The price sheet does not exist.
We can still ask the more useful economic question: can FLUX 3 reduce the total cost of producing a usable clip?
Cost per generated second is the most misleading number in AI video
AI video services commonly charge per second, per generation, or through credits. The number looks clean. The product a creator needs is not twenty seconds of pixels, though. It is a video that can be published.
A short clip with sound typically carries several layers of cost:
- visual generation
- narration or character dialogue
- lip synchronization
- ambient and event sound effects
- multi-shot continuity
- character and scene consistency
- retries and manual repair
If a model produces only visuals, a cheap generation price can still be followed by a speech model, a sound library, a lip-sync tool, and an editor. Every handoff adds latency, format conversion, and another place for the workflow to fail.
That is why I care less about whether FLUX 3 eventually charges cents per second or dollars per clip. A more useful formula is:
Cost per usable clip = total generation and post-production spend divided by the number of clips that pass review.
For content production, I would go one step further and measure cost per usable second. Ten lottery pulls that leave ten publishable seconds are not the same business as one twenty-second generation that is mostly usable.
Native audio is not another checkbox
FLUX 3 jointly learns images, video, and audio in one architecture. According to BFL, it can generate up to twenty seconds of video with sound from text, image, or video references. It supports text-to-video, image-to-video, video continuation, keyframe control, multilingual dialogue, and multi-shot sequences.
“Native audio” sounds like one extra row in a feature table. In practice, it targets one of the ugliest handoffs in video production.
A modular workflow generates the picture, records or synthesizes speech, applies lip sync, and adds event sounds. All four models may be good on their own and still fight when combined. The character closes her mouth before the line ends. A glass hits the floor and the crash arrives late. The camera moves outside while the indoor reverb keeps playing.
BFL disclosed a revealing number: audio represents less than 0.5% of the tokens in a 720p video with sound. Video prediction, by contrast, accounts for more than 95% of FLUX 3’s total training compute. Audio itself is not the expensive modality. The hard part is making sound and vision follow the same causal structure.
The potential saving from native audio is therefore not merely one fewer TTS call. It is avoiding full-pipeline rework when synchronization fails.
If speech, lip movement, impacts, and environmental changes line up in one generation, FLUX 3 could have a higher API price than a silent video model and still deliver a lower cost per usable clip. If the audio regularly needs replacement, “unified generation” just combines several lottery buttons into one larger lottery button.
Why twenty seconds matters more than it appears
Twenty seconds is not long. It is nowhere near the fantasy of generating a movie from one sentence. For short ads, product demonstrations, character dialogue, storyboards, and B-roll, however, it crosses a practical threshold.
Many models can produce an impressive five-second demo. The trouble begins when the clip must continue. Short segments require more generations and more stitching, creating more opportunities for clothing, lighting, faces, sound, and pacing to drift. Moving from five seconds to twenty is not only four times the duration. It may remove three edits and three groups of continuity risks.
BFL also says visual references can keep characters consistent across sequences that combine into several minutes of content. That claim needs real-world testing, but the direction is clear. BFL is not merely offering a longer clip. It is trying to deliver continuous material with fewer holes for external tools to patch.
In BFL’s preliminary human comparisons using ten-second, 720p videos with sound, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, and Seedance 2.0 and Gemini Omni Flash in 52% each. The reported numbers against Runway Gen-4.5 and Luma Ray 3.2 were 77% and 93%.
These results suggest that the early samples are competitive. They do not establish production economics. BFL itself says the model and evaluation harness are still in development. The numbers that determine the bill are still missing: failure rate, acceptance rate, queue time, concurrency limits, editability, and final pricing.
Do not celebrate an open-source price war yet
BFL’s impact on the image-model market makes an open-weight price squeeze easy to imagine. Open weights can enable self-hosting, competing API providers, quantization, distillation, and deep workflow integration. Once distribution broadens, platform margins become harder to defend.
FLUX 3 has not reached that point.
It is in Early Access. Video pricing is undisclosed. An open-weight plan is undisclosed. The accurate statement today is: a vendor with an open-weight track record has moved a unified audio-video model into early access. That puts pressure on Veo, Sora, Runway, Kling, and Seedance, but the pressure comes from the technical direction and future distribution options, not from an open-source price war that has already happened.
Even if weights are released later, video economics will not collapse as easily as they can for a small language model. BFL says video prediction consumed over 95% of training compute. Inference must also process large amounts of spatial and temporal information. Free weights do not mean free GPUs, and downloadable does not mean practical on a normal workstation.
Open distribution can more realistically cut three kinds of cost:
- Platform margin. Competing hosts compress the premium charged by a closed API.
- Workflow tax. ComfyUI and automation systems can optimize batching, caching, and asset reuse.
- Switching cost. Teams do not have to lock prompts, media, and pipelines into one vendor.
Compute cost remains. Pricing power weakens. That is the credible open-weight threat to Veo and Sora.
FLUX 3 x mimic reveals a larger ambition
The FLUX 3 x mimic announcement on the same day makes the launch more interesting.
BFL and mimic robotics connected the FLUX 3 video backbone to action prediction, creating FLUX-mimic and testing it in production tasks at Audi. Adding action prediction initially reduced video-generation quality by as much as 10%. After 3,500 training steps, quality recovered to its previous level.
That result suggests BFL is not merely building a cheaper video generator. Its bet is that one world model can understand pictures, sound, and action. Content generation simulates the world. Robot control acts inside it. Both use the same foundation.
For creators, this branch does not reduce the bill today. It does explain why BFL treats video training as the core investment. If a video model genuinely learns contact, mass, motion, and causality, the payoff is not only a smoother advertisement. The same foundation can support simulation, automation, and robot learning.
There is also a counterintuitive business possibility. Content generation may not become FLUX 3’s highest-margin business. Creator APIs could drive adoption and feedback, while industrial robotics and simulation pay the larger bills. If that model works, BFL may have room to price the video API more aggressively.
That is a business hypothesis, not an announced BFL pricing strategy.
How I would test whether it really cuts costs
When FLUX 3 becomes broadly available, I would not begin with the advertised generation price. I would measure five things:
- out of one hundred generations, how many can be used without major repair
- the average number of retries per accepted clip
- how often native audio needs new speech, sound effects, or lip sync
- whether one twenty-second clip is cheaper per usable result than four five-second clips
- how much human repair is needed to preserve a character across shots
Those numbers reveal whether FLUX 3 is a cost-saving tool or an expensive demo machine.
I would also run the same task through two pipelines: FLUX 3 as a unified system, and a modular stack consisting of video generation, TTS, lip sync, and sound effects. Both would use the same script, references, and acceptance criteria. The comparison should measure total money, elapsed time, and rework from script to approval, not the price printed beside the Generate button.
This resembles the argument I made yesterday about Opus 5: per-token pricing is giving way to cost per task. Video needs one more change of unit, from cost per second to cost per usable second.
What we can conclude now
How far will FLUX 3 cut the price tag on Veo and Sora? No one can give a credible number yet. BFL has not published one.
What we do know is that FLUX 3 moves the competition from “who produces the prettiest moving image” toward “who can deliver a more complete audio-visual asset in one pass.” If twenty seconds, native audio, multi-shot generation, and reference consistency survive real production, they can remove meaningful assembly and rework costs.
As for open-source price pressure, it is currently an option, not an exercised fact. We should wait for weights, license terms, and official pricing before declaring that Veo or Sora’s platform premium has been broken.
The metric I would watch is less glamorous and more honest: with a fixed budget, how many seconds of video are good enough to publish?