Flux 3 generates 20-second videos with native audio
Black Forest Labs releases Flux 3, a multimodal foundation model capable of generating up to 20-second videos with native audio, marking the company's first model to unify image, video, and audio generation.
Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available.
The company ultimately wants to build a world model and is already testing Flux 3 on robotics tasks. The article Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs appeared first on The Decoder.