Freiburg, Germany-based AI research lab Black Forest Labs (BFL) has dramatically expanded its FLUX family of generative models with the launch of FLUX 3, a groundbreaking multimodal frontier model capable of understanding and generating not only images but also combined audio and video clips up to 20 seconds from a single prompt. Crucially, BFL has architected FLUX 3 to extend this unified capability to the realms of robotic vision and action, signaling a significant step towards AI systems that can perceive, predict, and act across both physical and digital environments. This release marks BFL’s inaugural public foray into video generation, building upon its established reputation in image synthesis.
Unlike previous approaches that often stitch together separate models for image, video, and audio, FLUX 3 has been jointly trained across these diverse modalities. This fundamental architectural distinction is at the core of BFL’s vision, aiming to redefine how enterprises conceptualize and deploy AI. The company posits that creative generation, simulation, computer interaction, and robotics are not disparate applications but rather interconnected facets of a singular "visual intelligence" – models that possess the capacity to comprehend, forecast, and operate within the complexities of the real world and digital spaces.
FLUX 3 will be rolled out across four distinct product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the forthcoming open-source FLUX 3 Dev. Both FLUX 3 Video, which includes optional native audio generation, and FLUX 3 Action are now entering a gated "Early Access" program. While any interested party can apply, BFL will be conducting an approval process. Currently, there is no public API access available through BFL or its partners. However, the company anticipates the rollout of FLUX 3 Image in the coming weeks, with general availability to follow. This staggered release strategy mirrors the cautious launch approaches adopted by other leading AI frontier labs, such as Anthropic and OpenAI, though those instances were often attributed to security concerns or governmental directives, rather than BFL’s current phased market entry.
A notable omission from the initial FLUX 3 announcement is detailed pricing information, specific production service-level commitments, or comprehensive evaluation methodologies and benchmarks. Enterprise buyers will need to await further details to conduct thorough total cost of ownership calculations or independently verify the video comparison results. Furthermore, FLUX 3 is not launching with downloadable weights or an open-source license. BFL has indicated that faster, open-weight versions are slated for release later this year. The technical blog specifically highlights FLUX 3 Dev as offering "open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction," a considerably broader scope than any previous FLUX Dev release, which were exclusively focused on image generation. This delay in the release of open-weight models, a cornerstone of FLUX’s past adoption and developer community engagement, may be a point of disappointment for developers accustomed to immediate local deployment options.
FLUX 3 Claims Superiority in Video Generation Benchmarks, but Lacks Key Enterprise Adoption Data
Black Forest Labs has presented preliminary benchmark comparisons for FLUX 3, though the company stresses these are qualified as preliminary and full results with methodology will be published closer to general availability. In early head-to-head preference testing involving 10-second, 720p text-to-video clips with audio, FLUX 3 reportedly outperformed several competitors. BFL states it was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52% of comparisons.
However, a significant caveat accompanies these promising figures. BFL itself labels the chart displaying these results as a "preliminary evaluation of an early FLUX 3 candidate." This means the published numbers reflect a pre-release checkpoint, not necessarily the performance of the model currently entering early access. While this could indicate future improvements, it also means the data doesn’t precisely measure the model that end-users will experience.
The competitors against which FLUX 3 showed the most dominant results, Luma Ray 3.2 and Runway Gen-4.5, are established products but not necessarily the current leaders in independent video generation rankings. The comparisons against Seedance 2.0, which resulted in a 52% preference rate, are less impactful. Seedance 2.0’s international rollout was indefinitely postponed due to legal threats from major media corporations over alleged copyright infringement, leaving its current market relevance uncertain. A tie with a product facing such significant headwinds offers little competitive advantage.
The comparison with Google’s Gemini Omni Flash, also at 52%, is considerably more significant. Gemini Omni represents a direct analogue to FLUX 3’s ambition: multimodal input, video and audio-aware creation, and conversational editing. According to BFL’s own metrics, the two models are indistinguishable in terms of 10-second text-to-video quality. Google’s advantage here lies in the general availability of Omni via its Gemini API, priced at approximately $0.10 per second of generated 720p video, translating to roughly $1.00 for a 10-second clip.
A regional constraint for Gemini Omni Flash might offer a strategic opening for BFL, particularly in its home market and adjacent European regions. Omni Flash users in the European Economic Area, Switzerland, and the United Kingdom cannot currently edit uploaded video content; they are restricted to editing video generated by the model itself. This limitation could be a deciding factor for European enterprises seeking to leverage generative AI for editing existing footage.
For enterprises evaluating video generation models, a comparative overview of current offerings reveals key differentiators:
| Model | Max single-generation duration | Max resolution | Key constraints | Price per 10-sec clip (720p) | Price per 10-sec clip (1080p) | Price per 10-sec clip (4K) |
|---|---|---|---|---|---|---|
| FLUX 3 Video | 20 seconds | Not stated; 720p eval | Early access; no SLA/pricing | Not announced | Not announced | Not announced |
| HappyHorse 1.1 | 15 seconds | 1080p | No 4K; closed weights | Not published | Not published | n/a |
| Veo 3.1 | Per-second billing | 4K | Supports clip extension; preview | $4.00 | $4.00 | $6.00 |
| Veo 3.1 Fast | Per-second billing | 4K | Preview | $1.00 | $1.20 | $3.00 |
| Veo 3.1 Lite | Per-second billing | 1080p | No 4K, no clip extension; preview | $0.50 | $0.80 | n/a |
| Gemini Omni Flash | 10 seconds (3s minimum) | 720p at 24 FPS | Preview and no EU editing of uploaded video | $1.00 | n/a | n/a |
A Unified Architecture for Media Creation and Physical Interaction
FLUX 3 is built upon BFL’s proprietary "Self-Flow" technique, a method developed to unify multimodal understanding and generation within a single architectural framework, first announced in March 2026. The lab has significantly increased its compute power and data resources to train FLUX 3 holistically across video, images, and audio. Testing has indicated that video generation and action prediction can indeed be achieved using the same underlying architecture without compromising performance on either task.
Robin Rombach, co-founder and CEO of BFL, articulated this vision: "We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture. True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express."
He further emphasized the limitations of single-modality models, stating, "You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds."
BFL is targeting FLUX 3 at a broad spectrum of industries, including creative tooling, media, design, e-commerce, and physical AI. Its capabilities extend to video generation with synchronized audio, precise image editing, maintaining product and material consistency across motion sequences, multilingual generation, and robotic action prediction. Early testing partners include prominent companies such as Canva, Burda, Magnific (formerly Freepik), Krea, and Picsart.
For creative software companies, the appeal of FLUX 3 lies in potential consolidation. A unified foundation could streamline workflows by supporting storyboarding, image editing, product rendering, video variation, and localization without the need for constant asset and instruction translation between disparate AI models. Similarly, robotics teams stand to benefit from enhanced data efficiency. Models that already possess an intrinsic understanding of motion, object behavior, and physical transformations may require less task-specific robot training compared to systems that start from scratch.

FLUX 3 Video: Advancing Generative Video Capabilities
The video generation aspect of FLUX 3 is the most concretely detailed component of the launch, resolving earlier speculation. FLUX 3 is capable of generating video clips up to 20 seconds in length, complete with synchronized audio, in a single generation pass. This feature significantly expands creative possibilities, offering longer, more cohesive video sequences. For comparison, HappyHorse 1.0 supports up to 15 seconds of 1080p video with synchronized audio. While BFL has not yet specified the resolution for its 20-second clips, its published evaluations were conducted at 720p. Nevertheless, achieving a 20-second clip from a single prompt places FLUX 3 among the longest-duration generative video models, rivaling OpenAI’s now-discontinued Sora model in this regard.
BFL outlines the following key capabilities for FLUX 3 Video:
- Extended Generation Duration: Up to 20 seconds of continuous video from a single prompt.
- Native Audio Synchronization: All video outputs include synchronized audio.
- Character Consistency: Mechanisms to maintain character identity across generated sequences, crucial for narrative coherence.
- Multi-Shot Sequencing: The ability to generate sequences that can span several minutes by intelligently chaining together multiple generated clips, with visual references ensuring consistent character portrayal across scenes. This addresses a critical limitation that has historically hindered the widespread adoption of generative video in commercial pipelines: maintaining continuity across shots.
- Human Facial Expression Nuance: Particular strength in rendering realistic and varied human facial expressions.
- Audio-Visual Correlation: Associating sounds with their corresponding physical events in the video.
- Multilingual Output: Support for generating content in multiple languages.
The capability for multi-shot sequencing is particularly noteworthy for enterprise video teams. If FLUX 3 can reliably maintain continuity across longer sequences under production conditions, it could overcome a major hurdle in integrating generative video into commercial workflows. This focus on character consistency is a highly contested area in generative video development. HappyHorse 1.1, for instance, introduced "Reference-to-Video" (R2V), a feature that accepts multiple character reference images to stabilize identity across generated footage. Alibaba also claims zero-drift lip-sync and aims to eliminate artifacts common in AI-generated video, such as facial oiliness and over-sharpening. The battle for commercial viability in this sector hinges on achieving robust character consistency.
Beyond video, BFL reports significant advancements in FLUX 3 Image generation during mid-training evaluations. The model demonstrates improved handling of complex prompts and enhanced text generation, including high-accuracy text rendering in multiple languages. However, no specific image benchmarks or win rates have been published to date.
FLUX-mimic: Bridging Video Understanding and Robotic Action
BFL is actively demonstrating its unified-architecture thesis through FLUX-mimic, a novel video-action model developed in collaboration with Mimic Robotics, a Swiss firm and one of the initial recipients of early access. The technical documentation from BFL outlines two primary pathways for action prediction. The first involves integrating native action prediction directly into the FLUX 3 architecture, building upon the foundational Self-Flow work. The second pathway, exemplified by FLUX-mimic, leverages the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be fine-tuned with minimal task-specific data.
FLUX-mimic specifically combines the FLUX 3 multimodal backbone with Mimic Robotics’ expertise in robot learning and production deployment for dexterous manipulation. The system is engineered for general-purpose robotic manipulation, enabling robots to interpret visual scenes, predict the outcomes of their actions, and adapt to new tasks with significantly reduced task-specific data requirements. BFL and Mimic Robotics claim that depending on the complexity of the task, FLUX-mimic can be fine-tuned for a specific manipulation task using as little as 30 minutes of robot data, a stark contrast to traditional approaches that often necessitate 30 or more hours of training data.
Elvis Nava, CTO of Mimic Robotics, highlighted the significance of this advancement: "The hardest part of robotics is data. Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning."
BFL’s assertion that models trained solely on images cannot grasp a world that "moves, sounds, changes, and responds" aligns with similar claims made by Google for its Gemini Omni model. Google emphasizes Gemini’s "world knowledge," which integrates an understanding of physics with historical, scientific, and cultural context, enabling more realistic movements governed by real-world logic, including forces like gravity and kinetic energy. While both BFL and Google champion their models’ "physical understanding," there is currently no standardized benchmark to objectively measure this capability. Human preference ratings offer an indirect assessment, but a direct quantification of physical realism in generated content remains an open challenge.
Open Weights: A Catalyst for FLUX’s Industry Integration
Since its official launch in the summer of 2024, Black Forest Labs has carved out a significant niche in the AI industry, largely attributed to its commitment to open-sourcing high-quality AI image models. This approach has garnered widespread adoption among developers, creative professionals, and enterprises alike. The company’s founders, including Robin Rombach, Andreas Blattmann, and Patrick Esser, are recognized for their foundational contributions to technologies like VQGAN, latent diffusion, and notably, Stable Diffusion. Stable Diffusion, in particular, democratized AI generation capabilities for the masses and continues to power numerous AI image generators and commercial applications.
This open-source ethos has translated into substantial commercial distribution. FLUX models are now integrated into generative features within leading platforms such as Adobe Photoshop, Picsart, and Nous Research’s Hermes Agent. Renowned film director Martin Scorsese is also cited as a professional user of FLUX models. Wired magazine has recognized Black Forest Labs as a formidable competitor to Silicon Valley’s largest AI labs, despite its relatively smaller size, with FLUX models consistently ranking at the top of image generation benchmarks and achieving high download volumes on the Hugging Face AI code-sharing community. The company now operates a team of 100 professionals across its Freiburg and San Francisco offices.
BFL continued its pattern of releasing open-weight models with FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev, and associated control models, made available shortly after the company’s inception. These releases provided researchers and creative tool developers with direct access to downloadable checkpoints, enabling local inference and seamless integration with popular frameworks like Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released under an open-weight license for research and non-commercial use, with generated outputs permitted for commercial purposes under its specific terms.
This commitment persisted with the launch of FLUX.2 Dev in late 2025. This 32-billion-parameter open-weight model combined advanced generation and multi-reference editing capabilities. BFL positioned it as the most powerful open-weight image generation and editing model available at the time of its release, providing weights, reference inference code, and optimized implementations for consumer Nvidia GPUs.
FLUX 3 Dev is poised to elevate this commitment further. While previous Dev releases were exclusively image models, FLUX 3 Dev is described as a multimodal backbone encompassing video, audio, image, and action prediction. This implies a single licensing framework will govern the local deployment of a model capable of both content production and physical machinery control. BFL has yet to disclose details regarding licensing terms, parameter counts, quantization options, or hardware requirements for FLUX 3 Dev.
BFL frames the provision of open weights not merely as a gesture to the developer community but as a strategic enterprise feature. The company argues that open weights enable secure, low-latency local deployments crucial for applications such as robotic control systems, and empower teams to adapt FLUX 3 to their unique datasets, products, and workflows.
The significant financial backing behind FLUX 3 underscores the company’s ambitious trajectory. Black Forest Labs is currently valued at $3.25 billion and has successfully raised over $450 million from a distinguished roster of investors, including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva, and Deutsche Telekom’s T.Capital. This robust financial support positions BFL to continue its aggressive innovation in the rapidly evolving AI landscape.

