Model Watch
FLUX 3 Is Coming to Artifio.ai: Black Forest Labs' Multimodal Model, Explained
One model for video, images, and audio — 20-second clips with native sound, sharper stills, and open weights on the roadmap. Here's what was announced, and when you'll be able to use it on Artifio.

What's inside
- 01. What is FLUX 3?
- 02. One model, three senses
- 03. FLUX 3 Video: 20-second clips with native audio
- 04. How FLUX 3 stacks up so far
- 05. FLUX 3 Image: harder prompts, cleaner text
- 06. Robot arms and open weights
- 07. Self-Flow, in plain English
- 08. The FLUX 3 release timeline
- 09. FLUX 3 on Artifio.ai
- 10. What you can create today
- 11. FAQ
The short answer
What is FLUX 3?
On July 23, 2026, Black Forest Labs — the team behind the FLUX image models — announced FLUX 3, a multimodal foundation model that learns from images, video, and audio inside a single architecture. Not three models stitched together: one model that understands how things look, how they move, and how they sound.
That single backbone powers several products. FLUX 3 Video (video plus native audio, in early access now), FLUX 3 Image (stills and editing, early access opening in the coming weeks), FLUX 3 Action (motion prediction for robotics, through selected partners), and FLUX 3 Dev — a planned open-weight release of the multimodal backbone itself.
If you've used FLUX.1 Schnell or FLUX.2 on Artifio, this is the next generation of that same family — except this time the headline isn't a better image model. It's a model that treats images, video, and sound as one problem.
Why it matters
One model, three senses
The thinking behind FLUX 3 is simple to state: no single medium tells the whole story. A photo captures how a scene is arranged at one instant. Video adds time — how things move and how physics plays out. Audio carries the causal fingerprint of events: the thud has to match the impact, the voice has to match the lips.
Train a model on each of those separately and you get three narrow specialists. Train one model on all of them at once and each modality keeps the others honest — the sound must fit the motion, the motion must fit the scene, the next frame must follow from the last. Black Forest Labs describes this as building a model of the world rather than a model of pixels.
For creators, the practical payoff is coherence. AI video where the audio was bolted on afterwards always feels slightly wrong. A model that generated the picture and the sound together, from the same understanding of the scene, doesn't have that seam.

Headline capability
FLUX 3 Video: 20-second clips with native audio
The flagship capability is video generation up to 20 seconds long in a single pass — with sound generated natively alongside the picture, not layered on after. Dialogue, ambience, and effects come out of the same generation as the frames they belong to.
It's also not just text-to-video. FLUX 3 Video covers the full set of workflows working creators actually use:
- Text-to-video — describe the shot, get the clip with audio
- Image-to-video — animate a starting frame, or use images as visual references for style and character
- Video-to-video — carry a character or element from a source clip into a completely new scene
- Keyframe-to-video — define the start and end moments, let the model handle the transition
- Video + audio continuation — extend an existing clip, picture and sound together
- Multilingual dialogue — spoken lines in multiple languages, matched to the scene
- Multi-shot sequences — chain clips agentically into longer edits, with reference images keeping characters consistent across scenes
- Broad style range — camcorder realism, animation, cinematic looks, plus strong typography and animated design work
Early numbers
How FLUX 3 stacks up so far
Black Forest Labs published preliminary head-to-head preference results from 10-second, 720p text-to-video generations with audio. In those comparisons, human raters preferred FLUX 3 over Luma Ray 3.2 in 93% of match-ups and over Runway Gen-4.5 in 77%. Against the current heavyweights the margins were tighter but still ahead: up to 69% over Grok Imagine, 60% over Kling v3 Pro, and 52% over Seedance 2.0.
Two caveats worth keeping. These are the lab's own evaluations, run while the model is still in training — BFL itself says it expects further improvement before general release. And preference rates compress a lot of nuance: a model can win overall while trailing on specific shot types.
That said, the areas BFL calls out as already strong line up with what's hardest in AI video today: believable facial expressions, sound that matches physical events, and multilingual speech. Those are exactly the things that break immersion when they're wrong.
Stills too
FLUX 3 Image: harder prompts, cleaner text
The FLUX family made its name on image generation, and FLUX 3 doesn't leave that behind. The image side of the model handles synthesis and editing across a wide range of styles, aspect ratios, and resolutions — with two improvements BFL highlights over FLUX.2: notably better handling of complex, multi-part prompts, and high-accuracy text rendering in multiple languages.
Accurate in-image text has quietly become one of the biggest quality gaps between model generations. If you make thumbnails, posters, product shots, or anything with words in it, this is the upgrade you'll feel first.
FLUX 3 Image enters its own early access phase in the weeks after the video rollout.
Beyond content
Robot arms and open weights
The more surprising part of the announcement: the same video backbone is being used for action prediction — teaching robots to manipulate objects. BFL partnered with mimic robotics to build FLUX-mimic, a video-action model already being tested on real production tasks at Audi. The logic is that a model which has learned how the physical world moves from video is a head start for a model that has to act in it.
You won't use FLUX 3 Action to make content, but it matters as a signal: this is a genuine world model, not a video toy — and the physical-AI work funds and hardens the same backbone creators generate with.
The other roadmap item worth watching is FLUX 3 Dev: a planned open-weight release of the multimodal backbone covering video, audio, image, and action. Open weights are what turned FLUX.1 into an ecosystem — fine-tunes, LoRAs, and community tooling. An open multimodal backbone at this level would be a first.
Under the hood
Self-Flow, in plain English
FLUX 3 is built on an approach BFL calls Self-Flow — its method for aligning generation and understanding across modalities inside one architecture, scaled up with substantially more compute and data than earlier FLUX releases.
The published comparison against standard flow matching (the technique behind most current diffusion-style models) shows Self-Flow producing lower generation error across every modality tested, and higher success rates when the model is fine-tuned for manipulation tasks. In other words: the same trick that makes the videos better also makes the robot better — which is the whole thesis in one chart.
BFL says a fuller technical report on the approach is coming; we'll update this post when it lands.
When
The FLUX 3 release timeline
Black Forest Labs is rolling FLUX 3 out in stages, each gated by an early-access phase for feedback and safety testing. Here's the order of operations as announced:
- Now — FLUX 3 Video in early access (API and private weight access)
- Coming weeks — FLUX 3 Image early access for synthesis and editing
- Ongoing — FLUX 3 Action through selected research and commercial partners, starting with mimic robotics
- Later — FLUX 3 Dev: open-weight multimodal backbone for content creation and action prediction
Coming soon
FLUX 3 on Artifio.ai
Here's the part you actually came for: FLUX 3 is coming to Artifio. We already run the FLUX family — FLUX.1 Schnell, FLUX.1 Dev, and FLUX.2 are live in the catalogue today — and we're tracking the FLUX 3 early-access program with the goal of bringing FLUX 3 Video and FLUX 3 Image into the workspace as access opens up.
When it lands, it lands like every other model here: one wallet across image, video, and audio, priced per generation with no subscription, and a failed generation refunds itself automatically. You'll be able to run FLUX 3 next to Kling, Veo, Seedance, and Runway and pick the winner per shot — or chain it into an Artiflow pipeline where one model's output feeds the next.
We don't list a model until it has passed our end-to-end verification — submit, deliver, recover, refund — so we won't promise a date BFL hasn't given anyone. But the moment FLUX 3 is stable and available to platforms, you'll find it in the catalogue and announced on this blog.
Meanwhile
What you can create today
No need to sit on your hands during early access. FLUX.2 — the current generation of the same family — is live on Artifio right now for image work, alongside FLUX.1 Dev and the speed-focused FLUX.1 Schnell.
On the video side, the models FLUX 3 was benchmarked against are already in the catalogue: Kling, Veo, Seedance, Runway, and more, all under the same pay-per-generation wallet. If FLUX 3's early numbers hold, you'll be able to compare them side by side here the day it arrives — same prompt, same workspace, your call.
Be first in line
FLUX 3 lands here when access opens
Every new model we ship is announced on this blog and appears in the catalogue the day it passes verification. Until then, the rest of the FLUX family — and the video models FLUX 3 was benchmarked against — are one click away.
FLUX 3 FAQ
What is FLUX 3?
FLUX 3 is Black Forest Labs' multimodal foundation model, announced July 23, 2026. It learns from images, video, and audio in a single architecture and powers several products: FLUX 3 Video (video with native audio), FLUX 3 Image (image generation and editing), FLUX 3 Action (robotics), and a planned open-weight release called FLUX 3 Dev.
When is the FLUX 3 release date?
FLUX 3 Video entered early access on July 23, 2026. FLUX 3 Image opens its early-access phase in the following weeks, and the open-weight FLUX 3 Dev backbone comes later. Black Forest Labs has not announced a firm general-availability date — each capability ships after its early-access and safety-testing phase.
How long can FLUX 3 videos be?
Up to 20 seconds in a single generation, with audio generated natively alongside the picture. Longer, multi-shot sequences lasting several minutes are possible by chaining clips, with reference images keeping characters consistent across scenes.
Does FLUX 3 generate audio with its videos?
Yes — natively. Dialogue, ambience, and sound effects are generated in the same pass as the frames, by the same model, so the sound is derived from the same understanding of the scene as the picture. It also supports multilingual spoken dialogue.
How does FLUX 3 compare to Kling, Runway, and Luma?
In Black Forest Labs' preliminary evaluations of 10-second 720p text-to-video with audio, raters preferred FLUX 3 over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine in up to 69%, Kling v3 Pro in 60%, and Seedance 2.0 in 52%. These are the lab's own numbers from a model still in training, so treat them as directional.
Will FLUX 3 have open weights?
That's the plan. Black Forest Labs has announced FLUX 3 Dev, an open-weight release of the multimodal backbone covering video, audio, image, and action prediction. It arrives after the API products, following the same open-weight tradition as FLUX.1.
When is FLUX 3 coming to Artifio.ai?
As soon as platform access opens and the model passes our end-to-end verification. Artifio already runs FLUX.1 Schnell, FLUX.1 Dev, and FLUX.2, and we're tracking the FLUX 3 early-access program with the goal of adding FLUX 3 Video and FLUX 3 Image to the catalogue. New arrivals are announced on this blog.
What will FLUX 3 cost on Artifio?
Black Forest Labs hasn't published platform pricing yet, so we can't quote a number. What won't change is the model: on Artifio you pay per generation from one wallet that works across every image, video, and audio model — no subscription, and failed generations are refunded automatically.
Sources
Every date, capability, and evaluation figure in this article comes from Black Forest Labs' own FLUX 3 announcement — the primary source below. Secondary coverage is listed for additional context.
- Primary sourceFLUX 3 — Real World Models: Towards Multimodal Flow Models as the Backbone of Visual IntelligenceBlack Forest Labs
- CoverageFLUX 3 video model launch coverageMindStudio