


The FLUX 3 API brings Black Forest Labs' newest foundation model to your stack, a single Self-Flow architecture trained jointly on images, video, and audio. Generate clips up to 20 seconds carrying multilingual dialogue, sound effects, and ambience in one pass, or render text-to-image work with accurate typography. Atlas Cloud serves it on one OpenAI compatible key with pay-as-you-go per call pricing and Day-0 access.
Atlas Cloud provides you with the latest industry-leading creative models.
Five video endpoints sit behind one FLUX 3 API surface, and the table below shows what each one takes in, what it sends back, and what it costs per second.
| Modality | Description |
|---|---|
| FLUX 3 T2V API (Text to Video) | Text prompts become clips of 5 to 20 seconds with audio generated in the same pass, since generate_audio is enabled by default. Output runs at 720p or 1080p across seven fixed aspect ratios from 21:9 to 9:16 plus an auto option, and a seed value makes any result reproducible. Pricing is $0.17 per second of generated video. |
| FLUX 3 I2V API (Image to Video) | Supply a single PNG, JPEG, or WebP image and the model animates forward from that frame, taking motion and pacing from the prompt. Duration is selectable anywhere from 5 to 20 seconds, which suits product shots, character intros, and social cutdowns built on artwork you already own. Billing runs at $0.17 per second of generated video. |
| FLUX 3 Keyframes API (Keyframes to Video) | When a shot has to hit specific beats, up to 10 keyframe images can be pinned to exact frame positions along a 24 fps timeline. Each frame_index has to stay within duration multiplied by 24, which gives storyboard and previz teams direct control over what appears and when. The rate matches the other generation endpoints at $0.17 per second of generated video. |
| FLUX 3 FLF API (First and Last Frame to Video) | Need a clip to land on an exact closing image? Passing start_image_url and end_image_url locks both ends of the shot while the model fills in the motion between them. Logo reveals, transformation sequences, and scene handoffs that must match surrounding footage all fit here, at $0.17 per second of generated video. |
| FLUX 3 Extend API (Video Extension) | Existing footage carries the story forward on this endpoint: an MP4 under 50 MB and shorter than 15 seconds is continued from its final frames according to the prompt. Audio keeps generating alongside the picture, and the same 720p or 1080p output and 5 to 20 second range still apply. Extension is priced at $0.41 per second of generated video. |
Built by Black Forest Labs on one multimodal backbone, the FLUX 3 API returns up to twenty seconds of HD or Full HD video with dialogue, sound effects, and ambience generated in the same pass, and it takes direction on shots, keyframes, languages, and on-screen text along the way.
A single call returns 5 to 20 seconds of video with audio generated inside the same pass, never dubbed on afterward. Dialogue, sound effects, and ambient beds land on the exact frame where the event happens, and output arrives at 720p HD or 1080p Full HD. That timing is what makes a wok flare or a slammed door read as filmed rather than assembled.
Ask for a cut and the model delivers it. Scenes and camera angles change within one generation while the same character, wardrobe, and lighting logic carry across every shot. Longer pieces come from agentic clip chaining, which links separate generations into sequences that run for minutes. Storyboards that used to need four renders and an editor now come back as one continuous piece.
Text alone, a single still, a first and last frame pair, keyframes placed across the timeline, or an existing clip to continue: all five entry points are supported. Keyframes pin what happens and when, while continuation picks up from footage you already hold. Studios sitting on brand stills or half finished cuts can start from those assets instead of describing every shot from scratch.
Native dialogue spans English dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi, and more, with lip movement matched to whichever language you request. Typography renders as part of the scene as well, so signage and titles sit inside the frame instead of being composited later. One prompt can therefore ship a localized spot with its lettering already in place.
Atlas Cloud serves FLUX 3 through the same unified endpoint that fronts the rest of its catalog, so one key reaches video, image, and language models without a second integration. Billing stays pay as you go with no subscription and no seat fees, and you pay only for the generations you actually run. Swapping models means changing a string. Start building today.
Each row below runs a single identical prompt through the FLUX 3 API and two other video models on Atlas Cloud, so motion continuity, shot changes, and native audio can be judged on the same brief rather than on cherry picked demos.
Cinematic live action, rain slick Tokyo backstreet at night. A Shiba Inu in a tiny yellow raincoat rides a skateboard down the alley, weaving between puddles and steam vents. Open on a low angle tracking shot skimming the wet pavement just behind the board, neon reflections streaking past. The dog clips a stack of paper lanterns and they burst into the air, and a whip pan follows one spinning lantern as it tumbles toward a ramen stall. Cut to a drone shot pulling up and back while the Shiba slams a paw down, spins the board to a hard stop, and barks once at the laughing ramen chef, who flicks a slice of chashu that the dog catches midair. Warm stall light against cold neon, water spray lit from behind. Audio: hissing rain, urethane wheels rumbling over stone, a sizzling wok, one sharp bark, a train passing somewhere above. 16:9 aspect ratio.
Generated with BLACKFORESTLABS FLUX 3 on Atlas Cloud
Generated with Veo3.1 on Atlas Cloud
Generated with Kling v3.0 on Atlas Cloud
Hand painted anime style, bright noon above an endless sea of clouds. A teenage sky fisher braces on the deck of a wooden airship and casts a long line down into the cloud sea. Begin with a first person POV over the railing as the line whips out and vanishes into white. The rod snaps taut, and the camera cuts to a fast orbit around the deck while she is dragged across the planks, rope smoking through her gloves, one boot hooking a coil of rigging. She plants a foot on the railing and hauls back just as a koi the size of a whale breaches through the clouds behind her, scales throwing rainbow light across the sails. Finish on a low angle hero shot as the crew erupts cheering and cloud spray drifts through the frame. Painterly cel shading, warm rim light, visible brush texture. Audio: rushing wind, creaking rope, a rising orchestral swell, a deep watery boom on the breach. 16:9 aspect ratio.
Generated with BLACKFORESTLABS FLUX 3 on Atlas Cloud
Generated with Seedance 2.0 on Atlas Cloud
Generated with Kling v3.0 on Atlas Cloud
Whether the job is a social clip that runs twenty seconds with native audio, a keyframed storyboard, a localized ad cut, or a quick draft pass before the final render, the FLUX 3 API covers it on Atlas Cloud with pay-as-you-go pricing and Day-0 access.
Generate up to twenty seconds of video from a single text prompt, with speech, effects, and ambience produced alongside the frames. Social teams get a finished clip without a separate audio pass.
Ordered keyframes let the FLUX 3 API move a scene through defined moments, shifting camera angles and settings inside one generation. Storyboard artists can preview a full sequence before any shoot is booked.
Typography renders accurately inside generated scenes, and multilingual text handling improves on what earlier FLUX generations delivered. Packaging mockups, title cards, and localized banners come out of a single call ready for review.
Need the same spot in six markets? The FLUX 3 API generates synchronized multilingual dialogue with the video, so each cut carries its own localized voice track without a dubbing stage.
Existing footage can be carried into new contexts, extended with matching audio, or reshaped while its central elements stay intact. Agencies revive old campaign assets instead of commissioning another shoot.
Draft mode returns a fast HD preview at lower cost, and approved shots render again at HD or FHD through the FLUX 3 API. Iteration stays affordable when a concept needs twenty attempts.
Weigh clip length, output resolution, native audio, and per second cost to see where the FLUX 3 API pulls ahead and where another Atlas Cloud model may suit your pipeline better.
| Model | Max Clip Length | Max Output Resolution | Native Audio | Price per Second |
|---|---|---|---|---|
| FLUX 3 Video (Text-to-Video) | 20 s | FHD, up to 2 MP per frame | √ dialogue, SFX, ambience | $0.17 HD / $0.29 FHD |
| FLUX 3 Video (Video-to-Video) | 20 s | FHD, up to 2 MP per frame | √ dialogue, SFX, ambience | $0.41 HD / $0.53 FHD |
| Seedance 2.0 Text-to-Video | 15 s | 4K (3840 x 2160) | √ on by default | $0.112 |
| Kling V3.0 Turbo Text-to-Video | 15 s | 1080p | √ optional, off by default | $0.112 |
| Wan-2.7 Text-to-video | 15 s | 1080p native, 1440p-SR | √ SFX, music, ambience | $0.10 |
| Veo3.1 Text-to-video | 8 s | 1080p | √ synchronized audio | $0.20 |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced FLUX 3 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run FLUX 3, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
FLUX 3 is Black Forest Labs' multimodal foundation model, trained jointly across image, video, audio, and action prediction instead of being stitched together from separate systems. The FLUX 3 API exposes that model through Atlas Cloud, so a single request returns video with synchronized audio. One OpenAI-compatible key covers it alongside every other model in the catalog, billed per use.
Video is the capability that is generally available, covering text to video, image to video, keyframe-guided shots, and continuation of an existing clip, each with optional native audio. Black Forest Labs is releasing the family in phases, so FLUX 3 Image and FLUX 3 Action follow the video endpoints rather than shipping alongside them.
Create an account, generate an API key, and point your existing client at the FLUX 3 API endpoint. Requests follow the same OpenAI-compatible pattern used across the Atlas Cloud catalog, so there is no separate SDK to learn and no vendor lock-in to unwind later. Billing is pay-as-you-go per generation with no subscription to sign first. Start building today.
Yes, and it happens in the same generation pass rather than through a second model, which keeps dialogue, sound effects, and ambience aligned to the picture. FLUX 3 handles speech with lip sync across more than a dozen languages, including English, Chinese, Spanish, Japanese, and Hindi. Audio can be turned off per request when silent footage is all you need.
Generations run from 5 to 20 seconds, output at HD 720p with Full HD 1080p available when detail matters more than cost. Widescreen, square, and vertical framings are all supported, so the same prompt can be aimed at a landing page hero or a mobile feed. Longer stories are best built from several controlled clips rather than one oversized request.
Supply a starting image and the model animates it, or pin keyframes when a shot has to hit specific moments in a set order. Existing footage can also be continued, with motion and framing carried over from the tail of the input clip. Chain continuations sparingly, since visual consistency drifts as each extension builds on the previous one.
Draft mode returns a lower-cost HD preview so composition, pacing, and audio can be judged before a final render is paid for. Iterate on prompts there, then rerun the winning prompt at standard HD or Full HD without changing the request structure. Teams shipping many variants get the most value from this loop.
Video is billed per second of output, so a short clip costs a fraction of a long one and you pay only for what you actually generate. Resolution sets the rate, with Full HD above standard HD and Draft mode below both. Atlas Cloud adds no subscription, seat fee, or minimum spend on top. Start today.
Not yet. FLUX 3 Video is the released piece of the family, while FLUX 3 Image is in staged rollout and an open weights FLUX 3 Dev variant has been announced for later. Atlas Cloud adds each new FLUX 3 endpoint as it ships, so the same key keeps working when image generation arrives.
A joint model removes the stitching step: there is no second pass to time voice against mouth movement or to layer ambience under a finished cut. That matters most for dialogue scenes, where drift between tracks is immediately visible to viewers. Pipelines shrink from several services to one call, which is fewer failure points to monitor.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.