Seedance 2.5 Now Live — First on Atlas Cloud
Hero background 1Hero background 2Hero background 3

FLUX 3 API: 20-Second Video With Native Audio

The FLUX 3 API brings Black Forest Labs' newest foundation model to your stack, a single Self-Flow architecture trained jointly on images, video, and audio. Generate clips up to 20 seconds carrying multilingual dialogue, sound effects, and ambience in one pass, or render text-to-image work with accurate typography. Atlas Cloud serves it on one OpenAI compatible key with pay-as-you-go per call pricing and Day-0 access.

Explore the Leading FLUX 3

Atlas Cloud provides you with the latest industry-leading creative models.

Every FLUX 3 API Endpoint, Modality by Modality

Five video endpoints sit behind one FLUX 3 API surface, and the table below shows what each one takes in, what it sends back, and what it costs per second.

ModalityDescription
FLUX 3 T2V API (Text to Video)Text prompts become clips of 5 to 20 seconds with audio generated in the same pass, since generate_audio is enabled by default. Output runs at 720p or 1080p across seven fixed aspect ratios from 21:9 to 9:16 plus an auto option, and a seed value makes any result reproducible. Pricing is $0.17 per second of generated video.
FLUX 3 I2V API (Image to Video)Supply a single PNG, JPEG, or WebP image and the model animates forward from that frame, taking motion and pacing from the prompt. Duration is selectable anywhere from 5 to 20 seconds, which suits product shots, character intros, and social cutdowns built on artwork you already own. Billing runs at $0.17 per second of generated video.
FLUX 3 Keyframes API (Keyframes to Video)When a shot has to hit specific beats, up to 10 keyframe images can be pinned to exact frame positions along a 24 fps timeline. Each frame_index has to stay within duration multiplied by 24, which gives storyboard and previz teams direct control over what appears and when. The rate matches the other generation endpoints at $0.17 per second of generated video.
FLUX 3 FLF API (First and Last Frame to Video)Need a clip to land on an exact closing image? Passing start_image_url and end_image_url locks both ends of the shot while the model fills in the motion between them. Logo reveals, transformation sequences, and scene handoffs that must match surrounding footage all fit here, at $0.17 per second of generated video.
FLUX 3 Extend API (Video Extension)Existing footage carries the story forward on this endpoint: an MP4 under 50 MB and shorter than 15 seconds is continued from its final frames according to the prompt. Audio keeps generating alongside the picture, and the same 720p or 1080p output and 5 to 20 second range still apply. Extension is priced at $0.41 per second of generated video.

Video, Audio, and Direction in a Single FLUX 3 API Call

Built by Black Forest Labs on one multimodal backbone, the FLUX 3 API returns up to twenty seconds of HD or Full HD video with dialogue, sound effects, and ambience generated in the same pass, and it takes direction on shots, keyframes, languages, and on-screen text along the way.

Twenty Seconds of Picture and Sound in One Pass

A single call returns 5 to 20 seconds of video with audio generated inside the same pass, never dubbed on afterward. Dialogue, sound effects, and ambient beds land on the exact frame where the event happens, and output arrives at 720p HD or 1080p Full HD. That timing is what makes a wok flare or a slammed door read as filmed rather than assembled.

Scene and Camera Changes Inside One Generation

Ask for a cut and the model delivers it. Scenes and camera angles change within one generation while the same character, wardrobe, and lighting logic carry across every shot. Longer pieces come from agentic clip chaining, which links separate generations into sequences that run for minutes. Storyboards that used to need four renders and an editor now come back as one continuous piece.

Five Ways Into the FLUX 3 API

Text alone, a single still, a first and last frame pair, keyframes placed across the timeline, or an existing clip to continue: all five entry points are supported. Keyframes pin what happens and when, while continuation picks up from footage you already hold. Studios sitting on brand stills or half finished cuts can start from those assets instead of describing every shot from scratch.

Dialogue That Lands in Thirteen Languages

Native dialogue spans English dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi, and more, with lip movement matched to whichever language you request. Typography renders as part of the scene as well, so signage and titles sit inside the frame instead of being composited later. One prompt can therefore ship a localized spot with its lettering already in place.

The FLUX 3 API Behind One Atlas Cloud Key

Atlas Cloud serves FLUX 3 through the same unified endpoint that fronts the rest of its catalog, so one key reaches video, image, and language models without a second integration. Billing stays pay as you go with no subscription and no seat fees, and you pay only for the generations you actually run. Swapping models means changing a string. Start building today.

One Prompt, Three Models: FLUX 3 API Side by Side

Each row below runs a single identical prompt through the FLUX 3 API and two other video models on Atlas Cloud, so motion continuity, shot changes, and native audio can be judged on the same brief rather than on cherry picked demos.

Prompt

Cinematic live action, rain slick Tokyo backstreet at night. A Shiba Inu in a tiny yellow raincoat rides a skateboard down the alley, weaving between puddles and steam vents. Open on a low angle tracking shot skimming the wet pavement just behind the board, neon reflections streaking past. The dog clips a stack of paper lanterns and they burst into the air, and a whip pan follows one spinning lantern as it tumbles toward a ramen stall. Cut to a drone shot pulling up and back while the Shiba slams a paw down, spins the board to a hard stop, and barks once at the laughing ramen chef, who flicks a slice of chashu that the dog catches midair. Warm stall light against cold neon, water spray lit from behind. Audio: hissing rain, urethane wheels rumbling over stone, a sizzling wok, one sharp bark, a train passing somewhere above. 16:9 aspect ratio.

Generated with BLACKFORESTLABS FLUX 3 on Atlas Cloud

Generated with Veo3.1 on Atlas Cloud

Generated with Kling v3.0 on Atlas Cloud

Prompt

Hand painted anime style, bright noon above an endless sea of clouds. A teenage sky fisher braces on the deck of a wooden airship and casts a long line down into the cloud sea. Begin with a first person POV over the railing as the line whips out and vanishes into white. The rod snaps taut, and the camera cuts to a fast orbit around the deck while she is dragged across the planks, rope smoking through her gloves, one boot hooking a coil of rigging. She plants a foot on the railing and hauls back just as a koi the size of a whale breaches through the clouds behind her, scales throwing rainbow light across the sails. Finish on a low angle hero shot as the crew erupts cheering and cloud spray drifts through the frame. Painterly cel shading, warm rim light, visible brush texture. Audio: rushing wind, creaking rope, a rising orchestral swell, a deep watery boom on the breach. 16:9 aspect ratio.

Generated with BLACKFORESTLABS FLUX 3 on Atlas Cloud

Generated with Seedance 2.0 on Atlas Cloud

Generated with Kling v3.0 on Atlas Cloud

Where Teams Put the FLUX 3 API to Work

Whether the job is a social clip that runs twenty seconds with native audio, a keyframed storyboard, a localized ad cut, or a quick draft pass before the final render, the FLUX 3 API covers it on Atlas Cloud with pay-as-you-go pricing and Day-0 access.

Social Clips That Arrive With Their Own Sound

Generate up to twenty seconds of video from a single text prompt, with speech, effects, and ambience produced alongside the frames. Social teams get a finished clip without a separate audio pass.

Multi-Shot Storyboards Through the FLUX 3 API

Ordered keyframes let the FLUX 3 API move a scene through defined moments, shifting camera angles and settings inside one generation. Storyboard artists can preview a full sequence before any shoot is booked.

Typography and Text Inside the Frame

Typography renders accurately inside generated scenes, and multilingual text handling improves on what earlier FLUX generations delivered. Packaging mockups, title cards, and localized banners come out of a single call ready for review.

Localized Ad Cuts With the FLUX 3 API

Need the same spot in six markets? The FLUX 3 API generates synchronized multilingual dialogue with the video, so each cut carries its own localized voice track without a dubbing stage.

Second Life for Existing Footage

Existing footage can be carried into new contexts, extended with matching audio, or reshaped while its central elements stay intact. Agencies revive old campaign assets instead of commissioning another shoot.

Cheap Draft Passes on the FLUX 3 API

Draft mode returns a fast HD preview at lower cost, and approved shots render again at HD or FHD through the FLUX 3 API. Iteration stays affordable when a concept needs twenty attempts.

FLUX 3 API Side by Side With Other Video Models on Atlas Cloud

Weigh clip length, output resolution, native audio, and per second cost to see where the FLUX 3 API pulls ahead and where another Atlas Cloud model may suit your pipeline better.

ModelMax Clip LengthMax Output ResolutionNative AudioPrice per Second
FLUX 3 Video (Text-to-Video)20 sFHD, up to 2 MP per frame√ dialogue, SFX, ambience$0.17 HD / $0.29 FHD
FLUX 3 Video (Video-to-Video)20 sFHD, up to 2 MP per frame√ dialogue, SFX, ambience$0.41 HD / $0.53 FHD
Seedance 2.0 Text-to-Video15 s4K (3840 x 2160)√ on by default$0.112
Kling V3.0 Turbo Text-to-Video15 s1080p√ optional, off by default$0.112
Wan-2.7 Text-to-video15 s1080p native, 1440p-SR√ SFX, music, ambience$0.10
Veo3.1 Text-to-video8 s1080p√ synchronized audio$0.20

How to Use FLUX 3 on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use FLUX 3 on Atlas Cloud

Combining the advanced FLUX 3 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run FLUX 3, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

FLUX 3 API Questions, Answered for Developers

FLUX 3 is Black Forest Labs' multimodal foundation model, trained jointly across image, video, audio, and action prediction instead of being stitched together from separate systems. The FLUX 3 API exposes that model through Atlas Cloud, so a single request returns video with synchronized audio. One OpenAI-compatible key covers it alongside every other model in the catalog, billed per use.

Video is the capability that is generally available, covering text to video, image to video, keyframe-guided shots, and continuation of an existing clip, each with optional native audio. Black Forest Labs is releasing the family in phases, so FLUX 3 Image and FLUX 3 Action follow the video endpoints rather than shipping alongside them.

Create an account, generate an API key, and point your existing client at the FLUX 3 API endpoint. Requests follow the same OpenAI-compatible pattern used across the Atlas Cloud catalog, so there is no separate SDK to learn and no vendor lock-in to unwind later. Billing is pay-as-you-go per generation with no subscription to sign first. Start building today.

Yes, and it happens in the same generation pass rather than through a second model, which keeps dialogue, sound effects, and ambience aligned to the picture. FLUX 3 handles speech with lip sync across more than a dozen languages, including English, Chinese, Spanish, Japanese, and Hindi. Audio can be turned off per request when silent footage is all you need.

Generations run from 5 to 20 seconds, output at HD 720p with Full HD 1080p available when detail matters more than cost. Widescreen, square, and vertical framings are all supported, so the same prompt can be aimed at a landing page hero or a mobile feed. Longer stories are best built from several controlled clips rather than one oversized request.

Supply a starting image and the model animates it, or pin keyframes when a shot has to hit specific moments in a set order. Existing footage can also be continued, with motion and framing carried over from the tail of the input clip. Chain continuations sparingly, since visual consistency drifts as each extension builds on the previous one.

Draft mode returns a lower-cost HD preview so composition, pacing, and audio can be judged before a final render is paid for. Iterate on prompts there, then rerun the winning prompt at standard HD or Full HD without changing the request structure. Teams shipping many variants get the most value from this loop.

Video is billed per second of output, so a short clip costs a fraction of a long one and you pay only for what you actually generate. Resolution sets the rate, with Full HD above standard HD and Draft mode below both. Atlas Cloud adds no subscription, seat fee, or minimum spend on top. Start today.

Not yet. FLUX 3 Video is the released piece of the family, while FLUX 3 Image is in staged rollout and an open weights FLUX 3 Dev variant has been announced for later. Atlas Cloud adds each new FLUX 3 endpoint as it ships, so the same key keeps working when image generation arrives.

A joint model removes the stitching step: there is no second pass to time voice against mouth movement or to layer ambience under a finished cut. That matters most for dialogue scenes, where drift between tracks is immediately visible to viewers. Pipelines shrink from several services to one call, which is fewer failure points to monitor.

Explore More Families

Seedance 2.5

Seedance 2.5 API is now available on Atlas Cloud! It gives developers ByteDance's newest video model. It generates up to 30 seconds of native video in a single pass from text, a single image, or as many as 50 multimodal references, with synchronized audio and in-frame multilingual text. On Atlas Cloud you reach it through one key, with subject consistency and improved physics keeping long shots coherent.

View Family

MiniMax H3

The MiniMax H3 API opens MiniMax's general purpose multimodal video model, which reads text, images, video and audio as one context instead of one task at a time. Clips run 5 to 15 seconds at 24 FPS across aspect ratios from 21:9 to 9:16, and one prompt can swap characters, replace backgrounds, rewrite dialogue or clone a voice from a reference clip. Atlas Cloud serves it all through one OpenAI-compatible endpoint. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

The Seedance 2.0 API gives you production access to ByteDance's multimodal video model — quad-modal inputs (text, image, video, audio) and an industry-leading "Universal Reference" system that locks composition, camera movement, and character actions across shots. Integrate director-level control with one API call, a flat $0.09/s, instant key, and no waitlist — backed by enterprise-grade uptime and compliance. Seedance 2.0 Native 4K is now live!

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The Gemini Omni API brings Google DeepMind's multimodal video generation and editing model, introduced at Google I/O 2026, to your stack. Gemini Omni fuses Gemini's reasoning engine with generative media, accepting any mix of text, images, video, and audio to produce consistent, knowledge-grounded output. Refine results through natural conversation, swapping objects, rewriting scenes, and shifting styles while physics, characters, and continuity stay intact. Atlas Cloud serves the full Gemini Omni Flash lineup, text-to-video, image-to-video with up to 7 reference images, and reference-to-video, through one unified API with transparent per-second pricing from $0.112 and no subscription. Start building today.

View Family

Grok Imagine

The Grok Imagine API covers xAI's image, video, and speech models, from Image 2.0 to Video 1.5 and xAI TTS v1. Render 1K or 2K stills across 14 aspect ratios, push a scene to 15 seconds of 1080p motion, steer shots with up to 7 reference images, or narrate them in 20 languages. Atlas Cloud runs every mode on one endpoint, priced pay-as-you-go from $0.02 per image and $0.05 per second. Start building today.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

Alibaba

Atlas Cloud brings together Alibaba's full model lineup under one API: Qwen for language and image tasks, Wan for video generation up to 1080p. Access every model pay-as-you-go with no subscriptions. The Alibaba API is available via a single base URL using your existing OpenAI-compatible client.

View Family

OpenAI

Atlas Cloud gives you access to the full OpenAI API lineup, from GPT Image 2 for image generation to Sora 2 for video. Every model is available pay-as-you-go with no monthly commitment. Plug in with a single base URL swap using the OpenAI-compatible API.

View Family

One API for All Media AI.

Explore all models