рджреБрдирд┐рдпрд╛ рднрд░ рдореЗрдВ рд╕рдмрд╕реЗ рдХрдо рдХреАрдорддреЛрдВ рдкрд░ Seedance 2.0 Mini & Fast API тАФ рдЖрдзрд┐рдХрд╛рд░рд┐рдХ рдХреАрдордд рдкрд░ 68% рддрдХ рдХреА рдЫреВрдЯ
Atlas Cloud AI рдХреНрд░рд┐рдПрд╢рди рд╕реНрдЯреВрдбрд┐рдпреЛрдЕрдкрдиреЗ рднреАрддрд░ рдХреЗ рдирд┐рд░реНрджреЗрд╢рдХ рдХреЛ рдкрд╣рдЪрд╛рдиреЗрдВредрдмрдирд╛рдирд╛ рд╢реБрд░реВ рдХрд░реЗрдВ
Gemini Omni Flash Image-to-Video Developer
рдЗрдореЗрдЬ-рд╕реЗ-рд╡реАрдбрд┐рдпреЛ
DEV

Gemini Omni Flash Image-to-Video Developer API by Google

google/gemini-omni-flash/image-to-video-developer
Image-to-video-developer

Gemini Omni Flash is Google's multimodal video generation model. This image-to-video variant creates subject-consistent videos from up to 7 reference images combined with a text prompt, preserving visual identity across the full generated video.

Gemini Omni Flash Image-to-Video Developer рдХреЛ Google рджреНрд╡рд╛рд░рд╛ рд╡рд┐рдХрд╕рд┐рдд рдХрд┐рдпрд╛ рдЧрдпрд╛ рд╣реИред Atlas Cloud (Atlas Cloud AI LLC рджреНрд╡рд╛рд░рд╛ рд╕рдВрдЪрд╛рд▓рд┐рдд) рдЗрд╕ рддрдХ рдкрд╣реБрдБрдЪ рдкреНрд░рджрд╛рди рдХрд░рддрд╛ рд╣реИ, рдЗрд╕рдХрд╛ рд╕реНрд╡рд╛рдореА рдирд╣реАрдВ рд╣реИред рд╕рднреА рдЯреНрд░реЗрдбрдорд╛рд░реНрдХ рдЙрдирдХреЗ рд╕рдВрдмрдВрдзрд┐рдд рд╕реНрд╡рд╛рдорд┐рдпреЛрдВ рдХреА рд╕рдВрдкрддреНрддрд┐ рд╣реИрдВред

Gemini Omni Flash тАФ Image to Video (Developer)

Model ID: google/gemini-omni-flash/image-to-video-developer

Gemini Omni is Google's multimodal video generation model designed to create high-quality video content from diverse input types. This variant accepts a text prompt plus up to 7 reference images, enabling subject-consistent video generation where the visual identity of characters, objects, or scenes is anchored by real image references.


Overview

Gemini Omni brings together Google's deep knowledge of physics, narrative logic, biology, culture, and visual composition to produce contextually coherent videos. Rather than simple clip synthesis, the model reasons about scene dynamics, camera language, and temporal flow to produce results that feel intentional and cinematic.

With image inputs, the model extracts key visual features тАФ appearance, texture, structure, style тАФ and carries them faithfully into the generated video. This makes it well suited for character animation, product visualization, and style-guided generation.

The developer tier provides direct API access with full control over generation parameters including resolution, aspect ratio, duration, and random seed.


Key Capabilities

  • Image-guided generation тАФ Provide 1 to 7 reference images to anchor subjects, environments, or visual styles.
  • Subject consistency тАФ The model preserves key visual details from the reference images across the full video duration.
  • Rich prompt understanding тАФ Complement image references with a prompt of up to 20,000 characters describing actions, camera movements, lighting, and mood.
  • Multi-resolution output тАФ Generate at 720p, 1080p, or 4K.
  • Flexible aspect ratios тАФ 16:9 landscape or 9:16 portrait.
  • Controllable duration тАФ 4, 6, 8, or 10 seconds per generation.
  • Reproducible results тАФ Set a fixed seed to reproduce or iterate on a specific generation.

Input Parameters

ParameterTypeRequiredDefaultDescription
modelstringYesgoogle/gemini-omni-flash/image-to-video-developerModel identifier
promptstringYesтАФText description of the video. Max 20,000 characters.
imagesarrayYesтАФ1тАУ7 reference image URLs. Supported formats: PNG, JPEG, JPG, WebP. Max 20MB each.
durationintegerNo8Video length in seconds. Enum: 4, 6, 8, 10.
aspect_ratiostringNo16:9Output aspect ratio. Enum: 16:9, 9:16.
resolutionstringNo720pOutput resolution. Enum: 720p, 1080p, 4k.
seedintegerNo-1Random seed for reproducibility. -1 uses a random seed.

Image Input Notes

  • Accepts 1 to 7 images per request.
  • Supported codecs: PNG, JPEG, JPG, WebP.
  • Minimum image dimensions: 128├Ч128 pixels.
  • Each image must be under 20MB.

Use Cases

  • Character animation тАФ Bring a character photo or illustration to life with a text description of their action.
  • Product visualization тАФ Animate product images for marketing or e-commerce content.
  • Style transfer тАФ Feed a reference artwork or photograph to define the visual style of the generated video.
  • Scene composition тАФ Combine multiple reference images of different subjects to compose a coherent scene.
  • Storyboard-to-video тАФ Convert static storyboard frames into animated previews.

Pricing

Pricing is based on output resolution and video duration. Image inputs do not incur additional charges.

ResolutionFormulaExample (8s)
720p / 1080p$0.2 + duration ├Ч $0.1$1
4k$1 + duration ├Ч $0.1$1.8

Formula: (resolution == "4k" ? $1 : $0.2) + duration ├Ч $0.1

720p and 1080p are identically priced. The 0.2/0.2 / 1 term is a fixed base charge per generation; $0.1 is the per-second rate applied to the requested duration.

рд╕рдорд╛рди рдореЙрдбрд▓ рджреЗрдЦреЗрдВ

рд╣рд░ рдореАрдбрд┐рдпрд╛ AI рдХреЗ рд▓рд┐рдП рдПрдХ рд╣реА APIред

рд╕рднреА рдореЙрдбрд▓ рдПрдХреНрд╕рдкреНрд▓реЛрд░ рдХрд░реЗрдВ