APENAS DUAS SEMANAS | 20% DE DESCONTO no Seedream 5.0 Pro!
Início
Explorar
Google
Nano Banana 2 Lite
google/nano-banana-2-lite/reference-to-image
Nano Banana 2 Lite Reference-to-image
Imagem para Imagem

Nano Banana 2 Lite Reference-to-Image API by Google

google/nano-banana-2-lite/reference-to-image
Reference-to-image

Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) is Google's fastest, most cost-efficient image model, turning a source video clip plus a natural-language prompt (and optionally up to 14 reference images) into brand-new still images at low latency.

Google Nano Banana 2 Lite — Reference-to-Image (Video-to-Image)

Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image, gemini-3.1-flash-lite-image) is Google's fastest and most cost-efficient image model in the Nano Banana family. This variant is driven by a source video clip plus a natural-language prompt: the model uses the video's context as a multimodal reference, extracts its visual themes and key moments, and synthesizes brand-new still images from them — all with the low latency that defines the Lite tier.

It shares the underlying model with the text-to-image and edit variants — the difference is the input. Instead of a text prompt alone or a set of reference photos, you supply one video clip (and, optionally, up to 14 reference images) and describe the image you want. Video-to-image generation is a capability exclusive to the Nano Banana 2 / Gemini 3.1 Flash Image family, and Nano Banana 2 replaces the original Nano Banana (Gemini 2.5 Flash Image) with improved visual quality, stronger character consistency, and more legible in-image text — while being faster and cheaper.

🌟 Why it stands out

  • Video as a visual reference — Turn footage into stills. The model analyzes video frames to understand subjects, scenes, and key events, then generates images grounded in that context. See Google's video-to-image generation guide.
  • Flexible source input — Accepts a public YouTube URL or a direct HTTP video URL (up to 15 MB), with adjustable trim range and sampling FPS so you control which part of the clip is used.
  • Multimodal composition — Combine the video reference with up to 14 additional input images to blend subjects, transfer styles, or assemble scenes from multiple sources.
  • Strong character consistency — Preserves character identities and object fidelity drawn from the source video, so subjects stay recognizable.
  • Legible in-image text — Renders and localizes readable text directly within generated images for quick captioning and design work.
  • World knowledge — Understands scene structure and real-world context to keep generated frames coherent and plausible.
  • Fast and cost-efficient — The budget-friendly member of the Nano Banana family, built for high-volume generation at scale.

⚙️ How to use

  • Input: 1 source video_clip + a text prompt (optionally 1–14 reference image URLs)
  • Video source: public YouTube URL or HTTP video URL (HTTP video limited to 15 MB); configurable start/ends trim times (seconds) and fps (0–24)
  • Output: generated image (JPEG/PNG/WEBP), delivered as a URL (or BASE64 via the API)
  • Resolution: 1k
  • Aspect ratios: auto, 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, 4:1, 1:4, 8:1, 1:8
  • Thinking level: default, high, or minimal — controls how much internal reasoning the model performs before generating. Higher levels can improve quality on complex tasks at the cost of latency; minimal maximizes speed.
  • Works with natural prompts such as:
    • "Create a cinematic movie poster from the most dramatic moment in this clip."
    • "Generate a clean 16:9 video thumbnail highlighting the main subject."
    • "Produce a summary infographic of the key scenes in this video."
    • "Create new artwork inspired by the sunset scene, keeping the same character."

💰 Pricing

Pricing combines a flat per-image fee with a small per-request fee for supplying a video reference. When you pass a video clip, a single sku_video_clip charge is added on top of the per-image cost — it is a flat fee per request, regardless of clip length or number of frames sampled.

SKUDescriptionUnit Price
sku_1k1k image$0.04
sku_video_clipVideo clip reference (flat, per request)$0.035

💡 Best Use Cases

  • Video thumbnails — Generate high-quality, on-brand thumbnails from the footage itself.
  • Cinematic posters & key art — Turn a standout moment into a polished poster or promotional still.
  • Summary infographics — Distill a clip's key scenes into a single explanatory image.
  • Social media & content creation — Produce many still variations from a single video with minimal effort.
  • Concept art & storyboards — Derive new artwork inspired by specific video scenes while preserving subjects.

🔒 Content Authenticity

Generated images include C2PA content credentials and an imperceptible SynthID watermark by default, so outputs can be identified as AI-generated.

⚠️ Limitations

Optimized for speed and cost, this model may struggle with small faces, accurate spelling, and very fine details. Because output is grounded in a compressed video reference, fast-moving or low-quality footage can yield less precise results, and complex scenes may occasionally produce artifacts or disjointed compositions — for those cases, consider a higher-tier Nano Banana model.

📝 Notes

Please ensure your prompts and source video comply with Google's Safety Guidelines. If an error occurs, review your prompt, video clip, and input images for restricted content, adjust them, and try again.

Explorar Modelos Semelhantes

Uma API para toda a IA de mídia.

Explorar Todos os Modelos