Nano Banana Models

Google’s Nano Banana (Gemini 3 Image) series, featuring both standard and Pro models, combines deep semantic understanding with seamless integration for precise detail control. While the standard version delivers high-quality 1K outputs, Nano Banana Pro elevates professional workflows with versatile 1K/2K/4K resolution options with higher quality, making it the ideal solution for any creative or commercial application.

探索領先模型

Atlas Cloud 為您提供最新的行業領先創意模型。

NEW

文生圖

PRO

Nano Banana Pro Text-to-image

Nano Banana Pro is the next-generation Nano Banana image model, delivering sharper detail, richer color control, and faster diffusion for production-ready visuals.

Nano Banana Pro Edit

Nano Banana Pro Edit is an image editing tool built on the Nano Banana model family, designed for precise, AI-powered visual adjustments.

Nano Banana Text-to-image

Google's state-of-the-art image generation and editing model.

Nano Banana Edit

Google's state-of-the-art image generation and editing model.

From

$0.038/張

Nano Banana Models 的核心亮點

Atlas Cloud 為您提供業界領先的最新創意模型。

Photorealistic Quality

Generates crisp, high-resolution images with accurate lighting, textures, and detail for production use.

Fast, Lightweight Inference

Optimized architecture delivers rapid image generation on modest GPUs and edge hardware.

Fine-Grained Control

Supports styles, presets, and prompt controls so designers can quickly dial in the exact look they want.

Seamless Workflow Integration

Simple APIs and plugins connect Nano Banana to design tools, apps, and pipelines with minimal setup.

Cost-Efficient Creativity

Efficient diffusion kernels and smart caching keep generation costs low, so teams can experiment freely at scale.

Flexible Deployment Options

Flexible Deployment Options  Run in the cloud, on-prem, or in VPC environments.

峰值速度

最低成本

模態	描述
Nano Banana Pro T2I API(Text To Image)	Nano Banana Pro T2I API 提供領先業界的影像合成技術，將複雜的文字提示轉化為超逼真的視覺效果。它支援 1K、2K 和 4K 解析度，專為高保真創意資產、專業廣告以及每一個像素都至關重要的高級數位藝術而設計。
Nano Banana Pro Edit(Image To Image)	Nano Banana Pro Edit API 提供具有外科手術般精準度的進階圖像到圖像轉換功能。它支援高達 4K 的高解析度風格轉換和內容修改，確保迭代設計和高階修圖工作流程具有專業級的一致性和細節。
Nano Banana T2I API(Text To Image)	Nano Banana T2I API 為快速文字生成圖像提供了一種均衡且高效能的解決方案。它針對速度和可靠性進行了優化，使開發者能夠擴展社交媒體、網絡資產和動態行銷活動的視覺內容創作，並保持輸出的一致性。
Nano Banana Edit API(Image To Image)	Nano Banana Edit API 簡化了圖像到圖像的編輯流程，提供基於提示詞（prompt）的可靠修改。它是進行大量內容更新和靈活視覺實驗的理想工具，在這些場景中，效率和可靠的性能至關重要。

Nano Banana Models 新功能 + 展示

將先進模型與 Atlas Cloud 的 GPU 加速平台相結合，為圖像和視頻生成提供無與倫比的速度、可擴展性和創意控制。

使用 Nano Banana Pro API 實現完美角色一致性

在複雜場景中維持完美的視覺一致性，並能夠同時追蹤多達5個獨特角色。透過分析細微的身體特徵，Nano Banana Pro 確保了跨多次生成的角色外觀穩定性，使其成為連貫視覺敘事和系列化創意內容的首選工具。

使用 Nano Banana Pro API 進行超高畫質渲染

體驗原生 2K 輸出和先進 AI 驅動的 4K 升頻能力帶來的無與倫比的視覺清晰度。這種雙層渲染方法可生成具有清晰細節和豐富紋理的專業級資產，滿足高端商業設計和大型數位顯示器所需的嚴格品質標準。

使用 Nano Banana Pro API 進行全球多語言文字合成

實現完美的排版整合，支援 100 多種語言的精確文字渲染。從複雜的書寫系統到風格化字型，該模型消除了常見的 AI 文字瑕疵，為全球品牌推廣、在地化行銷素材和高傳真平面設計提供了無縫解決方案。

使用 Nano Banana Pro API 進行進階多影像合成

藉由融合多達 14 張參考影像來引導風格、結構與內容，解鎖精密的創意工作流程。這種強大的多層次融合功能，讓使用者能以極高的精確度合成複雜的視覺概念，為專業的情緒板製作與精細的概念藝術提供極致的靈活性。

使用 Nano Banana Models 可以做什麼

探索使用該模型家族可以構建的實際應用場景和工作流 — 從內容創作、自動化到生產級應用。

透過 Nano Banana API 實現無縫角色一致性

Nano Banana API 賦能創作者和開發者構建複雜的敘事世界，能夠同時保持多達 5 個獨特角色的完美視覺一致性。該 API 非常適合圖畫小說、連載故事和 IP 開發，可在各種環境和光照條件下保留複雜的面部特徵、服裝細節和風格特點——確保您整個創意項目的完美連續性。

使用 Nano Banana API 實現影棚級商業設計

為了具高影響力的行銷與全球品牌推廣，Nano Banana 透過原生 2K 渲染與先進的 4K AI 放大技術生成超清晰影像。此功能搭配超過 100 種語言的完美文字渲染，適合專業廣告、在地化行銷視覺效果及頂級產品設計。對於需要清晰排版與高保真紋理，以用於大規模數位與印刷展示的品牌而言，這是終極解決方案。

透過 Nano Banana API 進行複雜多參考合成

Nano Banana 支援精密的視覺工作流程，允許使用者融合多達 14 張不同的參考圖像，以深度影響風格、結構和構圖。此應用案例專為需要從多重來源合成複雜視覺創意的專業概念藝術家和世界建構者而設計。透過將多樣化的參考層與精確的提示詞控制相結合，該 API 為高階情緒板製作和複雜的概念藝術提供了無與倫比的靈活性。

模型對比

查看不同廠商的模型表現 — 對比效能、價格和獨特優勢，做出明智決策。

模型	參考圖像限制	輸出數量	解析度	縱橫比
Nano Banana Pro	10	1	4K, 2K, 1K	1:1 3:2 2:3 3:4 4:3 4:5 5:4 9:16 16:9 21:9
Nano Banana 2	14	1	4K, 2K, 1K	1:1 3:2 2:3 3:4 4:3 4:5 5:4 9:16 16:9 21:9
Seedream 5.0 Lite	14	1~15	2K~4K+	1:1 3:2 2:3 3:4 4:3 4:5 5:4 9:16 16:9 21:9
Qwen-Image	3	1~6	512P~2K	Width[512, 2048]px; Height[512, 2048]px
Wan 2.6 I2I(Image To Image)	4	1	580P~1080P+	1:1 3:2 2:3 3:4 4:3 4:5 5:4 9:16 16:9 21:9 9:21

如何在 Atlas Cloud 上使用 Nano Banana Models

幾分鐘即可上手 — 按照以下簡單步驟，透過 Atlas Cloud 平台整合和部署模型。

建立 Atlas Cloud 帳戶

在 atlascloud.ai 註冊並完成驗證。新用戶可獲得免費額度，用於探索平台和測試模型。

為何在 Atlas Cloud 使用 Nano Banana Models

將先進的 Nano Banana Models 模型與 Atlas Cloud 的 GPU 加速平台相結合，提供無與倫比的效能、可擴展性和開發體驗。

效能與靈活性

低延遲：
GPU 最佳化推理，實現即時回應。

統一 API：
一次整合，暢用 Nano Banana Models、GPT、Gemini 和 DeepSeek。

透明定價：
按 Token 計費，支援 Serverless 模式。

企業與規模

開發者體驗：
SDK、資料分析、微調工具和模板一應俱全。

可靠性：
99.99% 可用性、RBAC 權限控制、合規日誌。

安全與合規：
SOC 2 Type II 認證、HIPAA 合規、美國資料主權。

關於 Nano Banana Models 的常見問題

Nano Banana (Gemini 3 Flash Image) 是專為快速生成高品質 1K 影像而最佳化的標準模型。Nano Banana Pro 則是專為專業工作流程設計的進階版本，提供卓越的細節控制、原生 2K 渲染以及 4K 升頻功能。

對於複雜的構圖和風格轉換，Nano Banana Pro 支援最多 10 張參考圖像的多模態輸入。如果您希望輸入超過 10 張參考圖像並獲得更好的輸出品質，可以嘗試 Nano Banana 2（參考圖像限制：14）。

探索更多系列

Happy Horse 1.0

HappyHorse-1.0 is a unified multimodal AI video generation model that climbed to the top of the Artificial Analysis Video Arena blind-test leaderboard for both text-to-video and image-to-video generation. CNBC Alibaba Group confirmed ownership of HappyHorse, developed under its Alibaba Token Hub (ATH) business unit, where it leads benchmarks outperforming ByteDance's Seedance 2.0 and others. Caixin Global Led by Zhang Di — the former VP of Kuaishou who architected Kling AI — the 15-billion parameter model generates 1080p video with synchronized audio in a single pass using a unified transformer architecture that bypasses the multi-stage pipelines used by every major competitor.

檢視系列

Seedance 2.0 Models

Seedance 2.0（by Bytedance） is a multimodal video generation model that redefines "controllable creation," moving beyond the limitations of text or start/end frames. It supports quad-modal inputs—text, image, video, and audio—and introduces an industry-leading "Universal Reference" system. By precisely replicating the composition, camera movement, and character actions from reference assets, Seedance 2.0 solves critical issues with character consistency and physical coherence, empowering creators to act as true "directors" with deep control over their output.

檢視系列

GPT Image 2 Models

GPT Image 2 is a state-of-the-art multimodal foundation model engineered for exceptional text-to-image generation with unprecedented photorealism and creative versatility. Developed by OpenAI as the evolution of the DALL-E lineage, it transforms detailed natural language descriptions into hyper-realistic imagery at up to 4K resolution. With proprietary "Neural Rendering Engine" technology for precise visual control, GPT Image 2 delivers studio-quality results with accurate anatomy, lighting, and composition—making it the premier AI tool for professional creators, enterprises, and developers demanding production-ready visual assets.

檢視系列

Wan2.7 Models

Launching this March, Wan2.7 is the latest powerhouse in the Qwen ecosystem, delivering a massive upgrade in visual fidelity, audio synchronization, and motion consistency over version 2.6. This all-in-one AI video generator supports advanced features like first-and-last frame control, 3x3 grid synthesis, and instruction-based video editing. Outperforming competitors like Jimeng, Wan2.7 offers superior flexibility with support for real-person image inputs, up to five video references, and 1080P high-definition outputs spanning 2 to 15 seconds, making it the premier choice for professional digital storytelling and high-end content marketing.

檢視系列

Veo3.1 Models

Google DeepMind’s Veo 3.1 represents a paradigm shift in AI video generation, empowering creators with director-level narrative control and cinematic-grade audio quality that seamlessly integrates with its enhanced visual realism. By bridging the gap between imaginative concepts and photorealistic execution, this advanced model offers a transformative solution for a wide range of application scenarios, from professional filmmaking and high-end advertising to immersive digital content creation.

檢視系列

ERNIE Image Models

ERNIE-Image is an open-weight text-to-image model developed by the ERNIE-Image Team at Baidu, built on a single-stream Diffusion Transformer (DiT) with 8B parameters and paired with a lightweight Prompt Enhancer that rewrites short prompts into richer, more structured descriptions before passing them to the diffusion backbone. NYU Shanghai RITS Released on April 15, 2026 under the Apache 2.0 license, it transforms natural language descriptions into detailed imagery with particular strength in text rendering and structured layout generation. ERNIE-Image is designed not only for strong visual quality, but for controllability in practical generation scenarios where accurate content realization matters as much as aesthetics — making it well-suited for commercial posters, comics, multi-panel layouts, and other content creation tasks that require both visual quality and precise control.

檢視系列

GPT Image Models

The GPT Image Family is OpenAI's latest suite of multimodal image generation and editing models, built on the powerful GPT architecture. This family includes three tiers — GPT Image-1, GPT Image-1.5, and GPT Image-1 Mini — each available in both Text-to-Image and Image-to-Image variants. Combining GPT's world-class language understanding with DALL·E-class visual synthesis, these models deliver exceptional prompt adherence, photorealistic rendering, and creative versatility across illustration, photography, design, and visualization tasks. The series offers flexible pricing and quality tiers to match any workflow — from rapid prototyping and high-volume content production to professional-grade final deliverables. Whether you need ultra-fast iterations at minimal cost or maximum quality for brand campaigns, the GPT Image Family has a solution tailored to your needs.

檢視系列

Nano Banana2 Models

Nano Banana 2 (by Google), is a generative image model that perfectly balances lightning-fast rendering with exceptional visual quality. With an improved price-performance ratio, it achieves breakthrough micro-detail depiction, accurate native text rendering, and complex physical structure reconstruction. It serves as a highly efficient, commercial-grade visual production tool for developers, marketing teams, and content creators.

檢視系列

Seedream5.0 Models

Seedream 5.0, developed by ByteDance’s Jimeng AI, is a high-performance AI image generation model that integrates real-time search with intelligent reasoning. Purpose-built for time-sensitive content and complex visual logic, it excels at professional infographics, architectural design, and UI assistance. By blending live web insights with creative precision, Seedream 5.0 empowers commercial branding and marketing with a seamless, logic-driven workflow that turns sophisticated data into stunning, high-fidelity visuals.

檢視系列

Kling3.0 Models

Kuaishou’s flagship video generation suite, Kling 3.0, features two powerhouse models—Kling 3.0 (Upgraded from Kling 2.6) and Kling 3.0 Omni (Kling O3, Upgraded from Kling O1)—both offering high-fidelity native audio integration. While Kling 3.0 excels in intelligent cinematic storytelling, multilingual lip-syncing, and precision text rendering, Kling O3 sets a new standard for professional-grade subject consistency by supporting custom subjects and voice clones derived from video or image inputs. Together, these models provide a comprehensive solution tailored for cinematic narratives, global marketing campaigns, social media content, and digital skit production.

檢視系列

GLM LLM Models

GLM is a cutting-edge LLM series by Z.ai (Zhipu AI) featuring GLM-5, GLM-4.7, and GLM-4.6. Engineered for complex systems and long-horizon agentic tasks, GLM-5 outperforms top-tier closed-source models in elite benchmarks like Humanity’s Last Exam and BrowseComp. While GLM-4.7 specializes in reasoning, coding, and real-world intelligent agents, the entire GLM suite is fast, smart, and reliable, making it the ultimate tool for building websites, analyzing data, and delivering instant, high-quality answers for any professional workflow.

檢視系列

Open AI Model Families

Explore OpenAI’s language and video models on Atlas Cloud: ChatGPT for advanced reasoning and interaction, and Sora-2 for physics-aware video generation.

檢視系列

Open AI Model Families

Explore OpenAI’s language and video models on Atlas Cloud: ChatGPT for advanced reasoning and interaction, and Sora-2 for physics-aware video generation.

檢視系列

300+ 模型，即刻開啟，

探索全部模型

Nano Banana Models

探索領先模型

Nano Banana Pro Text-to-image

Nano Banana Pro Edit

Nano Banana Text-to-image

Nano Banana Edit

Nano Banana Models 的核心亮點

Photorealistic Quality

Fast, Lightweight Inference

Fine-Grained Control

Seamless Workflow Integration

Cost-Efficient Creativity

Flexible Deployment Options

峰值速度

Nano Banana Models 新功能 + 展示

使用 Nano Banana Pro API 實現完美角色一致性

使用 Nano Banana Pro API 進行超高畫質渲染

使用 Nano Banana Pro API 進行全球多語言文字合成

使用 Nano Banana Pro API 進行進階多影像合成

使用 Nano Banana Models 可以做什麼

透過 Nano Banana API 實現無縫角色一致性

使用 Nano Banana API 實現影棚級商業設計

透過 Nano Banana API 進行複雜多參考合成

模型對比

如何在 Atlas Cloud 上使用 Nano Banana Models

建立 Atlas Cloud 帳戶

為何在 Atlas Cloud 使用 Nano Banana Models

效能與靈活性

企業與規模

關於 Nano Banana Models 的常見問題

探索更多系列

Happy Horse 1.0

Seedance 2.0 Models

GPT Image 2 Models

Wan2.7 Models

Veo3.1 Models

ERNIE Image Models

GPT Image Models

Nano Banana2 Models

Seedream5.0 Models

Kling3.0 Models

GLM LLM Models

Open AI Model Families

Happy Horse 1.0

Seedance 2.0 Models

GPT Image 2 Models

Wan2.7 Models

Veo3.1 Models

ERNIE Image Models

GPT Image Models

Nano Banana2 Models

Seedream5.0 Models

Kling3.0 Models

GLM LLM Models

Open AI Model Families

300+ 模型，即刻開啟，

Join our Discord community