

MiniMax H3 API 開放 MiniMax 的通用多模態影片模型,可將文字、圖片、影片與音訊作為同一個脈絡來理解,而不是一次只處理一項任務。片段長度為 5 到 15 秒、24 FPS,支援從 21:9 到 9:16 的各種長寬比;只需一個 prompt,就能替換角色、更換背景、改寫對白,或從參考片段複製聲音。Atlas Cloud 透過單一 OpenAI-compatible endpoint 提供所有功能。立即開始建置。
Atlas Cloud 為您提供最新的行業領先創意模型。
每個 MiniMax H3 API endpoint 對文字、圖片、影片與音訊的讀取方式都不同,請瀏覽下方列項,將模態對應到你正在建置的內容。
| 模態 | 說明 |
|---|---|
| MiniMax H3 T2V API(文字轉影片) | 撰寫最多 7,000 個字元的提示詞,模型會回傳一段 5 到 15 秒、24 FPS、1440p 的影片,並在同一次生成中產生原生立體聲。每次請求都可在 21:9、16:9、4:3、1:1、3:4 和 9:16 之間設定長寬比,因此一次呼叫即可同時涵蓋電影感預告片與直式社群短片。 |
| MiniMax H3 I2V API(圖片轉影片) | 一張圖片可設定開場畫面,兩張圖片則可鎖定首尾畫面,由 H3 依輸入圖片的長寬比補齊中間動態。支援每邊 256 到 5760 像素的來源素材,讓主視覺、產品靜態圖與分鏡圖都能輕鬆動起來。 |
| MiniMax H3 Omni Reference API(參考轉影片) | 需要多個來源協同運作嗎?在 12 個檔案的上限內,單一請求最多可放入九張圖片、三段影片剪輯與三條音軌,並作為同一個多模態脈絡讀取。你可以在多個鏡頭之間延續角色、攝影機運動或聲音音色,也可以透過替換主體、背景與台詞來編輯既有片段。 |
MiniMax H3 API 回傳的每支片段最長可達 15 秒,具備 1440p 與原生立體聲,可由文字、圖片、影片與聲音參考的任意組合生成;同一次呼叫也能編輯你既有影片中的角色、場景與對白。
一次 MiniMax H3 API 請求最多可接受 9 張參考圖片、3 支影片片段與 3 條音軌,總上限為 12 個檔案。模型不會把它們當成彼此分離的欄位,而是將文字、圖片與聲音視為同一個脈絡,把照片中的角色、片段中的運鏡,以及音軌中的情緒整合進同一個場景。已有素材的團隊可以更直接地產出完成度高的鏡頭。
每個結果都自帶聲音。對白、環境音與音樂會在以 24 fps 算繪畫面的同一流程中,以原生立體聲生成,因此後續不需要再配樂或配音。由於時間點是在生成鏡頭時一併決定,腳步聲、語音與剪接點都能與動作保持同步。短廣告與社群短片產出後即可發布。
想把貓換成狗、替換綠幕,或重寫一句對白嗎?MiniMax H3 API 可將這類編輯套用到你提供的影片,涵蓋角色、物件、背景、燈光與特效;未提及的部分則會盡量貼近原始素材。提示詞最多可達 7000 字元,因此可在一次處理中疊加十多項變更。這讓已核准剪輯版本的反覆調整變得切實可行。
將聲音樣本與圖片或影片一併送入,生成角色就能以該音色說話。每次請求最多允許 3 個音訊參考,每個長度介於 2 到 15 秒,且音訊必須一律搭配視覺輸入,不能單獨提交。既有對白也能被替換,並可調整演出語氣以相互匹配。系列內容可在每一集維持同一個可辨識的聲音。
片段長度介於 5 到 15 秒;1440p 模式會在 16:9 到 9:16 之間將短邊設為 1440 pixels,或在較寬比例下達到約 3.7 megapixels,例如 21:9 的 2976 × 1248。可選擇 6 種比例,從電影感的 21:9 到直式 9:16,MiniMax H3 API 也可以替你自動選擇。這個範圍無需切換模型,就能涵蓋預告片、產品循環影片與直式短劇。
如果來源檔直接來自拍攝設備或剪輯時間軸,可以原樣送入:H.264 與 H.265 影片、JPG、PNG、WEBP、HEIC 與 HEIF 靜態圖片,以及 WAV 與 MP3 音訊。單檔限制為影片 50MB、圖片 30MB、音訊 15MB;以 URL 傳遞素材則可讓請求保持在 64MB body 限制內。在 Atlas Cloud 上,整套流程可透過一把 OpenAI-compatible key 執行,並採隨用隨付計費。
此組中的每段片段都來自同一段完全相同的提示詞,分別送至 MiniMax H3 API 以及 Atlas Cloud 上託管的另外兩個影片模型,因此可在不改動任何一個字的情況下,比較動態、聲音與指令遵循度。
15 秒,16:9 橫向短影片。深夜自助洗衣店的真人實拍影像,融合手繪發光動畫,形成混合媒材畫面。一間小型自助洗衣店,螢光燈微微閃爍;店內有運轉中的洗衣機、塑膠洗衣籃和一張老舊長椅,地上躺著一隻襪子。整個空間很安靜,帶著淡淡的懷舊氛圍。畫面具有單手手持手機拍攝的質感,明顯有攝影機晃動;白色螢光燈讓曝光在明暗之間起伏;玻璃表面帶有環境反射;鏡頭靠近物體時會有對焦延遲。畫面不應像商業廣告那樣精緻有序——整體感覺應像真實的紀錄片式隨手捕捉,彷彿你深夜偶然闖入,追逐某種奇異、夢境般的幻象時順手拍下。
Generated with MiniMax H3 on Atlas Cloud
Generated with Seedance 2.0 on Atlas Cloud
Generated with Wan-2.7 on Atlas Cloud
第一人稱視角 · 視線高度 · 手持遊戲鏡頭 場景:鏡頭模擬玩家操作一款現代戰爭 FPS 遊戲,雙手持突擊步槍,沿著軍事基地外圍緩慢推進。玩家沿著掩體旁的道路向前移動,準星掃過前方通道;短暫停頓後,朝遠處目標射出幾發子彈,接著繼續向前推進——就像一般玩家的實際遊玩畫面。 光線:現代軍事基地的冷色調自然光,與煙霧和槍口火光交織。畫面寫實且清晰,金屬武器、戰場塵土與霧霾都具有 AAA 遊戲級質感。 運鏡:隨著玩家移動,鏡頭帶有輕微的手持晃動——先是緩慢前進,接著小幅左右掃視觀察,開火時有細微後座力震動,最後再穩定地繼續向前推進。
Generated with MiniMax H3 on Atlas Cloud
Generated with Seedance 2.0 on Atlas Cloud
Generated with Wan-2.7 on Atlas Cloud
從品牌影片、直式短劇,到產品剪輯、遊戲視覺與既有素材的精準編修,MiniMax H3 API 透過一次多模態請求涵蓋各種情境,並回傳內建原生立體聲音訊的影片。
將分鏡畫面與鏡頭清單放進一次呼叫,即可產出 1440p 預告片、TVC 廣告與時尚活動影片,並以 24 FPS 呈現。每個結果都隨附原生立體聲音訊,讓品牌團隊可直接審看完成版剪輯。
直式 9:16 輸出可涵蓋劇本化戲劇場景,從古裝懸疑到家庭衝突,並在同一次流程中完成對白配音。打造短劇片庫的工作室,不必預約演員或攝影棚,就能取得 15 second 的開場鉤子。
如果某個鏡頭需要特定臉孔與動作,最多可用九張圖片、三段影片與三段音訊片段引導一次呼叫。角色身分、攝影運鏡與聲音音色都能在各集之間保持一致。
事後需要修改嗎?既有素材可透過提示詞進行編修:替換主體、更換背景、調整光線,或重寫一句口白,而不必重拍任何一格畫面。
產品照片可以變成動態內容:一張參考圖即可轉成 360 degree 展示、功能解說影片或付費社群短片。從 21:9 到 9:16 的長寬比,讓同一組素材可供應所有版位。
風格化輸出足以支援遊戲 CG、角色 PV、動漫片頭與介面展示,讓選單、HUD 元素與文字疊加都維持可讀性。美術團隊可在正式製作前用它驗證概念。
將 MiniMax H3 API 與 Atlas Cloud 上託管的其他影片模型並列比較,在選定 endpoint 前,先了解輸入模態、參考檔案限制、長度、解析度與音訊輸出的實際差異。
| 模型 | 輸入模態 | 參考檔案上限 | 輸出時長 | 最高解析度 | 原生音訊 |
|---|---|---|---|---|---|
| MiniMax H3 | 文字、圖片、影片、音訊 | 9 張圖片、3 支影片、3 段音訊片段,總計 12 個檔案 | 5s 至 15s | 1440p at 24 FPS | √ 每個輸出皆提供原生立體聲音訊 |
| Seedance 2.0 Reference-to-Video | 文字、圖片、影片、音訊 | 9 張圖片、3 支影片、3 段音訊片段 | 最長 15s | 720p | √ 單次生成即可產生立體聲對白、音效與音樂 |
| Veo3.1 Reference-to-video | 文字與圖片 | 3 張參考圖片 | 4s、6s 或 8s | 僅限 8s 長度支援 4K | √ 對白、環境音與音效可對齊時間軸 |
| Wan-2.7 Reference-to-video | 文字、圖片、影片、音訊 | 5 張圖片或影片片段,另加一段語音片段 | 2s 至 15s | 1080p | √ 可生成音樂與音效,或使用你自己的音訊驅動對嘴同步 |
| Kling v3.0 Pro Image-to-Video | 文字與圖片 | - | 最長 15s | 1080p | √ 支援五種語言的多語對白與對嘴同步 |
幾分鐘即可上手 — 按照以下簡單步驟,透過 Atlas Cloud 平台整合和部署模型。
在 atlascloud.ai 註冊並完成驗證。新用戶可獲得免費額度,用於探索平台和測試模型。
將先進的 MiniMax H3 模型與 Atlas Cloud 的 GPU 加速平台相結合,提供無與倫比的效能、可擴展性和開發體驗。
低延遲:
GPU 最佳化推理,實現即時回應。
統一 API:
一次整合,暢用 MiniMax H3、GPT、Gemini 和 DeepSeek。
透明定價:
按 Token 計費,支援 Serverless 模式。
開發者體驗:
SDK、資料分析、微調工具和模板一應俱全。
可靠性:
99.99% 可用性、RBAC 權限控制、合規日誌。
安全與合規:
SOC 2 Type II 認證、HIPAA 合規、美國資料主權。
The MiniMax H3 API gives developers programmatic access to MiniMax H3, an open general purpose multimodal video model that treats text, images, video, and audio as one shared context. Rather than splitting generation, editing, and reference into separate task models, H3 reads the full input set and returns a finished clip with sound. On Atlas Cloud it runs behind a single API key with pay-as-you-go pricing.
Brand films, trailers, vertical short drama, product and ecommerce spots, game and UI motion demos, and stylized animation all sit inside its range. Because the model handles on screen text, subtitles, and brand assets, teams also use it for concept validation, storyboard previews, and visual pitches before committing production budget.
Create an Atlas Cloud account, generate an API key, then send a prompt plus any reference files to the video generation endpoint and poll for the finished result. Passing media as hosted URLs is recommended over inline uploads, since the request body is capped at 64MB. Start building today.
Billing is pay-as-you-go, so you pay per call instead of buying a subscription or a seat license. Cost tracks what you actually render, which means resolution and clip length drive the total for a batch. Check the model page for the current per generation rate before planning large volume runs.
Yes. Every H3 result is delivered with sound in native stereo, so dialogue, effects, and ambience arrive in the same pass as the picture. Supply a reference audio clip and the model can carry that timbre onto a character, which removes a separate voice synthesis step from the pipeline.
Clips run from 5 to 15 seconds at 24 FPS. In 1440p mode the short side renders at 1440 pixels for ratios between 16:9 and 9:16, and outside that band the frame holds roughly 3.7M pixels in total, for example 2976 by 1248 at 21:9. Text to video and omni reference requests accept 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, while first and last frame requests inherit the aspect ratio of the input image.
Up to nine images, three video segments, and three audio clips can travel with one prompt, capped at twelve files in total. Video and audio each stay within 15 seconds combined, and audio must accompany an image or a video rather than arrive on its own. Prompts reach 7000 characters, which leaves room for shot by shot direction covering camera, performance, and sound.
Accepted inputs include H.264 and H.265 video, JPG, JPEG, PNG, WEBP, HEIC, and HEIF images, and WAV or MP3 audio, with AAC or MP3 for the audio track inside a video file. Per file ceilings are 50MB for video, 30MB for images, and 15MB for audio. Because the request body is limited to 64MB, hosted URLs remain the safer route for heavy assets.
Editing is one of its core modes. Send a source clip with instructions and the MiniMax H3 API can swap characters or objects, replace backgrounds and lighting, layer in visual effects, and rewrite dialogue while keeping untouched regions stable. Compound instructions are handled in one request, so several changes land together instead of across repeated round trips.
Most video models take one prompt plus one image and return a silent clip. H3 instead consumes a mixed set of references, reads character, motion, camera, and sound intent across all of them, then returns a clip with native audio. If your pipeline currently stitches a video model, a voice model, and an editor together, H3 collapses those stages into a single call.
指南、教學與產品動態,助你充分發揮 Atlas Cloud 的價值。