
Kling 1.6 是快手的视频生成模型,旨在将提示词和参考图像转化为连贯的短视频。它提升了对运动、时序动作和镜头运动的响应能力,同时增强了主体一致性、色彩准确度、光照动态效果和细节呈现。通过 Atlas Cloud,使用一个兼容 OpenAI 的密钥和透明的按量计费价格即可接入该模型系列。立即开始构建。
Kling 1.6 由 Kuaishou 开发。Atlas Cloud(由 Atlas Cloud AI LLC 运营)仅提供接入服务,并不拥有该模型。所有商标均归其各自所有者所有。
了解各个 Kling 1.6 Endpoint 如何处理文本或图像输入,以及其在运动表现、连贯性、细节和效率方面经过验证的优势分别适用于哪些场景。
| 模态 | 描述 |
|---|---|
| Kling 1.6 Multi I2V Pro API (Multi Image To Video) | 将图像转换为视频,支持多个主体,并提升画面连贯性和高级运动追踪的准确性。此 Pro Endpoint 适用于复杂场景,尤其适合需要协调主体运动和保持视觉一致性的任务。 |
| Kling 1.6 Multi I2V Standard API (Multi Image To Video) | 针对注重成本效益的多主体生成,此 Endpoint 可将图像转换为视频,并在速度与细节之间取得平衡。适用于基础场景动画以及包含多个主体的常规内容制作。 |
| Kling 1.6 T2V Standard API (Text To Video) | 将文本提示词转换为短视频,呈现稳定的运动效果,并可靠地遵循提示词要求。此入门级 Endpoint 适用于概念可视化、社交媒体内容和简单的提示词驱动序列。 |
| Kling 1.6 I2V Pro API (Image To Video) | 以静态图像为基础,Pro Endpoint 可生成运动衔接更流畅、纹理更加逼真的视频。适用于精致的产品视觉素材、角色镜头和电影感图像动画。 |
| Kling 1.6 I2V Standard API (Image To Video) | 选择此轻量级 Endpoint,通过基础生成流程将图像转换为视频。其低成本定位使其适合简单动画、早期创意测试和大批量内容制作。 |
Kling 1.6 通过一个 Atlas Cloud API,整合文本、单图和最多四张参考图工作流,支持生成 5 或 10 秒视频、可控的提示词遵循度、可选负面提示词、指定宽高比,以及 Standard 或 Pro 访问权限。
Kling 1.6 通过专用变体支持文生视频和图生视频。你可以从文字描述的场景开始,也可以让源图像动起来,再使用最多 2,500 个字符的提示词引导动作。Standard 模型优先兼顾成本效率,而 Pro 图像变体则注重更流畅的动作融合和更逼真的纹理表现。这一系列能力适用于概念测试、社交媒体短片和精致的视觉片段制作。
使用 Multi I2V Standard 或 Pro 变体时,可上传 1 到 4 张参考图。模型会将其作为视觉参考,同时由提示词引导主体、动作和场景发展。Pro 针对更强的多主体一致性和更准确的动作跟踪进行了优化,适合角色互动、产品组合,以及需要将多个视觉元素融合在一起的场景。
使用 I2V Pro 变体设置首张图像和可选的结束图像,从而控制片段的起始和落点。每张图像可以是 JPG、JPEG 或 PNG 格式,大小不超过 10 MB,且至少为 300 × 300 像素。在这两个视觉锚点之间添加动作指令。这项控制功能尤其适合规划好的转场、姿势变化和产品展示。
流畅动作是 Pro 图像变体的核心重点之一,其 Atlas Cloud 配置中记录了升级后的融合能力和更逼真的纹理表现。0 到 1 的引导强度可调节提示词遵循度,而负面提示词则有助于避开不需要的元素。使用这些控制项,可以在自然运动与严格定向的动作之间取得平衡。这种组合适合要求较高的角色、织物和镜头运动场景。
Kling 1.6 全系列均支持选择 5 或 10 秒的输出时长。文本和多图变体还支持 16:9、9:16 和 1:1 宽高比,因此同一工作流即可面向宽屏、竖屏或方形版位。2,500 个字符的提示词上限,为主体、动作、光照和镜头指令留下了充足空间。这些选项帮助团队规划适配不同平台的素材,而无需更换模型系列。
通过一个 Atlas Cloud 视频生成 API 和按量付费模式,即可运行所列出的所有 Kling 1.6 变体。Standard 变体采用经过验证的原始基础价格,每次运行 $0.056;Pro 变体每次运行 $0.098。根据成本、细节和动作质量之间所需的平衡选择合适层级。这种配置支持快速试验和生产流水线,无需分别集成不同服务商。
See how Kling 1.6 and two alternative video models interpret identical prompts across realistic action and stylized storytelling.
A cinematic 8–10 second miniature live-action sequence inside a glassblowing workshop at midnight: a young artisan continuously rotates a blowpipe as a white-hot glass bubble rapidly expands, its perfectly round form anchoring the composition. Begin with an extreme macro orbit around the spinning molten glass, capturing transparent amber filaments stretching, viscous surface tension, heat shimmer, tiny sparks, and realistic internal refraction. The swelling bubble suddenly slips from the pipe, strikes the silver-gray metal table, and rolls fast; drop into a table-level high-speed tracking shot alongside it as it wobbles, deforms, sheds glowing threads, and reflects cobalt-blue moonlight against the furnace-orange glow. Whip-pan to the alarmed artisan lunging across the bench and catching the runaway glass with a wet wooden paddle at the last instant—an explosive hiss sends physically accurate steam swirling through the frame. As the steam clears, reveal the glass miraculously frozen into a small transparent pufferfish with delicate glass fins and consistent internal bubbles; it gently puffs its cheeks once, a playful final beat. Seamless continuous motion, coherent subject transformation, realistic glass viscosity, collisions, thermal glow, sparks, refraction, caustics, heat distortion, and steam physics; layered amber, cobalt blue, and silver-gray palette; furnace orange key light, cool moon-blue rim light, shallow depth of field, tactile high-end practical miniature filmmaking, photoreal cinematic texture, no slow motion, no static filler, no screens, software interfaces, dashboards, progress bars, charts, captions, text, logos, or watermarks. Synchronized audio: roaring furnace, rotating pipe scrape, sharp metallic clink, rolling glass rattle, urgent footstep, loud wet hiss, then a tiny crystalline puff; tense percussive rhythm ending on a whimsical glass chime. 16:9 aspect ratio.
Generated with Kling v1.6 Multi i2v Pro on Atlas Cloud
Generated with Seedance 2.0 Image-to-Video on Atlas Cloud
Generated with Kling v1.6 i2v Standard on Atlas Cloud
A tense 9-second micro-story in the cramped back kitchen of a late-night Hong Kong noodle shop: a young chef urgently rescues a ramen order, moving continuously with precise, believable hand choreography. Begin with an extreme low-angle lateral tracking shot skimming across the flour-dusted cutting board as he snaps his wrists and throws a long bundle of noodles high into the air; chase the twisting strands through dense steam, with individual noodles flexing, stretching, and naturally occluding his hands and hanging cookware. Whip-pan to a tight stove-side angle as the noodle bundle nearly drops into a licking gas flame; at the last instant he lunges forward and catches it cleanly in a wire skimmer, the mesh bending under its weight, then pivots in one fluid motion and plunges the noodles into violently boiling broth, sending realistic droplets, bubbles, and oily ripples across the pot. Snap to an overhead top-down shot as the noodles unfurl into a neat blooming spiral in the soup; through the service hatch, waiting diners lean in and burst into delighted applause while the chef flashes a breathless grin. Maintain exact continuity of the same chef, clothing, utensils, noodle bundle, kitchen geography, and motion across every cut. Tight deep staging constantly coordinates the chef, noodles, flames, pots, and foreground utensils; tactile layers of airborne flour, wet tile reflections, glistening oil, condensation, and rolling steam. Photorealistic Hong Kong cinema aesthetic, handheld kinetic energy, crisp natural motion blur, warm tungsten-orange practical lights clashing with cool cyan ceramic tiles, rich contrast, subtle 35mm film grain, realistic skin and food texture. Synchronized sound: knife-board clatter, gas flame roar, rushing steam, skimmer clang, boiling broth splash, then a sharp burst of applause; fast percussive kitchen rhythm, no dialogue. No slow motion, frozen poses, empty establishing shots, montage gaps, jumpy continuity, extra fingers, malformed hands, duplicated limbs, broken utensils, clipping, teleporting noodles, rubbery motion, impossible fluid behavior, floating objects, UI, screens, dashboards, progress bars, charts, captions, subtitles, logos, or visible text. 16:9 aspect ratio.
Generated with Kling v1.6 Multi i2v Pro on Atlas Cloud
Generated with Seedance 2.0 Image-to-Video on Atlas Cloud
Generated with Kling v1.6 i2v Standard on Atlas Cloud
Kling 1.6 可将提示词和源图像转化为短视频创意,适用于分镜、产品营销、多主体场景、动画艺术作品、社交媒体多版本内容和角色叙事。
将文字概念转化为短视频草稿,实现稳定运动和可靠的提示词对齐。导演、代理机构和产品团队可以在投入完整制作前,预览节奏、动作和视觉方向。
通过更流畅的运动融合和更逼真的纹理,为产品静态图像添加动态效果。营销团队无需安排新的拍摄,即可将目录图片转化为精致的发布短片、功能展示和营销素材。
融合来自不同图像的主体,同时保持更强的场景连贯性,并准确跟踪主体运动。可大规模制作品牌内容和社交媒体营销活动所需的角色互动、群像时刻或生活方式叙事。
利用 Pro 图像转视频变体更流畅的运动融合和更出色的纹理真实感,为静态艺术作品添加动画效果。艺术家和游戏团队可以基于现有视觉素材创作动画概念、角色动作片段或氛围场景研究。
从提示词或源图像开始,生成具有稳定运动或平滑融合运动效果的短视频创意。创作者可以为社交渠道开发多个吸引点、视觉风格和营销方向。
当多个主体需要共享同一场景时,Multi 变体会优先保证场景连贯性,并提供高级运动跟踪。可将其用于角色组合、宠物互动、群体场景或基于图像构建的叙事测试。
对比 Atlas Cloud 上提供的 Kling 1.6 不同版本与其他视频模型,涵盖输入工作流、片段时长、原生音频和标准价格。
| 模型 | 输入工作流 | 输出时长 | 原生音频 | 标准价格 |
|---|---|---|---|---|
| Kling v1.6 Multi i2v Pro | 提示词 + 1 至 4 张图片 | 5 或 10 秒 | - | $0.098/run |
| Kling v1.6 Multi i2v Standard | 提示词 + 1 至 4 张图片 | 5 或 10 秒 | - | $0.056/run |
| Kling v1.6 t2v Standard | 文本提示词 | 5 或 10 秒 | - | $0.056/run |
| Kling v1.6 i2v Pro | 提示词 + 起始帧 + 可选结束帧 | 5 或 10 秒 | - | $0.098/run |
| Kling v1.6 i2v Standard | 提示词 + 起始帧 | 5 或 10 秒 | - | $0.056/run |
| Wan-3.0 Image-to-video | 提示词 + 起始帧 + 可选结束帧 | 2 至 30 秒 | √ | $0.05/second |
| MiniMax H3 Image-to-Video | 提示词 + 起始帧 + 可选结束帧 | 4 至 15 秒 | √ | $0.038/second |
几分钟即可上手 — 按照以下简单步骤,通过 Atlas Cloud 平台集成和部署模型。
在 atlascloud.ai 注册并完成验证。新用户可获得免费额度,用于探索平台和测试模型。
将先进的 Kling 1.6 模型与 Atlas Cloud 的 GPU 加速平台相结合,提供无与伦比的性能、可扩展性和开发体验。
低延迟:
GPU 优化推理,实现实时响应。
统一 API:
一次集成,畅用 Kling 1.6、GPT、Gemini 和 DeepSeek。
透明定价:
按 Token 计费,支持 Serverless 模式。
开发者体验:
SDK、数据分析、微调工具和模板一应俱全。
可靠性:
99.99% 可用性、RBAC 权限控制、合规日志。
安全与合规:
SOC 2 Type II 认证、HIPAA 合规、美国数据主权。
Kling 1.6 是快手开发的视频生成模型系列。在 Atlas Cloud 上,它通过五个 Standard 和 Pro 端点支持文本生视频、单图生视频和多图生视频。其不同变体在运动稳定性、提示词对齐、更顺滑的运动融合、纹理真实感和多主体一致性方面各具优势。
你可以将文本提示词转换为原创视频场景,为单张起始图像添加动画,或组合多张参考图像生成多主体构图。根据所使用的端点,你可以生成时长为 5 或 10 秒的视频片段。
文本驱动生成请选择 T2V Standard;如果要基于单张图像进行高性价比动画生成,请选择 I2V Standard。I2V Pro 更注重顺滑的运动融合和更真实的纹理。对于多个主体或参考图像,请根据成本和一致性要求选择 Multi I2V Standard 或 Multi I2V Pro。
在 Atlas Cloud 控制台中创建 API 密钥,然后携带 Bearer token 向 POST /api/v1/model/generateVideo 发送 JSON 请求。请求中需包含准确的模型标识符、提示词,以及基于图像的端点所要求的图像。保存返回的 prediction ID,并使用它获取异步结果。
每个端点都要求提供模型标识符和最多 2,500 个字符的提示词。文本生视频支持宽高比、时长、引导系数和负面提示词;单图端点则要求提供 JPG、JPEG 或 PNG 格式的起始图像。多图端点接受 1 到 4 张参考图像,并支持设置时长、宽高比和负面提示词。
Atlas Cloud 的五个端点都支持生成 5 或 10 秒的视频,文档中规定的默认时长为 5 秒。文本生视频和多图端点支持 16:9、9:16 或 1:1。单图工作流会根据所提供的源图像确定画面构图,不提供相同的宽高比参数。
按标准目录价计算,T2V Standard、I2V Standard 和 Multi I2V Standard 端点的费用均为每生成一秒视频 $0.056。I2V Pro 和 Multi I2V Pro 的费用为每生成一秒视频 $0.098,采用按量计费。
首先确认 Bearer token、准确的模型标识符、必填提示词以及所有必需的图像字段。重新提交无效请求前,请检查文档规定的提示词长度、图像格式、文件大小、分辨率、时长和参数范围。由于生成过程是异步的,请保留 prediction ID,并在状态为 processing 期间继续查询结果端点。
指南、教程与产品动态,助你充分发挥 Atlas Cloud 的价值。