Hero background 1Hero background 2Hero background 3

HappyHorse 1.1 Multi Image Video Generation

HappyHorse 1.1 is Alibaba's video generation family for developers working from text, source frames, or visual references. It turns prompts into clips, animates a chosen first frame with optional guidance, and can use one to nine reference images for precise visual direction. Atlas Cloud brings these workflows into one integration with transparent pay-as-you-go pricing. Start building today.

Happyhorse 1.1 is developed by Alibaba. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.

Explore the Leading Happyhorse 1.1

Atlas Cloud provides you with the latest industry-leading creative models.

Compare HappyHorse 1.1 Video Generation Modes

Choose among text, image, and reference inputs based on how you want to guide each video.

ModalityDescription
HappyHorse 1.1 T2V API (Text to Video)Turn text prompts into videos at 480P, 720P, or 1080P with flexible aspect ratios and durations from 3 to 15 seconds. Use this endpoint for concept visualization, social content, and video production without a source image.
HappyHorse 1.1 I2V API (Image to Video)Starting with a first-frame image, this endpoint creates a 3 to 15 second video with optional prompt guidance and 480P, 720P, or 1080P output. It suits product animation, character motion, and visual storytelling built from an existing composition.
HappyHorse 1.1 R2V API (Reference to Video)Guide video generation with one to nine reference images plus a text prompt, selecting 480P, 720P, or 1080P output, flexible aspect ratios, and a 3 to 15 second duration. Choose it for reference-led scenes, character-focused clips, and projects requiring stronger visual direction.

Three Inputs, One HappyHorse 1.1 Toolkit

HappyHorse 1.1 turns prompts, first frames, or up to nine reference images into videos lasting 3 to 15 seconds, with 480P, 720P, and 1080P output, nine ratios for text and reference modes, seed control, and pay-as-you-go access through Atlas Cloud.

HappyHorse 1.1 Text-to-Video

Start with a text prompt of up to 2,500 characters and generate a complete clip without an image input. HappyHorse 1.1 supports 480P, 720P, or 1080P output and lets you set duration from 3 to 15 seconds. Choose from nine aspect ratios for widescreen, vertical, square, portrait, or ultrawide delivery. This workflow fits concept visualization, ads, and short narrative scenes.

First-Frame Animation

Turn one first-frame image into motion, with an optional prompt that guides action, style, or scene details. The image-to-video endpoint accepts JPEG, JPG, PNG, or WEBP files up to 20 MB, with both dimensions at least 300 pixels. Select 480P, 720P, or 1080P and a 3 to 15 second duration. It is suited to product shots, portraits, and storyboard keyframes.

HappyHorse 1.1 Reference Control

Guide a new video with one to nine reference images plus a required text prompt. Each JPEG, JPG, PNG, or WEBP reference can be up to 20 MB, and its shorter side must be at least 400 pixels. Set resolution, duration, aspect ratio, and seed for the output. This mode gives character, product, and art direction workflows more source material to work from.

Formats for Every Destination

Frame each text or reference-driven generation for its destination with nine ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, and 21:9. Resolution spans 480P, 720P, and 1080P, while duration ranges from 3 to 15 seconds. Combine these controls to produce clips for mobile feeds, web players, square placements, portraits, or ultrawide compositions without changing endpoints.

HappyHorse 1.1 Seed Control

Use the seed parameter alongside prompt, resolution, ratio, and duration settings when you need tighter control over generation inputs. HappyHorse 1.1 accepts seed values from -1 through 2,147,483,647, with -1 requesting a random seed. Keep the other settings explicit as you iterate. This parameterized workflow helps developers organize creative testing and compare variations with a consistent request structure.

One Key, Three Workflows

Access all three HappyHorse 1.1 workflows through Atlas Cloud with one OpenAI-compatible key and pay-as-you-go billing. Standard pricing starts at $0.07 per second, and billing scales with usage instead of a subscription. Switch among text-to-video, image-to-video, and reference-to-video model paths while keeping the same integration pattern. This setup suits developers building multi-mode video products without separate provider accounts.

HappyHorse 1.1 in a Three Model Prompt Showdown

See how HappyHorse 1.1 and two comparable video models interpret the same prompts across cinematic action and stylized fantasy storytelling.

Prompt

An 8–10-second continuous, photorealistic kinetic food-cinema sequence aboard a dining car racing through a violent alpine blizzard. A skilled train chef sways naturally with the carriage’s brutal side-to-side motion while cooking breakfast: he flips a thin golden crêpe high with one hand, slides across the narrow aisle to catch a rolling orange with his free hand, then strikes the descending crêpe with the rim of his pan, launching it in a clean, physically accurate arc through an opening inter-carriage door. Begin with an extreme close-up tracking laterally beside roaring stove flames, sizzling batter, and droplets of hot butter splashing across polished steel; sweep into a fast half-orbit around the chef as he counterbalances every jolt, coat and apron reacting realistically; then rush through the doorway in a precise follow shot locked onto the airborne crêpe as it crosses the flexing connector into the next carriage, narrowly clearing a waiter, and lands perfectly on a surprised passenger’s breakfast plate. At the exact landing beat, a vast avalanche surges past the windows outside as the train escapes it, creating a sharp visual climax. Warm tungsten lamps and orange firelight collide with icy blue dawn through snow-streaked windows; layered orange-gold food tones, silver-gray metal, and cold blue wilderness; deep, narrow central perspective, repeating doorframes within doorframes, strong parallax, subtle handheld vibration, crisp motion clarity, seamless subject continuity and spatial coherence across both carriages, realistic balance, fluid splashes, fabric inertia, collisions, and projectile physics. Synchronized sound design: rail clatter and carriage groans establish the rhythm, pan hiss and butter pops accelerate it, the orange rolls in sync with metallic clicks, the pan strike lands as a bright percussive beat, followed by a soft plate slap and a thunderous avalanche roar; no slow motion, no static filler, no screens, interfaces, dashboards, progress bars, charts, captions, subtitles, logos, or visible text. Premium cinematic realism, dynamic food commercial, immersive one-take illusion, 16:9 aspect ratio.

Generated with HappyHorse-1.1 Text-to-video on Atlas Cloud

Generated with Kling V3.0 Turbo Text-to-Video on Atlas Cloud

Generated with HappyHorse-1.0 Text-to-video on Atlas Cloud

Prompt

A high-saturation, graphic 3D animated 9-second action sequence at a container terminal at dusk: a coral-red grand piano breaks loose and hurtles down a steel loading ramp, its small caster wheels rattling and bouncing under convincing weight, while a forklift driver in cobalt-blue coveralls races after it, weaving through mustard-yellow cranes and constantly rising and lowering shipping containers that briefly form a shifting emergency corridor. Begin with an extreme low-angle tracking shot inches from the spinning wheel and violently jolting piano leg, sparks and metallic impacts perfectly synchronized with the sound; crane rapidly upward into a top-down view revealing the chase route as a geometric maze of coral-red, cobalt-blue, and mustard-yellow color blocks; then whip-pan and orbit to the front of the speeding forklift as it closes the gap. At the dock edge, the driver thrusts the forks beneath the piano, lifts its heavy body at the final instant, pivots in one forceful rotation, and lands it securely with a deep synchronized clang. The lid snaps open from inertia; on the last beat, a seagull drops onto the keys and strikes one crisp final note. Continuous high-density motion, physically credible momentum, collisions, suspension bounce, shifting loads, and stable character and object identity throughout; hard-edged shadows, bold geometric composition, stylized cinematic 3D animation, sharp dusk rim light, rhythmic industrial percussion built from wheel clatter, hydraulic hisses, container impacts, engine strain, gull wings, and the final piano note. No slow motion, no idle shots, no cuts that break spatial continuity, no screens, software interfaces, dashboards, progress bars, charts, captions, labels, logos, watermarks, or explanatory text. 16:9 aspect ratio.

Generated with HappyHorse-1.1 Text-to-video on Atlas Cloud

Generated with Kling V3.0 Turbo Text-to-Video on Atlas Cloud

Generated with HappyHorse-1.0 Text-to-video on Atlas Cloud

HappyHorse 1.1 Across Creative Workflows

From campaign concepts and product animation to reference-guided brand videos and developer tools, HappyHorse 1.1 turns prompts and images into adaptable short video content.

HappyHorse 1.1 Campaign Concepts

Turn written concepts into videos at 480P, 720P, or 1080P with flexible aspect ratios. Creative teams can produce campaign concepts, social posts, and story pitches without source footage or filming.

Still Images in Motion

Begin with a first-frame image, then guide its animation with an optional prompt across 3 to 15 seconds. Product marketers can turn still assets into launch clips, display videos, or catalog motion.

Reference-Guided Brand Variations

Using one to nine reference images, the model shapes videos around supplied visual inputs at up to 1080P. Brand teams can develop campaign variations guided by approved character, product, or environment imagery.

Multi-Format Content Pipelines

Choose flexible aspect ratios and three resolution tiers to prepare clips for different placements from a single generation workflow. Developers can serve portrait, landscape, and other layouts across content products and campaigns.

HappyHorse 1.1 Clip Planning

Need a quick storyboard beat or a longer product moment? Durations from 3 to 15 seconds let creative tools support rapid concepts, promotional clips, and concise visual sequences within a single interface.

Flexible Video App Inputs

Build video applications around text, first-frame image, and multi-reference inputs through one model family. Users can choose the starting material that best fits ideation, animation, or brand-guided creation for each workflow.

How HappyHorse 1.1 Stacks Up for Video Generation

Compare HappyHorse 1.1 generation modes with leading Atlas Cloud video models across supported inputs, clip duration, resolution, reference capacity, and listed base price.

ModelSupported InputsOutput DurationResolutionMax Reference ImagesListed Base Price
HappyHorse-1.1 Text-to-videoText prompt3 to 15 seconds480P, 720P, 1080P-$0.07, unit not listed
HappyHorse-1.1 Image-to-videoFirst-frame image and optional prompt3 to 15 seconds480P, 720P, 1080P-$0.07, unit not listed
HappyHorse-1.1 Reference-to-videoText prompt and reference images3 to 15 seconds480P, 720P, 1080P9$0.07, unit not listed
Wan-3.0 Reference-to-videoText, image, video, audio, document, or webpage references2 to 30 seconds or smart duration480P, 720P, 1080P10$0.05, unit not listed
Seedance 2.5 Reference-to-VideoText, image, video, and audio referencesUp to 30 seconds-30$0.167, unit not listed
MiniMax H3 Reference-to-VideoText, image, video, and audio references4 to 15 seconds768p default, up to 2K9$0.038/second

How to Use Happyhorse 1.1 on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use Happyhorse 1.1 on Atlas Cloud

Combining the advanced Happyhorse 1.1 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run Happyhorse 1.1, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

Building with HappyHorse 1.1: Common Questions

HappyHorse 1.1 is an Alibaba video generation model available on Atlas Cloud for turning text, a first frame, or reference images into short video clips. The family provides separate text-to-video, image-to-video, and reference-to-video endpoints.

You can generate a scene from a text prompt, animate a still image, or guide a new clip with multiple visual references. These workflows cover prompt-only videos, first-frame animations, and reference-guided content.

Create an Atlas Cloud API key and send an authenticated POST request to /api/v1/model/generateVideo with the chosen model ID and required inputs. The response includes a task ID that you can use to check the prediction status until the output is ready.

Select alibaba/happyhorse-1.1/text-to-video for prompt-based generation, alibaba/happyhorse-1.1/image-to-video for first-frame animation, or alibaba/happyhorse-1.1/reference-to-video for multi-image guidance. Each mode has its own input schema.

For text-to-video and reference-to-video, choose 480p, 720p, or 1080p, a duration from 3 to 15 seconds, and one of nine supported aspect ratios. Image-to-video offers the same resolution and duration ranges while using the uploaded image as the first frame.

Reference-to-video accepts one to nine image URLs plus a prompt describing how the images should guide the result. JPEG, JPG, PNG, and WEBP files are supported, with a 20 MB limit per image and a minimum shorter side of 400 pixels.

Atlas Cloud lists a standard price of $0.07 per second for each of the three HappyHorse 1.1 endpoints. Billing is pay-as-you-go, so the generation cost scales with the requested clip duration.

Use the seed parameter on text-to-video or reference-to-video when you want to control generation randomness. Atlas Cloud accepts explicit integer values or -1 when a random seed is preferred.

Each request produces a clip between 3 and 15 seconds. Image-to-video uses one first-frame image, while reference-to-video accepts no more than nine reference images.

Explore More Families

Seedance 2.5

Seedance 2.5 API is now available on Atlas Cloud! It gives developers ByteDance's newest video model. It generates up to 30 seconds of native video in a single pass from text, a single image, or as many as 50 multimodal references, with synchronized audio and in-frame multilingual text. On Atlas Cloud you reach it through one key, with subject consistency and improved physics keeping long shots coherent. (Update: Seedance 2.5 1080P API Is Available NOW!)

View Family

Wan 3.0

Wan 3.0 API is the next generation of Alibaba's Wan video family, built to push long-form generation, multi-reference control, and audiovisual quality to new heights. Atlas Cloud already hosts Wan 2.7, 2.6, and 2.5, and Wan 3.0 runs on the same unified key with no separate setup. Start building today. Scroll down to the showcase to see what Wan 3.0 can create.

View Family

MiniMax H3

MiniMax H3 is MiniMax's multimodal video family for text, image, and reference guided creation. Across supported routes, it preserves subjects from reference media, offers flexible aspect ratios, and pairs generated sound with visuals through H3 Developer, with output profiles selected by endpoint. Atlas Cloud unifies the family behind one OpenAI-compatible key with transparent pay-as-you-go pricing from the standard rate of $0.038 per second. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

Seedance 2.0 is ByteDance’s production video model for precise shot creation. Turn prompts into video, animate a first-frame image with optional last-frame guidance, or shape results with reference media and optional web search. Atlas Cloud brings these workflows into one unified API with transparent pay-as-you-go pricing and one OpenAI-compatible key. Start building today.

View Family

GPT Image 2.5

The gpt-image-2.5 family from OpenAI gives developers a choice of Flare and Sunburst for production image workflows. Render at arbitrary resolutions up to 3840x2160 and select from five quality tiers, including xhigh and max, to match specific output requirements. Atlas Cloud provides ready-to-use REST inference with no cold starts and standard pricing from $0.004 per generation. Start building today.

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

Gemini Omni is Google DeepMind's natively multimodal video family for prompt guided creation and editing. Generate from text, animate still images, or revise existing footage with optional visual references while preserving untouched content. Flexible resolution, aspect ratio, and duration controls support varied production workflows. Atlas Cloud unifies access with transparent pay-as-you-go pricing. Start building today.

View Family

Grok Imagine

Grok Imagine Image is xAI's family for generating polished visuals and revising one or more reference images through natural language instructions. Its standard and quality endpoints cover text to image creation, single image changes, and indexed multi-image composition. Atlas Cloud brings these workflows into one API, with standard generation and editing priced at $0.02 per image. Start building today.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

One API for All Media AI.

Explore all models