


HappyHorse 1.1 is Alibaba's video generation family for developers working from text, source frames, or visual references. It turns prompts into clips, animates a chosen first frame with optional guidance, and can use one to nine reference images for precise visual direction. Atlas Cloud brings these workflows into one integration with transparent pay-as-you-go pricing. Start building today.
Happyhorse 1.1 is developed by Alibaba. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
Atlas Cloud provides you with the latest industry-leading creative models.
Choose among text, image, and reference inputs based on how you want to guide each video.
| Modality | Description |
|---|---|
| HappyHorse 1.1 T2V API (Text to Video) | Turn text prompts into videos at 480P, 720P, or 1080P with flexible aspect ratios and durations from 3 to 15 seconds. Use this endpoint for concept visualization, social content, and video production without a source image. |
| HappyHorse 1.1 I2V API (Image to Video) | Starting with a first-frame image, this endpoint creates a 3 to 15 second video with optional prompt guidance and 480P, 720P, or 1080P output. It suits product animation, character motion, and visual storytelling built from an existing composition. |
| HappyHorse 1.1 R2V API (Reference to Video) | Guide video generation with one to nine reference images plus a text prompt, selecting 480P, 720P, or 1080P output, flexible aspect ratios, and a 3 to 15 second duration. Choose it for reference-led scenes, character-focused clips, and projects requiring stronger visual direction. |
HappyHorse 1.1 turns prompts, first frames, or up to nine reference images into videos lasting 3 to 15 seconds, with 480P, 720P, and 1080P output, nine ratios for text and reference modes, seed control, and pay-as-you-go access through Atlas Cloud.
Start with a text prompt of up to 2,500 characters and generate a complete clip without an image input. HappyHorse 1.1 supports 480P, 720P, or 1080P output and lets you set duration from 3 to 15 seconds. Choose from nine aspect ratios for widescreen, vertical, square, portrait, or ultrawide delivery. This workflow fits concept visualization, ads, and short narrative scenes.
Turn one first-frame image into motion, with an optional prompt that guides action, style, or scene details. The image-to-video endpoint accepts JPEG, JPG, PNG, or WEBP files up to 20 MB, with both dimensions at least 300 pixels. Select 480P, 720P, or 1080P and a 3 to 15 second duration. It is suited to product shots, portraits, and storyboard keyframes.
Guide a new video with one to nine reference images plus a required text prompt. Each JPEG, JPG, PNG, or WEBP reference can be up to 20 MB, and its shorter side must be at least 400 pixels. Set resolution, duration, aspect ratio, and seed for the output. This mode gives character, product, and art direction workflows more source material to work from.
Frame each text or reference-driven generation for its destination with nine ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, and 21:9. Resolution spans 480P, 720P, and 1080P, while duration ranges from 3 to 15 seconds. Combine these controls to produce clips for mobile feeds, web players, square placements, portraits, or ultrawide compositions without changing endpoints.
Use the seed parameter alongside prompt, resolution, ratio, and duration settings when you need tighter control over generation inputs. HappyHorse 1.1 accepts seed values from -1 through 2,147,483,647, with -1 requesting a random seed. Keep the other settings explicit as you iterate. This parameterized workflow helps developers organize creative testing and compare variations with a consistent request structure.
Access all three HappyHorse 1.1 workflows through Atlas Cloud with one OpenAI-compatible key and pay-as-you-go billing. Standard pricing starts at $0.07 per second, and billing scales with usage instead of a subscription. Switch among text-to-video, image-to-video, and reference-to-video model paths while keeping the same integration pattern. This setup suits developers building multi-mode video products without separate provider accounts.
See how HappyHorse 1.1 and two comparable video models interpret the same prompts across cinematic action and stylized fantasy storytelling.
An 8–10-second continuous, photorealistic kinetic food-cinema sequence aboard a dining car racing through a violent alpine blizzard. A skilled train chef sways naturally with the carriage’s brutal side-to-side motion while cooking breakfast: he flips a thin golden crêpe high with one hand, slides across the narrow aisle to catch a rolling orange with his free hand, then strikes the descending crêpe with the rim of his pan, launching it in a clean, physically accurate arc through an opening inter-carriage door. Begin with an extreme close-up tracking laterally beside roaring stove flames, sizzling batter, and droplets of hot butter splashing across polished steel; sweep into a fast half-orbit around the chef as he counterbalances every jolt, coat and apron reacting realistically; then rush through the doorway in a precise follow shot locked onto the airborne crêpe as it crosses the flexing connector into the next carriage, narrowly clearing a waiter, and lands perfectly on a surprised passenger’s breakfast plate. At the exact landing beat, a vast avalanche surges past the windows outside as the train escapes it, creating a sharp visual climax. Warm tungsten lamps and orange firelight collide with icy blue dawn through snow-streaked windows; layered orange-gold food tones, silver-gray metal, and cold blue wilderness; deep, narrow central perspective, repeating doorframes within doorframes, strong parallax, subtle handheld vibration, crisp motion clarity, seamless subject continuity and spatial coherence across both carriages, realistic balance, fluid splashes, fabric inertia, collisions, and projectile physics. Synchronized sound design: rail clatter and carriage groans establish the rhythm, pan hiss and butter pops accelerate it, the orange rolls in sync with metallic clicks, the pan strike lands as a bright percussive beat, followed by a soft plate slap and a thunderous avalanche roar; no slow motion, no static filler, no screens, interfaces, dashboards, progress bars, charts, captions, subtitles, logos, or visible text. Premium cinematic realism, dynamic food commercial, immersive one-take illusion, 16:9 aspect ratio.
Generated with HappyHorse-1.1 Text-to-video on Atlas Cloud
Generated with Kling V3.0 Turbo Text-to-Video on Atlas Cloud
Generated with HappyHorse-1.0 Text-to-video on Atlas Cloud
A high-saturation, graphic 3D animated 9-second action sequence at a container terminal at dusk: a coral-red grand piano breaks loose and hurtles down a steel loading ramp, its small caster wheels rattling and bouncing under convincing weight, while a forklift driver in cobalt-blue coveralls races after it, weaving through mustard-yellow cranes and constantly rising and lowering shipping containers that briefly form a shifting emergency corridor. Begin with an extreme low-angle tracking shot inches from the spinning wheel and violently jolting piano leg, sparks and metallic impacts perfectly synchronized with the sound; crane rapidly upward into a top-down view revealing the chase route as a geometric maze of coral-red, cobalt-blue, and mustard-yellow color blocks; then whip-pan and orbit to the front of the speeding forklift as it closes the gap. At the dock edge, the driver thrusts the forks beneath the piano, lifts its heavy body at the final instant, pivots in one forceful rotation, and lands it securely with a deep synchronized clang. The lid snaps open from inertia; on the last beat, a seagull drops onto the keys and strikes one crisp final note. Continuous high-density motion, physically credible momentum, collisions, suspension bounce, shifting loads, and stable character and object identity throughout; hard-edged shadows, bold geometric composition, stylized cinematic 3D animation, sharp dusk rim light, rhythmic industrial percussion built from wheel clatter, hydraulic hisses, container impacts, engine strain, gull wings, and the final piano note. No slow motion, no idle shots, no cuts that break spatial continuity, no screens, software interfaces, dashboards, progress bars, charts, captions, labels, logos, watermarks, or explanatory text. 16:9 aspect ratio.
Generated with HappyHorse-1.1 Text-to-video on Atlas Cloud
Generated with Kling V3.0 Turbo Text-to-Video on Atlas Cloud
Generated with HappyHorse-1.0 Text-to-video on Atlas Cloud
From campaign concepts and product animation to reference-guided brand videos and developer tools, HappyHorse 1.1 turns prompts and images into adaptable short video content.
Turn written concepts into videos at 480P, 720P, or 1080P with flexible aspect ratios. Creative teams can produce campaign concepts, social posts, and story pitches without source footage or filming.
Begin with a first-frame image, then guide its animation with an optional prompt across 3 to 15 seconds. Product marketers can turn still assets into launch clips, display videos, or catalog motion.
Using one to nine reference images, the model shapes videos around supplied visual inputs at up to 1080P. Brand teams can develop campaign variations guided by approved character, product, or environment imagery.
Choose flexible aspect ratios and three resolution tiers to prepare clips for different placements from a single generation workflow. Developers can serve portrait, landscape, and other layouts across content products and campaigns.
Need a quick storyboard beat or a longer product moment? Durations from 3 to 15 seconds let creative tools support rapid concepts, promotional clips, and concise visual sequences within a single interface.
Build video applications around text, first-frame image, and multi-reference inputs through one model family. Users can choose the starting material that best fits ideation, animation, or brand-guided creation for each workflow.
Compare HappyHorse 1.1 generation modes with leading Atlas Cloud video models across supported inputs, clip duration, resolution, reference capacity, and listed base price.
| Model | Supported Inputs | Output Duration | Resolution | Max Reference Images | Listed Base Price |
|---|---|---|---|---|---|
| HappyHorse-1.1 Text-to-video | Text prompt | 3 to 15 seconds | 480P, 720P, 1080P | - | $0.07, unit not listed |
| HappyHorse-1.1 Image-to-video | First-frame image and optional prompt | 3 to 15 seconds | 480P, 720P, 1080P | - | $0.07, unit not listed |
| HappyHorse-1.1 Reference-to-video | Text prompt and reference images | 3 to 15 seconds | 480P, 720P, 1080P | 9 | $0.07, unit not listed |
| Wan-3.0 Reference-to-video | Text, image, video, audio, document, or webpage references | 2 to 30 seconds or smart duration | 480P, 720P, 1080P | 10 | $0.05, unit not listed |
| Seedance 2.5 Reference-to-Video | Text, image, video, and audio references | Up to 30 seconds | - | 30 | $0.167, unit not listed |
| MiniMax H3 Reference-to-Video | Text, image, video, and audio references | 4 to 15 seconds | 768p default, up to 2K | 9 | $0.038/second |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced Happyhorse 1.1 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run Happyhorse 1.1, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
HappyHorse 1.1 is an Alibaba video generation model available on Atlas Cloud for turning text, a first frame, or reference images into short video clips. The family provides separate text-to-video, image-to-video, and reference-to-video endpoints.
You can generate a scene from a text prompt, animate a still image, or guide a new clip with multiple visual references. These workflows cover prompt-only videos, first-frame animations, and reference-guided content.
Create an Atlas Cloud API key and send an authenticated POST request to /api/v1/model/generateVideo with the chosen model ID and required inputs. The response includes a task ID that you can use to check the prediction status until the output is ready.
Select alibaba/happyhorse-1.1/text-to-video for prompt-based generation, alibaba/happyhorse-1.1/image-to-video for first-frame animation, or alibaba/happyhorse-1.1/reference-to-video for multi-image guidance. Each mode has its own input schema.
For text-to-video and reference-to-video, choose 480p, 720p, or 1080p, a duration from 3 to 15 seconds, and one of nine supported aspect ratios. Image-to-video offers the same resolution and duration ranges while using the uploaded image as the first frame.
Reference-to-video accepts one to nine image URLs plus a prompt describing how the images should guide the result. JPEG, JPG, PNG, and WEBP files are supported, with a 20 MB limit per image and a minimum shorter side of 400 pixels.
Atlas Cloud lists a standard price of $0.07 per second for each of the three HappyHorse 1.1 endpoints. Billing is pay-as-you-go, so the generation cost scales with the requested clip duration.
Use the seed parameter on text-to-video or reference-to-video when you want to control generation randomness. Atlas Cloud accepts explicit integer values or -1 when a random seed is preferred.
Each request produces a clip between 3 and 15 seconds. Image-to-video uses one first-frame image, while reference-to-video accepts no more than nine reference images.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.