Seaside town, text only
Text to video · 10s · 1080pA pastel seaside town at morning, fishing boats returning to the harbor, gulls circling, a baker opening his shop shutters. Slow pan. Waves, gulls, shutter rattle.
Use this promptThe multi-reference engine. Combine up to nine images (a person, an outfit, a product, a place) into one coherent shot with believable physics and sound, at up to 1080p.



19 credits
for 5s at 720p (≈ $3.61)
110s
across 118 finished clips, last 60 days
79/100
#3 of 9 engines, blind benchmark
Live figures from the GoCrazy Video Index, refreshed hourly. Dollar amounts use the best plan rate.
Every clip below was generated on GoCrazyAI with the exact prompt and settings shown. Hover or tap to play.
A pastel seaside town at morning, fishing boats returning to the harbor, gulls circling, a baker opening his shop shutters. Slow pan. Waves, gulls, shutter rattle.
Use this promptHundreds of paper lanterns rise into a night sky over a river, reflections shimmering. Slow tilt up. Soft crowd murmur and water.
Use this promptA red train winds through snowy mountains, steam trailing, the camera flying alongside. Train rhythm and wind.
Use this promptHappyHorse 1.1 is a video model served on Alibaba’s cloud, and it is the engine in the studio built for composition. Most engines animate one picture; HappyHorse takes up to nine reference images and assembles them into a single scene, keeping each element recognisable.
That makes it the right tool for shots with several ingredients: a character in a specific outfit holding a specific product in a specific place. Give each ingredient its own reference image, say in the prompt how they come together, and it handles the rest.
It also works from a single photo or from text alone, renders in 720p or 1080p for 5, 10 or 15 seconds, and generates sound with the picture.
Credits, from the live price list. You always see the exact cost before you generate.
| Resolution | 5s | 10s | 15s |
|---|---|---|---|
| 720p | 19≈ $3.61 | 30≈ $5.70 | 45≈ $8.55 |
| 1080p | 29≈ $5.51 | 40≈ $7.60 | 55≈ $10.45 |
Dollar figures use the best plan rate. See plans and credit packs.
Up to nine reference images combined into a single coherent scene. The one engine designed for assembling a look from parts.
Faces, clothing and product details from the references carry through the motion instead of drifting.
Fabric, hair, liquids and objects move with weight, which keeps product and fashion shots convincing.
Every clip comes with audio that fits the scene, so it is ready to post without an extra step.
No references to hand? Describe the scene and pick 16:9, 9:16 or 1:1.
Render in full HD for delivery, or 720p while you are still exploring.
HappyHorse 1.1 is preselected when you arrive from this page. Add a start image, or leave it empty to generate from text, or attach references.
Describe the subject, the action, the camera and the light in one or two sentences. Start from a tested prompt below.
Choose duration, resolution and aspect ratio. The credit cost updates as you change them, and failed renders are refunded.
Six prompts written for how HappyHorse 1.1 reads direction. The full guide has 30 more, sorted by use case.
The woman from image 1 wears the jacket from image 2 and walks through the street market from image 3, glancing at the stalls. Walking-pace tracking shot. Market chatter and footsteps.
Upload one reference per ingredient, in the order you name them.
The man from image 1 picks up the bottle from image 2 on the kitchen counter from image 3, turns it to read the label and smiles. Medium shot, warm morning light. A soft clink.
The woman from image 1 and the man from image 2 meet on the bridge from image 3, shake hands and start walking together. Wide shot, golden hour. River and distant traffic.
The dress from the reference sways as the woman turns slowly on a rooftop, fabric catching the wind. Low angle, sunset light. Wind and city hum.
A pastel seaside town at morning, fishing boats returning to the harbor, gulls circling, a baker opening his shop shutters. Slow pan. Waves, gulls, shutter rattle.
No image needed. Pick 16:9, 9:16 or 1:1.
The dog from image 1 runs across the park from image 2 and jumps up to greet the woman from image 3, who laughs and kneels down. Low tracking shot. Barking, laughter, leaves.
Same studio, same credits. Switch engines per clip; here is when another one fits better.
| Engine | Longest clip | Audio | Credits / 5s | Median render |
|---|---|---|---|---|
| Seedance 2.5 | 30s | Native | 30 | 385s |
| Seedance 2.5 Turbo | 30s | Native | 16 | — |
| Wan 3.0 Prime | 30s | Native | 10 | 88s |
| MiniMax H3 | 15s | Native | 4 | 108s |
| Seedance 2.0 Mini | 15s | Native | 10 | 227s |
| Seedance 2.0 Fast | 15s | Native | 15 | 213s |
| Seedance 2.0 | 15s | Native | 20 | 242s |
| Wan 2.7 | 15s | Native | 17 | 114s |
| HappyHorse 1.1 | 15s | Native | 19 | 110s |
Credits for a 5-second clip at 720p (or the engine's nearest tier). Full method on the Video Index.
HappyHorse 1.1 is an AI video model served on Alibaba’s cloud that builds a single scene from up to nine reference images, with sound. It also generates from one photo or from text alone.
Up to nine. The first image is the opening frame; the others guide the characters, clothing, products and setting.
720p or 1080p, for 5, 10 or 15 seconds. Extend a clip afterwards to make it longer.
Text-to-video supports 16:9, 9:16 and 1:1. Image-based clips follow the start image.
It is billed in credits by length and resolution. The table on this page comes from the live price list, and failed renders are refunded automatically.
Yes, audio is generated with every clip.
Yes, of yourself or of adults who have agreed to appear. Uploads are checked automatically. GoCrazyAI is for adults (18+) only.
Opens AI Video Pro with HappyHorse 1.1 selected. Paste a prompt from this page and adjust the length before you generate.
Last updated: