Put headphones on before you scroll to the clip below. In this category it actually matters.
A knife comes down through a pomegranate carved out of ruby glass. The shell splits, a crystalline crack, then a few dozen glass seeds tumble across wet slate. Here is the part that is different from almost every AI ASMR clip you have scrolled past this year: nobody added that sound. There was no sound library, no editor, no timeline. The picture and the audio came out of one generation, in one call.
For the last year AI ASMR has looked satisfying and sounded slightly wrong, and the reason was never the visuals. It was the last step of the pipeline. Gluing audio onto a silent clip has been a manual job, done by hand, every single time. The category grew so fast around that gap that a small industry of dedicated AI ASMR generator sites appeared purely to automate the stitching, one of them advertising "binaural 3D audio" right on the homepage. The demand was proven long ago. It was just being served by assembly.
So I ran the thing properly. One reproducible workflow, six clips, and then the part nobody in this category ever does: I pulled the left and right channels apart in ffmpeg and measured them. Some of what I expected turned out to be wrong, and that turned out to be the most useful thing in the article.
Key takeaways
- H3 renders the 24 fps picture and a 32 kHz AAC stereo track in the same pass. The audio is generated, not dubbed on afterwards.
- The stereo is genuinely stereo, not dual mono wearing a stereo costume. But you cannot prompt a pan. I asked for a hard left-to-right sweep and got a 1.9 dB difference, which is nothing.
- Stereo width comes from the content, not from your wording. The rain clip, where I asked for no direction at all, came back with a genuinely wide field.
- Locking the first frame keeps the shot across resolutions, but 768P and 2K are still two different takes. The draft previews the look and the sound spec, not the exact choreography.
- A 5 second 768P clip bills $0.50. The same clip at 2K bills $0.70. The whole article cost $4.27.
Listen First: A MiniMax H3 ASMR Video, Cut in One Pass
Five seconds, 2K, 24 fps, 32 kHz stereo, one generation, $0.70. Sound on: the squeak of the edge biting in, the crack, then the seeds. None of it was added afterwards. Generated with minimax/h3/image-to-video.
One prompt. One model. One call. Everything below is how that clip was made, what it cost, and which parts of the marketing story survive contact with a waveform.
Why AI ASMR Blew Up, and Why Most of It Fails the Headphone Test
The trend has a birthday. AI ASMR started spreading in June 2025, right after Veo 3 landed, and the format went nuclear within days. A lava mukbang from @asmraiworks pulled over 11.3 million views in a week; a glass cutting clip did 5.3 million in eleven days (Know Your Meme, 2025). Newsrooms picked it up as a curiosity, with surreal visuals and molten lava eaten with chopsticks doing most of the heavy lifting (The Express Tribune, 2025).
Then people put headphones on, and the mood changed. TechRadar's writer watched a pile of the viral ones and came away describing a sheen of artificiality, "lacking the errors and imprecision that are the hallmark of human-made ASMR" (TechRadar, June 2025). Fans quoted in that piece said the tingle triggers were technically present and it still was not the same.
There is a mechanical reason, and it is not mystical. ASMR is the one content category where the content is the sound. Everywhere else audio is a garnish. Here it is the product. And the dominant workflow was:
- generate a silent clip with a video model,
- hunt a sound library for a crack, a crunch, a rain bed,
- line the sound up to the picture in an editor,
- hand-pan and hand-mix it if you care, which almost nobody does at volume,
- export.
Step 3 is where it dies. A canned crack fired two frames late reads as fake before you can explain why. Multiply that by a daily posting schedule and the errors compound, because the alignment is being redone by a human every single time.
That gap is why an entire cottage industry of AI ASMR generator sites exists. At least three are actively marketing themselves and one leads with binaural 3D audio as a headline feature. They are not solving a fake problem. They are patching a real hole: the video models underneath them do not produce sound.
Which makes one leaderboard split worth a look. On Artificial Analysis's video editing board, the one scored with audio, MiniMax H3 sits at #1 with 1,126 Elo, ahead of Gemini Omni Flash at 1,119 and roughly a hundred points clear of Dreamina Seedance 2.0 at 1,027 (Artificial Analysis, August 2026). On the image-to-video boards, where audio is not part of the score, several of those models sit within a couple of points of each other. The gap opens up when sound counts. For ASMR, sound is all that counts.
The MiniMax H3 ASMR Video Workflow, and What Each Step Costs
Three endpoints, one browser tab. I ran everything in this article on Atlas Cloud, so every figure below came off the bill rather than a rate card.
| Job in this article | Model | Rate |
|---|---|---|
| First frame for the showcase | OpenAI GPT Image 2 (text-to-image) | listed from $0.009/image, billed $0.1745 at high and 2048x1152 |
| Draft and keeper | MiniMax H3 Image-to-Video | $0.10/s at 768P, $0.14/s at 2K |
| Feeding a real recording in | MiniMax H3 Reference-to-Video | same two tiers |
| The three templates | MiniMax H3 Text-to-Video | same two tiers |
| The old way, audio half only | ByteDance Seed Audio 1.0 | $0.143/min |
The listing shows "From $0.1/SEC" on all three H3 endpoints. That is the 768P floor. 2K bills at $0.14/s and that number appears nowhere except your invoice, so plan around the tier you actually request.
Table 1: one pass versus the stitched pipeline.
| MiniMax H3, native audio | Silent model plus post audio | |
|---|---|---|
| Steps to a finished clip | 1 | 4 to 5 |
| Tools involved | 1 | video model + sound library + editor + export |
| How picture and sound stay in sync | same generation, same timeline | you drag it until it looks right |
| Who mixes and pans | nobody, you take what the model gives | you, by hand, per sound |
| Human time per clip | roughly zero after the prompt | 15 to 40 minutes, honestly |
| Compute per 5s clip | $0.50 at 768P | video model + $0.143/min of audio generation |
I am deliberately not quoting a rival video platform's per-second rate, because compute is not where the difference lives. Audio generation runs $0.143 a minute, which is pennies. The real cost of the stitched pipeline is 20 minutes of human attention per clip, and that cost does not fall as you scale. It multiplies. A daily-posting faceless account is exactly where 20 minutes a clip quietly eats the whole business model.
The honest counterweight: hand-stitching buys you control. You can put a sound anywhere in the stereo field and trim it to the frame. H3 gives you the sync for free and, as the measurements below show, does not give you that control back.
Table 2: the three H3 endpoints, the parameters that actually matter.
| image-to-video | text-to-video | reference-to-video | |
|---|---|---|---|
| Main input | image (first frame) | prompt only | refers[] |
| resolution | 768P / 2K, default 2K | same | same |
| duration | any whole second 4 to 15, default 8 | same | same |
| ratio | decided by the first frame | must be set explicitly, adaptive is rejected | decided by the references |
| refers limits | not used together with image | not used | up to 9 images, 3 videos, 3 audio |
| Delivered 16:9 output | 1344x768 at 768P, 2560x1440 at 2K | same | same |
| Audio track | AAC stereo 32 kHz over 24 fps video | same | same |
Two parameter traps worth writing on your hand. The parameter is called ratio, not aspect_ratio, and the wrong key leaves it unset so text-to-video rejects the request. And image plus refers in the same call does not error: it completes, bills in full, and silently throws one of the two inputs away.
MiniMax H3 ASMR Video Tutorial: Cutting a Glass Pomegranate
Glass fruit cutting is the format everybody recognises, so the reader knows instantly what they are looking at. Glass apples and glass mangoes have been done to death, though. A pomegranate is better for one specific reason: when the shell splits, hundreds of glass seeds spill out and roll, so the sound has duration and structure instead of being one centred impact. If a model is faking its audio, a scatter is where it shows.
Step 1: Build the Base Frame With GPT Image 2
Lock the composition before you spend anything on video. Settings: quality: high, size: 2048x1152 (16:9), output_format: png.
Plain1Extreme close-up macro photograph of a whole pomegranate carved entirely from thick 2translucent ruby-red glass, resting on a wet black slate slab. The glass shell is 8mm 3thick with visible internal bubbles and chipped facets; hundreds of tiny glass seeds 4glow inside like garnet crystals. A polished Japanese santoku knife lies at the left 5edge of the frame, blade tip just touching the glass skin. One hard key light from the 6upper right rakes across the slab and throws sharp red caustics onto the wet stone. 7100mm macro, f/2.8, shallow depth of field, dark moody background. 8No hands, no text, no watermark, no logo.

GPT Image 2 playground on Atlas Cloud with quality set to high and size 16:9 2048x1152, the glass pomegranate rendered in the OUTPUT panel
The real run. Quality high at 2048x1152, and the page quotes $0.1745, which is what the invoice later said too.

The generated base frame: a pomegranate carved from thick ruby glass on wet black slate, a santoku knife entering from the left edge
The frame handed to H3. Generated with openai/gpt-image-2/text-to-image.
Step 2: Write the Sound Half of the MiniMax H3 ASMR Video Prompt
Most people write an ASMR prompt as a picture description with the word "ASMR" bolted on the front, then wonder why the model scores it with lo-fi piano. H3 takes audio direction seriously if you give it any. Structure the audio half in five slots:
- Material. What is making the sound. Glass, wax, molten rock, water on glass. Be specific about thickness and wetness, they change timbre.
- Action and its shape in time. Not just "cuts" but "presses down in one slow continuous stroke". The model uses this to place events.
- Mic perspective.
close-mic,intimate, recording distance cues. This sets how dry and near the sound feels, and it works well. - Onomatopoeia sequence. Spell the sound events out in order: squeak, then crack, then dozens of small impacts, then silence. This is the slot that does the most work.
- Negatives.
no music, no voice, no narration, no room reverb, no ambience bed. Without these H3 will helpfully add a score, because most video in its training data has one.
I also wrote channel directions into slot 4, asking for the crack in the left channel and the seeds rolling left to right. Keep reading for what that actually produced.
Step 3: Run the 768P Draft
Draft at 768P. The audio track is the same specification at both tiers, so the entire creative judgement, which in ASMR is a listening judgement, can be made at the cheap rate.
Settings on minimax/h3/image-to-video: image = the Step 1 output, resolution: 768P, duration: 5. The ratio comes from the first frame, so leave it alone.
Plain1The santoku blade enters from the left and presses down through the glass pomegranate 2in one slow continuous stroke, travelling left to right across the frame. The shell 3splits with a bright crystalline crack and a spray of tiny glass seeds tumbles onto the 4wet slate, scattering toward the right edge. Camera locked off, macro, single take, no cuts. 5 6Audio: ASMR, close-mic, intimate, no music, no voice, no narration, no room reverb. 7A dry sustained glass squeak as the edge bites in, then one sharp high crack panned to 8the LEFT channel, followed by dozens of small tinkling seed impacts rolling from the 9LEFT channel across to the RIGHT channel as they scatter. A faint blade-on-stone scrape 10at the very end, then silence. No ambience bed, no hum, no background music.

MiniMax H3 image-to-video playground on Atlas Cloud with the base frame loaded, duration 5, and a completed clip in the OUTPUT panel
The image-to-video run: first frame loaded, duration 5, OUTPUT completed. Note the Resolution select is sitting on the page default of 2K. Switch it to 768P for drafts, which is what the clip below was rendered at.
The 768P draft, $0.50. Softer picture than the hero clip, same audio specification. This is the take the whole decision gets made on.
Step 4: Measure the Stereo Before You Trust It
Do not take my word for the sound, and do not take the model's. Pull the channels apart and look. This is local ffmpeg, no API call, no cost:
Bash1ffmpeg -i 02-glass-pomegranate-768p-draft.mp4 \ 2 -filter_complex "showwavespic=s=1480x240:split_channels=1:colors=#d93a30|#2f7ed8" \ 3 -frames:v 1 stereo-proof.png

Four-panel waveform analysis comparing the glass pomegranate clip and the rain clip, showing left and right channels plus the side channel for each
Left channel over right channel for two clips, plus the side channel (left minus right) at matched gain underneath. This is the whole argument in one picture.
Here is what came back, and part of it is not what I expected.
The stereo is real. The side channel is not empty on any clip I generated. A dual mono file dressed up as stereo would show nothing there. This one has content.
But you cannot prompt a pan. I asked, in capital letters, for the crack in the LEFT channel and the seeds rolling LEFT to RIGHT. Measured: channel correlation 0.97, side channel sitting 17.7 dB below mid, and the largest left-to-right level difference in any quarter-second window is 1.9 dB. That is not a pan. That is a centred image with a little natural width. The knife travels across the frame; the sound stays put.
Width comes from the content, not from your wording. The rain clip in the next section, where I requested no direction at all, came back at correlation 0.32 with the side channel only 2.9 dB below mid. That is a genuinely wide, decorrelated field. Diffuse ambience gets width automatically. Discrete impacts stay near the centre no matter what you write.
So the honest claim for this model is not spatial. It is temporal. You get sound that was generated alongside the picture and lands on the frame it belongs to, for free, forever, without an editor. If your format truly needs a hard pan, you still have to do that yourself in post, and you should stop writing channel names into prompts because they are doing nothing.
Now re-run the keeper: same endpoint, same first frame, same prompt, change one field to resolution: 2K. One more thing worth knowing before you assume the draft is a preview.

Two rows of five frames comparing the 768P and 2K runs at the same timestamps, identical at frame zero and diverging from about 2.5 seconds
Same first frame, same prompt, both tiers. Frame 0 matches exactly. By 2.5s the shell breaks differently and the two clips have become different takes.
Locking the first frame holds the composition, the framing, the lighting and the colour across both tiers, which is why you should always use image-to-video rather than text-to-video for this. It does not give you the same take. The 2K run re-rolls the action. Treat the draft as a preview of the look and the sound, not of the exact choreography.
Step 5: Add a Real Recording as a MiniMax H3 ASMR Video Reference
If you have a sound you actually want, hand it over. reference-to-video takes up to 9 images, 3 videos and 3 audio clips, and it pulls timbre and room character from an audio reference instead of inventing them. I used a public-domain-friendly glass breaking recording by Gravity Sound from Wikimedia Commons (CC BY 4.0), three takes concatenated to 4.5 seconds.
Three rules, and I found the last two the hard way by having requests rejected:
- Audio cannot go alone. The endpoint spells it out: at least one image or video is required, and audio by itself is not allowed. Pair it with a frame.
- The audio reference must be 2 to 15 seconds. My first attempt used a 1.49 second clip and came back with
audio duration 1486 ms, expected [2000, 15000] ms. That range is not on the model page. - The file format is checked. A base64 payload declared as
audio/mpegwas rejected withaudio format ".mpeg" not allowed. A.wavpayload went through. Rejected requests are not billed, but they do cost you a round trip.
Settings: refers: [{url: <base frame>, type: "image"}, {url: <a 2 to 15s recording>, type: "audio"}], resolution: 768P, duration: 5.
Plain1Same locked-off macro shot: the blade slices the glass pomegranate from left to right 2and the seeds scatter. Match the timbre, brightness and room character of the reference 3audio, same recording distance, same dryness. ASMR close-mic, no music, no voice.

MiniMax H3 reference-to-video playground on Atlas Cloud with two reference materials loaded, an audio file in slot 1 and the base frame in slot 2
Reference Materials carrying the audio file and the image together, 2 of 9 used. Drop the image and the request will not run at all.
The reference-guided take, $0.50. Same shot, timbre steered by a real recording rather than invented from the prompt. Audio reference: Glass breaking by Gravity Sound, CC BY 4.0.
Three MiniMax H3 ASMR Video Templates: Cutting, Chewing, Rain
Three formats cover most of what performs in this category. Each is a full paste-and-run prompt with settings and the clip it produced. Deliberately no glass, no red, no slate anywhere below: if every clip on your account looks the same, the algorithm treats it as the same clip.
Table 3: the same five slots, three different jobs.
| Slot | Template 1, The Cut | Template 2, The Chew | Template 3, The Rain |
|---|---|---|---|
| Material | dripping honeycomb, hot steel blade | glowing molten rock, stone bowl | rain on car window glass |
| Action shape | blade sinks front to back, honey pulls down | chopsticks lift, then two bites | no subject action, droplets only |
| Mic perspective | close-mic, very dry, no room | extremely close, intimate | outside the glass, cabin muffled |
| Sound sequence | wax crack, honey stretch, drip on plate | molten sizzle, wet chew, glassy crunch, exhale | steady rain bed, droplet hits, distant tyre hiss |
| Negatives | no music, no voice, no room reverb | no words, no music, no narration | no music, no thunder, no voice, no wipers |
Template 1, The Cut. Honeycomb, amber and gold, 9:16, 768P, 6 seconds, ratio: "9:16".
Plain1Extreme macro, locked-off overhead shot of a thick slab of golden honeycomb on a pale 2ceramic plate, honey already pooling around it. A hot steel blade lowers into the wax 3and sinks through it front to back in one slow continuous press; the cut faces slump 4and a heavy thread of honey stretches and drips onto the plate. Warm amber key light 5from the left, shallow depth of field, single take, no cuts, no hands in frame. 6 7Audio: ASMR, close-mic, very dry, no room reverb, no music, no voice, no narration. 8A soft crackling of wax splitting under the blade, then a thick low stretching pull as 9the honey lengthens, then two heavy wet drips landing on the ceramic plate. Nothing else.
Template 1, $0.60. Wax crack, honey stretch, drip. Generated with minimax/h3/text-to-video.
Template 2, The Chew. Lava mukbang, the format that pulled 11.3 million views in a week, 9:16, 768P, 8 seconds, ratio: "9:16".
Plain1Vertical close-up: a pair of black lacquer chopsticks lifts a glowing clump of molten 2orange lava out of a dark stone bowl and carries it toward the camera, out of frame at 3the bottom. The lava is self-illuminating, casting orange light on everything around it; 4the rest of the room is almost black. Steam curls off the surface. Locked-off camera, 5single take, no cuts. 6 7Audio: ASMR, extremely close-mic, intimate, no words, no music, no narration, no room tone. 8A low molten sizzle and bubbling as the lava is lifted, then a soft wet chew, then a 9brighter glassy crunch on the second bite, then one slow satisfied exhale. No speech.
Template 2, $0.80. Sizzle, chew, crunch, exhale, and not a single word. Generated with minimax/h3/text-to-video.
Template 3, The Rain. Night car window, cold blue and neon, 16:9, 768P, 10 seconds, ratio: "16:9".
This is the only one of the three with no subject action, and it teaches the most about negatives. With nothing happening on screen, H3 reaches for a music bed unless you forbid one explicitly. It is also the clip that came back with by far the widest stereo field, without being asked. Sleep-oriented content wants length, so this runs near the top of the useful duration range.
Plain1Macro shot through the inside of a car window at night, heavy rain running down the 2outside of the glass. Individual droplets hit, hesitate, then streak downward and merge. 3Beyond the glass, out-of-focus neon signage breaks into cold blue and magenta bokeh. 4No wipers, no people, no movement inside the car. Camera locked off, single continuous 5take, shallow depth of field, no cuts. 6 7Audio: ASMR, close-mic on the outside of the glass, cabin slightly muffled. 8A steady even rain bed with clearly separated individual droplet impacts on the glass, 9plus a faint distant hiss of tyres on a wet road. No music, no thunder, no voice, 10no narration, no wiper noise.
Template 3, $1.00. Ten seconds of pure ambience, no soundtrack because the prompt banned one. Generated with minimax/h3/text-to-video.
What a MiniMax H3 ASMR Video Actually Costs
Here is the receipt for this article, not an estimate.
| Line item | Count | Billed |
|---|---|---|
| GPT Image 2 base frame, high, 2048x1152 | 1 | $0.17 |
| 2K keeper, 5s, image-to-video | 1 | $0.70 |
| 768P draft, 5s, image-to-video | 1 | $0.50 |
| 768P reference-to-video, 5s | 1 | $0.50 |
| Template clips at 768P, 6s + 8s + 10s | 3 | $2.40 |
| Rejected reference requests | 2 | $0.00 |
| Total | $4.27 |
Six publishable clips and one base frame. Scale it to a schedule: one 8 second vertical clip a day at 768P is $0.80, so roughly $24 a month for a daily-posting faceless account, before you spend a minute in an editor.
Two billing notes. GPT Image 2's listed $0.009 is a token-tier floor; a high-quality 2048x1152 render lands at $0.1745, nearly twenty times that, so budget from the tier you actually use. And H3's price field backfills on a delay: poll until status is completed and it is frequently still undefined. Poll the prediction id again a moment later for the real figure. That second poll is the only way to build a receipt like this one instead of a guess.
One Honest Note on ASMR Audio and Rights
Do not upload somebody else's ASMR recording as a refers audio file. The reference genuinely migrates timbre into the output, and running a recording through a model does not change who owns it. The reference used here is CC BY 4.0 and credited above, which took about two minutes to source.
Do not build ASMR around a recognisable real person's voice. Every template in this article is deliberately voiceless, which is both the genre convention and the least complicated path.
Calling H3 through an API is not affected by the community licence attached to the open weights. That licence, which covers self-hosting, currently excludes the EU, UK, Korea and the US, and a formal licensing channel is open for those territories. Worth knowing, not worth worrying about if you are hitting the endpoint.
MiniMax H3 ASMR Video: Frequently Asked Questions
Does a MiniMax H3 ASMR video generate its own sound, or do I add it afterwards?
The model generates it. Picture and audio come out of the same pass: 24 fps video with an AAC stereo track at 32 kHz in the delivered file. There is no separate audio step, no sound library and no alignment work, which is the entire reason this endpoint is interesting for ASMR specifically.
Can I make a MiniMax H3 ASMR video pan from left to right?
Not by asking. I wrote explicit channel directions into the prompt and measured the result: correlation 0.97 between channels and a maximum level difference of 1.9 dB in any quarter-second window, which is no pan at all. The track is genuinely stereo and diffuse content like rain comes back with a wide field, but discrete impacts stay centred regardless of wording. If your format needs a hard pan, do it in post.
Can I upload my own recording as a MiniMax H3 ASMR video reference?
Yes, through reference-to-video, which accepts up to 3 audio references alongside up to 9 images and 3 videos. Audio cannot be submitted on its own, so pair it with at least one image or video. The audio also has to be between 2 and 15 seconds, and the format is validated: a .wav payload worked where an audio/mpeg one was rejected. Rejected requests are not billed.
Is a 768P MiniMax H3 ASMR video's audio worse than 2K?
No. Both tiers deliver the same AAC stereo 32 kHz specification. Only the picture changes, and it changes a lot: 16:9 comes back as 1344x768 at 768P versus 2560x1440 at 2K. Since the creative judgement in ASMR is a listening judgement, draft at $0.10/s and re-run keepers at $0.14/s. Just remember the 2K run is a new take, not an upscale of the draft.
How long should a MiniMax H3 ASMR video be, and what aspect ratio?
Duration accepts any whole second from 4 to 15. Cutting and chewing clips work at 5 to 8 seconds because the payoff is one event; sleep-oriented ambience wants 10 or more. For Shorts, Reels and TikTok use 9:16, and remember that text-to-video requires ratio to be set explicitly and rejects adaptive. On image-to-video the first frame decides the shape for you.
How much does one MiniMax H3 ASMR video cost?
A 5 second clip bills $0.50 at 768P and $0.70 at 2K. An 8 second vertical clip at 768P is $0.80. Posting one of those a day is about $24 a month in compute, which is the number worth comparing against the 20 minutes an editor would take per clip on the stitched workflow.
Why does my MiniMax H3 ASMR video come out with background music I never asked for?
Because you did not forbid it. Most video in any model's training data has a music bed, so the default behaviour is to add one, especially when nothing is happening on screen. Put the negatives in the audio half of the prompt explicitly: no music, no voice, no narration, no room reverb, no ambience bed. Ambient templates need this more than action ones do.






