Seedance 2.5 Now Live — First on Atlas Cloud

MiniMax H3 ASMR Video: I Cut a Glass Pomegranate, Then Measured Whether the Stereo Was Real

Most AI ASMR glues stock sound onto silent video. A MiniMax H3 ASMR video renders 32 kHz stereo in one pass. I measured the channels, with prompts and costs.

Put headphones on before you scroll to the clip below. In this category it actually matters.

A knife comes down through a pomegranate carved out of ruby glass. The shell splits, a crystalline crack, then a few dozen glass seeds tumble across wet slate. Here is the part that is different from almost every AI ASMR clip you have scrolled past this year: nobody added that sound. There was no sound library, no editor, no timeline. The picture and the audio came out of one generation, in one call.

For the last year AI ASMR has looked satisfying and sounded slightly wrong, and the reason was never the visuals. It was the last step of the pipeline. Gluing audio onto a silent clip has been a manual job, done by hand, every single time. The category grew so fast around that gap that a small industry of dedicated AI ASMR generator sites appeared purely to automate the stitching, one of them advertising "binaural 3D audio" right on the homepage. The demand was proven long ago. It was just being served by assembly.

So I ran the thing properly. One reproducible workflow, six clips, and then the part nobody in this category ever does: I pulled the left and right channels apart in ffmpeg and measured them. Some of what I expected turned out to be wrong, and that turned out to be the most useful thing in the article.

Key takeaways

  • H3 renders the 24 fps picture and a 32 kHz AAC stereo track in the same pass. The audio is generated, not dubbed on afterwards.
  • The stereo is genuinely stereo, not dual mono wearing a stereo costume. But you cannot prompt a pan. I asked for a hard left-to-right sweep and got a 1.9 dB difference, which is nothing.
  • Stereo width comes from the content, not from your wording. The rain clip, where I asked for no direction at all, came back with a genuinely wide field.
  • Locking the first frame keeps the shot across resolutions, but 768P and 2K are still two different takes. The draft previews the look and the sound spec, not the exact choreography.
  • A 5 second 768P clip bills $0.50. The same clip at 2K bills $0.70. The whole article cost $4.27.

Listen First: A MiniMax H3 ASMR Video, Cut in One Pass

Five seconds, 2K, 24 fps, 32 kHz stereo, one generation, $0.70. Sound on: the squeak of the edge biting in, the crack, then the seeds. None of it was added afterwards. Generated with minimax/h3/image-to-video.

One prompt. One model. One call. Everything below is how that clip was made, what it cost, and which parts of the marketing story survive contact with a waveform.

Why AI ASMR Blew Up, and Why Most of It Fails the Headphone Test

The trend has a birthday. AI ASMR started spreading in June 2025, right after Veo 3 landed, and the format went nuclear within days. A lava mukbang from @asmraiworks pulled over 11.3 million views in a week; a glass cutting clip did 5.3 million in eleven days (Know Your Meme, 2025). Newsrooms picked it up as a curiosity, with surreal visuals and molten lava eaten with chopsticks doing most of the heavy lifting (The Express Tribune, 2025).

Then people put headphones on, and the mood changed. TechRadar's writer watched a pile of the viral ones and came away describing a sheen of artificiality, "lacking the errors and imprecision that are the hallmark of human-made ASMR" (TechRadar, June 2025). Fans quoted in that piece said the tingle triggers were technically present and it still was not the same.

There is a mechanical reason, and it is not mystical. ASMR is the one content category where the content is the sound. Everywhere else audio is a garnish. Here it is the product. And the dominant workflow was:

  1. generate a silent clip with a video model,
  2. hunt a sound library for a crack, a crunch, a rain bed,
  3. line the sound up to the picture in an editor,
  4. hand-pan and hand-mix it if you care, which almost nobody does at volume,
  5. export.

Step 3 is where it dies. A canned crack fired two frames late reads as fake before you can explain why. Multiply that by a daily posting schedule and the errors compound, because the alignment is being redone by a human every single time.

That gap is why an entire cottage industry of AI ASMR generator sites exists. At least three are actively marketing themselves and one leads with binaural 3D audio as a headline feature. They are not solving a fake problem. They are patching a real hole: the video models underneath them do not produce sound.

Which makes one leaderboard split worth a look. On Artificial Analysis's video editing board, the one scored with audio, MiniMax H3 sits at #1 with 1,126 Elo, ahead of Gemini Omni Flash at 1,119 and roughly a hundred points clear of Dreamina Seedance 2.0 at 1,027 (Artificial Analysis, August 2026). On the image-to-video boards, where audio is not part of the score, several of those models sit within a couple of points of each other. The gap opens up when sound counts. For ASMR, sound is all that counts.

The MiniMax H3 ASMR Video Workflow, and What Each Step Costs

Three endpoints, one browser tab. I ran everything in this article on Atlas Cloud, so every figure below came off the bill rather than a rate card.

Job in this articleModelRate
First frame for the showcaseOpenAI GPT Image 2 (text-to-image)listed from $0.009/image, billed $0.1745 at high and 2048x1152
Draft and keeperMiniMax H3 Image-to-Video$0.10/s at 768P, $0.14/s at 2K
Feeding a real recording inMiniMax H3 Reference-to-Videosame two tiers
The three templatesMiniMax H3 Text-to-Videosame two tiers
The old way, audio half onlyByteDance Seed Audio 1.0$0.143/min

The listing shows "From $0.1/SEC" on all three H3 endpoints. That is the 768P floor. 2K bills at $0.14/s and that number appears nowhere except your invoice, so plan around the tier you actually request.

Table 1: one pass versus the stitched pipeline.

MiniMax H3, native audioSilent model plus post audio
Steps to a finished clip14 to 5
Tools involved1video model + sound library + editor + export
How picture and sound stay in syncsame generation, same timelineyou drag it until it looks right
Who mixes and pansnobody, you take what the model givesyou, by hand, per sound
Human time per cliproughly zero after the prompt15 to 40 minutes, honestly
Compute per 5s clip$0.50 at 768Pvideo model + $0.143/min of audio generation

I am deliberately not quoting a rival video platform's per-second rate, because compute is not where the difference lives. Audio generation runs $0.143 a minute, which is pennies. The real cost of the stitched pipeline is 20 minutes of human attention per clip, and that cost does not fall as you scale. It multiplies. A daily-posting faceless account is exactly where 20 minutes a clip quietly eats the whole business model.

The honest counterweight: hand-stitching buys you control. You can put a sound anywhere in the stereo field and trim it to the frame. H3 gives you the sync for free and, as the measurements below show, does not give you that control back.

Table 2: the three H3 endpoints, the parameters that actually matter.

image-to-videotext-to-videoreference-to-video
Main inputimage (first frame)prompt onlyrefers[]
resolution768P / 2K, default 2Ksamesame
durationany whole second 4 to 15, default 8samesame
ratiodecided by the first framemust be set explicitly, adaptive is rejecteddecided by the references
refers limitsnot used together with imagenot usedup to 9 images, 3 videos, 3 audio
Delivered 16:9 output1344x768 at 768P, 2560x1440 at 2Ksamesame
Audio trackAAC stereo 32 kHz over 24 fps videosamesame

Two parameter traps worth writing on your hand. The parameter is called ratio, not aspect_ratio, and the wrong key leaves it unset so text-to-video rejects the request. And image plus refers in the same call does not error: it completes, bills in full, and silently throws one of the two inputs away.

MiniMax H3 ASMR Video Tutorial: Cutting a Glass Pomegranate

Glass fruit cutting is the format everybody recognises, so the reader knows instantly what they are looking at. Glass apples and glass mangoes have been done to death, though. A pomegranate is better for one specific reason: when the shell splits, hundreds of glass seeds spill out and roll, so the sound has duration and structure instead of being one centred impact. If a model is faking its audio, a scatter is where it shows.

Step 1: Build the Base Frame With GPT Image 2

Lock the composition before you spend anything on video. Settings: quality: high, size: 2048x1152 (16:9), output_format: png.

Plain
1Extreme close-up macro photograph of a whole pomegranate carved entirely from thick
2translucent ruby-red glass, resting on a wet black slate slab. The glass shell is 8mm
3thick with visible internal bubbles and chipped facets; hundreds of tiny glass seeds
4glow inside like garnet crystals. A polished Japanese santoku knife lies at the left
5edge of the frame, blade tip just touching the glass skin. One hard key light from the
6upper right rakes across the slab and throws sharp red caustics onto the wet stone.
7100mm macro, f/2.8, shallow depth of field, dark moody background.
8No hands, no text, no watermark, no logo.

AI image generator interface showing a text prompt and generated image

GPT Image 2 playground on Atlas Cloud with quality set to high and size 16:9 2048x1152, the glass pomegranate rendered in the OUTPUT panel

The real run. Quality high at 2048x1152, and the page quotes $0.1745, which is what the invoice later said too.

Knife blade cutting a translucent red glass pomegranate on dark stone

The generated base frame: a pomegranate carved from thick ruby glass on wet black slate, a santoku knife entering from the left edge

The frame handed to H3. Generated with openai/gpt-image-2/text-to-image.

Step 2: Write the Sound Half of the MiniMax H3 ASMR Video Prompt

Most people write an ASMR prompt as a picture description with the word "ASMR" bolted on the front, then wonder why the model scores it with lo-fi piano. H3 takes audio direction seriously if you give it any. Structure the audio half in five slots:

  1. Material. What is making the sound. Glass, wax, molten rock, water on glass. Be specific about thickness and wetness, they change timbre.
  2. Action and its shape in time. Not just "cuts" but "presses down in one slow continuous stroke". The model uses this to place events.
  3. Mic perspective. close-mic, intimate, recording distance cues. This sets how dry and near the sound feels, and it works well.
  4. Onomatopoeia sequence. Spell the sound events out in order: squeak, then crack, then dozens of small impacts, then silence. This is the slot that does the most work.
  5. Negatives. no music, no voice, no narration, no room reverb, no ambience bed. Without these H3 will helpfully add a score, because most video in its training data has one.

I also wrote channel directions into slot 4, asking for the crack in the left channel and the seeds rolling left to right. Keep reading for what that actually produced.

Step 3: Run the 768P Draft

Draft at 768P. The audio track is the same specification at both tiers, so the entire creative judgement, which in ASMR is a listening judgement, can be made at the cheap rate.

Settings on minimax/h3/image-to-video: image = the Step 1 output, resolution: 768P, duration: 5. The ratio comes from the first frame, so leave it alone.

Plain
1The santoku blade enters from the left and presses down through the glass pomegranate
2in one slow continuous stroke, travelling left to right across the frame. The shell
3splits with a bright crystalline crack and a spray of tiny glass seeds tumbles onto the
4wet slate, scattering toward the right edge. Camera locked off, macro, single take, no cuts.
5
6Audio: ASMR, close-mic, intimate, no music, no voice, no narration, no room reverb.
7A dry sustained glass squeak as the edge bites in, then one sharp high crack panned to
8the LEFT channel, followed by dozens of small tinkling seed impacts rolling from the
9LEFT channel across to the RIGHT channel as they scatter. A faint blade-on-stone scrape
10at the very end, then silence. No ambience bed, no hum, no background music.

AI video generator interface showing a glass pomegranate video

MiniMax H3 image-to-video playground on Atlas Cloud with the base frame loaded, duration 5, and a completed clip in the OUTPUT panel

The image-to-video run: first frame loaded, duration 5, OUTPUT completed. Note the Resolution select is sitting on the page default of 2K. Switch it to 768P for drafts, which is what the clip below was rendered at.

The 768P draft, $0.50. Softer picture than the hero clip, same audio specification. This is the take the whole decision gets made on.

Step 4: Measure the Stereo Before You Trust It

Do not take my word for the sound, and do not take the model's. Pull the channels apart and look. This is local ffmpeg, no API call, no cost:

Bash
1ffmpeg -i 02-glass-pomegranate-768p-draft.mp4 \
2  -filter_complex "showwavespic=s=1480x240:split_channels=1:colors=#d93a30|#2f7ed8" \
3  -frames:v 1 stereo-proof.png

Waveform comparison of stereo channel separation in two ASMR audio clips

Four-panel waveform analysis comparing the glass pomegranate clip and the rain clip, showing left and right channels plus the side channel for each

Left channel over right channel for two clips, plus the side channel (left minus right) at matched gain underneath. This is the whole argument in one picture.

Here is what came back, and part of it is not what I expected.

The stereo is real. The side channel is not empty on any clip I generated. A dual mono file dressed up as stereo would show nothing there. This one has content.

But you cannot prompt a pan. I asked, in capital letters, for the crack in the LEFT channel and the seeds rolling LEFT to RIGHT. Measured: channel correlation 0.97, side channel sitting 17.7 dB below mid, and the largest left-to-right level difference in any quarter-second window is 1.9 dB. That is not a pan. That is a centred image with a little natural width. The knife travels across the frame; the sound stays put.

Width comes from the content, not from your wording. The rain clip in the next section, where I requested no direction at all, came back at correlation 0.32 with the side channel only 2.9 dB below mid. That is a genuinely wide, decorrelated field. Diffuse ambience gets width automatically. Discrete impacts stay near the centre no matter what you write.

So the honest claim for this model is not spatial. It is temporal. You get sound that was generated alongside the picture and lands on the frame it belongs to, for free, forever, without an editor. If your format truly needs a hard pan, you still have to do that yourself in post, and you should stop writing channel names into prompts because they are doing nothing.

Now re-run the keeper: same endpoint, same first frame, same prompt, change one field to resolution: 2K. One more thing worth knowing before you assume the draft is a preview.

Comparison of two video resolutions cutting a glass pomegranate

Two rows of five frames comparing the 768P and 2K runs at the same timestamps, identical at frame zero and diverging from about 2.5 seconds

Same first frame, same prompt, both tiers. Frame 0 matches exactly. By 2.5s the shell breaks differently and the two clips have become different takes.

Locking the first frame holds the composition, the framing, the lighting and the colour across both tiers, which is why you should always use image-to-video rather than text-to-video for this. It does not give you the same take. The 2K run re-rolls the action. Treat the draft as a preview of the look and the sound, not of the exact choreography.

Step 5: Add a Real Recording as a MiniMax H3 ASMR Video Reference

If you have a sound you actually want, hand it over. reference-to-video takes up to 9 images, 3 videos and 3 audio clips, and it pulls timbre and room character from an audio reference instead of inventing them. I used a public-domain-friendly glass breaking recording by Gravity Sound from Wikimedia Commons (CC BY 4.0), three takes concatenated to 4.5 seconds.

Three rules, and I found the last two the hard way by having requests rejected:

  • Audio cannot go alone. The endpoint spells it out: at least one image or video is required, and audio by itself is not allowed. Pair it with a frame.
  • The audio reference must be 2 to 15 seconds. My first attempt used a 1.49 second clip and came back with audio duration 1486 ms, expected [2000, 15000] ms. That range is not on the model page.
  • The file format is checked. A base64 payload declared as audio/mpeg was rejected with audio format ".mpeg" not allowed. A .wav payload went through. Rejected requests are not billed, but they do cost you a round trip.

Settings: refers: [{url: <base frame>, type: "image"}, {url: <a 2 to 15s recording>, type: "audio"}], resolution: 768P, duration: 5.

Plain
1Same locked-off macro shot: the blade slices the glass pomegranate from left to right
2and the seeds scatter. Match the timbre, brightness and room character of the reference
3audio, same recording distance, same dryness. ASMR close-mic, no music, no voice.

AI video generator interface showing prompt input and generated pomegranate video

MiniMax H3 reference-to-video playground on Atlas Cloud with two reference materials loaded, an audio file in slot 1 and the base frame in slot 2

Reference Materials carrying the audio file and the image together, 2 of 9 used. Drop the image and the request will not run at all.

The reference-guided take, $0.50. Same shot, timbre steered by a real recording rather than invented from the prompt. Audio reference: Glass breaking by Gravity Sound, CC BY 4.0.

Three MiniMax H3 ASMR Video Templates: Cutting, Chewing, Rain

Three formats cover most of what performs in this category. Each is a full paste-and-run prompt with settings and the clip it produced. Deliberately no glass, no red, no slate anywhere below: if every clip on your account looks the same, the algorithm treats it as the same clip.

Table 3: the same five slots, three different jobs.

SlotTemplate 1, The CutTemplate 2, The ChewTemplate 3, The Rain
Materialdripping honeycomb, hot steel bladeglowing molten rock, stone bowlrain on car window glass
Action shapeblade sinks front to back, honey pulls downchopsticks lift, then two bitesno subject action, droplets only
Mic perspectiveclose-mic, very dry, no roomextremely close, intimateoutside the glass, cabin muffled
Sound sequencewax crack, honey stretch, drip on platemolten sizzle, wet chew, glassy crunch, exhalesteady rain bed, droplet hits, distant tyre hiss
Negativesno music, no voice, no room reverbno words, no music, no narrationno music, no thunder, no voice, no wipers

Template 1, The Cut. Honeycomb, amber and gold, 9:16, 768P, 6 seconds, ratio: "9:16".

Plain
1Extreme macro, locked-off overhead shot of a thick slab of golden honeycomb on a pale
2ceramic plate, honey already pooling around it. A hot steel blade lowers into the wax
3and sinks through it front to back in one slow continuous press; the cut faces slump
4and a heavy thread of honey stretches and drips onto the plate. Warm amber key light
5from the left, shallow depth of field, single take, no cuts, no hands in frame.
6
7Audio: ASMR, close-mic, very dry, no room reverb, no music, no voice, no narration.
8A soft crackling of wax splitting under the blade, then a thick low stretching pull as
9the honey lengthens, then two heavy wet drips landing on the ceramic plate. Nothing else.

Template 1, $0.60. Wax crack, honey stretch, drip. Generated with minimax/h3/text-to-video.

Template 2, The Chew. Lava mukbang, the format that pulled 11.3 million views in a week, 9:16, 768P, 8 seconds, ratio: "9:16".

Plain
1Vertical close-up: a pair of black lacquer chopsticks lifts a glowing clump of molten
2orange lava out of a dark stone bowl and carries it toward the camera, out of frame at
3the bottom. The lava is self-illuminating, casting orange light on everything around it;
4the rest of the room is almost black. Steam curls off the surface. Locked-off camera,
5single take, no cuts.
6
7Audio: ASMR, extremely close-mic, intimate, no words, no music, no narration, no room tone.
8A low molten sizzle and bubbling as the lava is lifted, then a soft wet chew, then a
9brighter glassy crunch on the second bite, then one slow satisfied exhale. No speech.

Template 2, $0.80. Sizzle, chew, crunch, exhale, and not a single word. Generated with minimax/h3/text-to-video.

Template 3, The Rain. Night car window, cold blue and neon, 16:9, 768P, 10 seconds, ratio: "16:9".

This is the only one of the three with no subject action, and it teaches the most about negatives. With nothing happening on screen, H3 reaches for a music bed unless you forbid one explicitly. It is also the clip that came back with by far the widest stereo field, without being asked. Sleep-oriented content wants length, so this runs near the top of the useful duration range.

Plain
1Macro shot through the inside of a car window at night, heavy rain running down the
2outside of the glass. Individual droplets hit, hesitate, then streak downward and merge.
3Beyond the glass, out-of-focus neon signage breaks into cold blue and magenta bokeh.
4No wipers, no people, no movement inside the car. Camera locked off, single continuous
5take, shallow depth of field, no cuts.
6
7Audio: ASMR, close-mic on the outside of the glass, cabin slightly muffled.
8A steady even rain bed with clearly separated individual droplet impacts on the glass,
9plus a faint distant hiss of tyres on a wet road. No music, no thunder, no voice,
10no narration, no wiper noise.

Template 3, $1.00. Ten seconds of pure ambience, no soundtrack because the prompt banned one. Generated with minimax/h3/text-to-video.

What a MiniMax H3 ASMR Video Actually Costs

Here is the receipt for this article, not an estimate.

Line itemCountBilled
GPT Image 2 base frame, high, 2048x11521$0.17
2K keeper, 5s, image-to-video1$0.70
768P draft, 5s, image-to-video1$0.50
768P reference-to-video, 5s1$0.50
Template clips at 768P, 6s + 8s + 10s3$2.40
Rejected reference requests2$0.00
Total$4.27

Six publishable clips and one base frame. Scale it to a schedule: one 8 second vertical clip a day at 768P is $0.80, so roughly $24 a month for a daily-posting faceless account, before you spend a minute in an editor.

Two billing notes. GPT Image 2's listed $0.009 is a token-tier floor; a high-quality 2048x1152 render lands at $0.1745, nearly twenty times that, so budget from the tier you actually use. And H3's price field backfills on a delay: poll until status is completed and it is frequently still undefined. Poll the prediction id again a moment later for the real figure. That second poll is the only way to build a receipt like this one instead of a guess.

One Honest Note on ASMR Audio and Rights

Do not upload somebody else's ASMR recording as a refers audio file. The reference genuinely migrates timbre into the output, and running a recording through a model does not change who owns it. The reference used here is CC BY 4.0 and credited above, which took about two minutes to source.

Do not build ASMR around a recognisable real person's voice. Every template in this article is deliberately voiceless, which is both the genre convention and the least complicated path.

Calling H3 through an API is not affected by the community licence attached to the open weights. That licence, which covers self-hosting, currently excludes the EU, UK, Korea and the US, and a formal licensing channel is open for those territories. Worth knowing, not worth worrying about if you are hitting the endpoint.

MiniMax H3 ASMR Video: Frequently Asked Questions

Does a MiniMax H3 ASMR video generate its own sound, or do I add it afterwards?

The model generates it. Picture and audio come out of the same pass: 24 fps video with an AAC stereo track at 32 kHz in the delivered file. There is no separate audio step, no sound library and no alignment work, which is the entire reason this endpoint is interesting for ASMR specifically.

Can I make a MiniMax H3 ASMR video pan from left to right?

Not by asking. I wrote explicit channel directions into the prompt and measured the result: correlation 0.97 between channels and a maximum level difference of 1.9 dB in any quarter-second window, which is no pan at all. The track is genuinely stereo and diffuse content like rain comes back with a wide field, but discrete impacts stay centred regardless of wording. If your format needs a hard pan, do it in post.

Can I upload my own recording as a MiniMax H3 ASMR video reference?

Yes, through reference-to-video, which accepts up to 3 audio references alongside up to 9 images and 3 videos. Audio cannot be submitted on its own, so pair it with at least one image or video. The audio also has to be between 2 and 15 seconds, and the format is validated: a .wav payload worked where an audio/mpeg one was rejected. Rejected requests are not billed.

Is a 768P MiniMax H3 ASMR video's audio worse than 2K?

No. Both tiers deliver the same AAC stereo 32 kHz specification. Only the picture changes, and it changes a lot: 16:9 comes back as 1344x768 at 768P versus 2560x1440 at 2K. Since the creative judgement in ASMR is a listening judgement, draft at $0.10/s and re-run keepers at $0.14/s. Just remember the 2K run is a new take, not an upscale of the draft.

How long should a MiniMax H3 ASMR video be, and what aspect ratio?

Duration accepts any whole second from 4 to 15. Cutting and chewing clips work at 5 to 8 seconds because the payoff is one event; sleep-oriented ambience wants 10 or more. For Shorts, Reels and TikTok use 9:16, and remember that text-to-video requires ratio to be set explicitly and rejects adaptive. On image-to-video the first frame decides the shape for you.

How much does one MiniMax H3 ASMR video cost?

A 5 second clip bills $0.50 at 768P and $0.70 at 2K. An 8 second vertical clip at 768P is $0.80. Posting one of those a day is about $24 a month in compute, which is the number worth comparing against the 20 minutes an editor would take per clip on the stitched workflow.

Why does my MiniMax H3 ASMR video come out with background music I never asked for?

Because you did not forbid it. Most video in any model's training data has a music bed, so the default behaviour is to add one, especially when nothing is happening on screen. Put the negatives in the audio half of the prompt explicitly: no music, no voice, no narration, no room reverb, no ambience bed. Ambient templates need this more than action ones do.

Latest Models

One API for All Media AI.

Explore all models