New AI models in AI Studio

Try Now
PRIZMAD

New in AI Studio: MiniMax H3 — Reference-Grade Video Generation, Inside Prizmad

MiniMax H3 is live in Prizmad's AI Studio: 5-15s video from text, images, or references, up to 4K, 9:16 and 21:9 ratios, from 60 tokens.

Prizmad Team4 min read
A stylized vertical video frame with a product in motion and cinematic controls, representing MiniMax H3 generation in AI Studio
On this page

Prizmad's AI Studio added another frontier video model this week: MiniMax H3, live since September 4 on every paid plan. Where the avatar pipeline turns a product URL into a finished UGC-style ad, MiniMax H3 is the asset engine beside it — high-resolution motion clips generated from a text prompt, a product image, or a full set of brand references, at resolutions up to 4K and aspect ratios that actually match ad placements.

What it actually does

MiniMax H3 runs in three modes, each built for a different starting point:

  • Text-to-Video — describe the subject, scene, action, camera movement, lighting, and even audio direction (room ambience, a product-impact sound at the hero frame); the model generates the clip from scratch. Output runs 5–15 seconds at 480P (fastest), 768P native HD, or 2K/4K (upscaled from a 768P base).
  • Image-to-Video — upload a Start Image (and optionally an End Image for a controlled first-to-last transition). The output follows your image's own aspect ratio automatically — a clean product shot becomes a moving hero clip without fighting a frame mismatch.
  • Reference-to-Video — the mode that matters most for brand work. Upload up to 9 reference images, 3 reference videos, and 3 reference audio files (2–15 seconds each, combined video and audio capped at 15 seconds each), then address them in the prompt as Image 1, Image 2, Video 1, and Audio 1. The model holds the product, the presenter, the motion style, and the voice feel consistent across the generation.

Why reference mode is the ecommerce hook

Most video models give you one chance at consistency: describe the scene, hope the product renders correctly, regenerate when the label drifts. MiniMax H3's reference mode is built for the opposite workflow — you give the model the ground truth up front.

The practical pattern for an ad team: Image 1 is the product shot from your catalog, Image 2 is the spokesperson or the brand's visual style, Video 1 is the camera motion you liked from a previous winning clip, Audio 1 is the voice or sound direction. The output inherits the product's actual packaging, the presenter's actual face, and the pacing of footage you already know converts — rather than a model's guess at what your bottle looks like.

The first 5 reference images are free on every reference generation; each additional reference image adds a small token cost (16 tokens), and video/audio references meter into the same flat per-second rate as the base clip.

Ad-native output controls

Three details are worth calling out for anyone who has fought a video tool over format:

  • 9:16 and 21:9 are native options. Vertical for Reels, Shorts, and TikTok; 21:9 for cinematic in-stream or YouTube bumper-adjacent placements; 16:9, 4:3, 1:1, and 3:4 in between. Text-to-Video lets you pick the exact frame; Reference-to-Video defaults to Adaptive (the model chooses) but lets you lock a ratio when orientation matters.
  • Prompt Expansion is a dial, not a mystery. Disabled, fast, balanced (default), or quality — you control how much effort the model spends rewriting your prompt before generating. Quality spends longer for a richer prompt when the scene is complex; Disabled keeps your wording literal when you need the output to match a spec exactly.
  • A seed field is exposed for reproducible variant testing — run the same prompt and settings with a fixed seed and compare output takes that start from the same noise, which is closer to a controlled creative test than a blind reroll.

What it costs

Token cost updates automatically by duration and resolution. Source-verifiable examples from the tool's own math (1 token = $0.01 of provider cost ×2):

  • 60 tokens — the default: 768P, 5 seconds, text-to-video. That's about $0.78 at Launch pricing ($39/mo for 3,000 tokens), $0.74 on Starter ($99/mo for 8,000), $0.43 on Pro ($249/mo for 35,000).
  • 180 tokens — a 15-second clip at 768P (the longest single take at native HD).
  • 390 tokens — 15 seconds upscaled to 2K.
  • 480 tokens — 15 seconds at 4K.

Renders are asynchronous with an estimated ≈2–6 minute turnaround for the default configuration (queue dependent) — single-digit minutes for a shot that, on the raw model API, would mean managing your own queue and per-second invoices. Inside Prizmad it draws from the same token pool as the avatar pipeline and the other 30+ AI Studio tools: no separate signup, no API key, no surprise line item.

Where it fits in the ad pipeline

MiniMax H3 is a raw generation tool, not a finished-ad factory — the honest division of labor is:

  1. Generate the motion assets: a product reveal at 4K for the hero frame, a lifestyle clip in 9:16 for the hook, a camera-motion take you want to reuse.
  2. Keep the product honest with Image-to-Video from your own catalog shot, or lock brand consistency with Reference-to-Video.
  3. Feed the clips into the URL-to-ad pipeline where the avatar delivers the script, or use them as B-roll inside any ad you're assembling in AI Studio.

The bigger picture is the same one we covered when Seedance 2.5 joined AI Studio: frontier video models are becoming a commodity layer, and the difference for an ecommerce advertiser is what sits around the model — the product data, the script, the avatar, the ratios, and a pricing model you can predict before you click generate. MiniMax H3 is another step in that direction: try it on your own product shots — 60 tokens is a cheap first look at whether reference-driven motion holds your brand's details.

Generate Your First Ad in 5 Minutes

Paste a product URL. Prizmad writes the script, picks the avatar, renders the voiceover with lip-sync, adds subtitles and music, and ships a TikTok / Meta / YouTube-ready mp4. No camera, no editor.