MiniMax H3 Ref2V (Ref2VA)

MiniMax H3 Ref2V:
References In.
Audio-Video Out.

Ref2V is how most people search for MiniMax H3's reference-to-audio-video mode: you upload reference images, optional clips, and optional voice audio, and H3 performs a new scene with them as the cast — and generates the stereo soundtrack in the same pass. No subscription, pay per clip with credits.

Reference slots9 images + 3 videos + 3 audio
AudioNative stereo, same pass
Output768P or 2K, 4-15s
PricingCredits, no subscription
MiniMax H3 Ref2V (Ref2VA)
Smooth motion
Rich details
Cinematic lighting
Built for creators, storytellers, and teams who want production-ready results in minutes.

Model Options

Choose the workflow that fits your scene

Start with a prompt, add references when needed, and pick the generation path that matches the output.

MiniMax H3 Ref2VA

The reference-to-audio-video mode of MiniMax H3, exposed on Moosky as an omni-reference composer. Feed it a reference sheet — face images, wardrobe or product shots, motion clips, voice takes — label each reference in the prompt, and H3 keeps the identity, the styling, and the voice while performing your new scene with synchronized stereo sound.

Up to 12 reference files (9 img / 3 video / 3 audio) Native 32 kHz stereo dialogue + ambience 768P or 2K, 16:9 or 9:16, 4-15 seconds Stable dialogue in 11 languages
Open the H3 Ref2V composer

Other reference-to-video models

Ref2V is MiniMax H3's take on R2V. When a shot needs heavier motion or different framing, your same credits run Wan 3.0 R2V or Seedance 2.0 R2V — the full lineup and how to pick between them lives on the reference-to-video hub.

One credit balance Switch engines per shot Cost shown before every run Commercial use included
Browse the R2V hub

Why use it on Moosky AI

Powerful tools. Seamless experience.

Video and audio in one pass

Ref2V's defining trick: the A in Ref2VA. Dialogue, ambience, and score cues render synchronized with the picture — no separate TTS or dubbing step.

Omni reference sheet

Up to 9 images, 3 video clips, and 3 standalone audio files per job (12 files max) — identity, wardrobe, style, motion, and voice all in one bundle.

Character-consistent scenes

The same face, outfit, and product across takes. H3's default strength is talking heads and dialogue scenes that have to stay recognizably the same person.

Voice carried by reference

Attach a voice clip as a reference and the performed line keeps that timbre — across languages, with lip sync baked in.

11-language dialogue

Stable speech in Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.

Pay per clip

H3 Ref2V runs on Moosky credits (100 per $1). The exact cost of your references, length, and resolution is quoted in the composer before you generate.

How It Works

Three simple steps to finished output

1

Build the reference sheet

2

Direct the scene

3

Render video + audio

Credit-Based Generation

Simple, honest pricing

No subscriptions. No hidden fees. Pay only for the credits you use.

Pay as you goOnly pay when you generate.
No monthly feesNo subscriptions. No commitments.
Credits never expireUse your credits whenever you want.
Commercial useUse outputs for personal or commercial projects.
View credits & pricing

Frequently Asked Questions

What does Ref2V mean on MiniMax H3?

Ref2V is shorthand for reference-to-video: you give the model references instead of just a prompt, and it generates a clip conditioned on them. On MiniMax H3 the mode is formally Ref2VA — reference to audio-video — because the same pass also produces synchronized stereo sound (dialogue, ambience, music cues).

What can I feed the Ref2V composer?

Up to 9 reference images, 3 reference videos, and 3 standalone audio clips, 12 files total. Each video or audio clip should run 2-15 seconds. Standalone audio must accompany at least one image or video — audio can't be the only input.

How do I control what each reference does?

Order them intentionally and label them in the prompt: Image 1 is the face, Image 2 the wardrobe, Video 1 the camera move, Audio 1 the voice. H3 reads those labels and assigns each reference a role — identity, style, motion, or sound.

How is H3 Ref2V different from image-to-video?

Image-to-video animates one still — the frame itself becomes the shot. Ref2V conditions on identity instead: place the same person or product in new scenes, angles, and actions and they stay consistent take after take, and with H3 they speak with a consistent voice too.

How much does MiniMax H3 Ref2V cost on Moosky?

Generations are billed in credits (100 credits per $1). H3 pricing scales with output length, resolution (768P vs 2K), and input materials — the first 5 images are free, extra images and reference-video duration are billed. The composer shows the exact quote before you run, and failed runs are refunded.

Disclosure

Moosky AI is an independent platform provider and reseller of AI model access, including access to MiniMax H3. Moosky AI is not officially affiliated with, sponsored by, or endorsed by MiniMax. This page and its examples may include AI-assisted and synthetic content.