MiniMax H3 on EVO-X2: FL2VA Turn Keyframes, 4.5s Band, and Audio Pitch Control
Contents
Following the previous MiniMax H3 test, where we ran pruned GGUF models on the EVO-X2 to generate audio-video from a single Kana-chan character illustration and two reference images, this follow-up evaluates first and last frame keyframing, an extended 4.5-second dialogue, a 4-piece rock band, and pitch conditioning using reference audio.
Test Environment
Generations were executed on the same GMKtec EVO-X2 ROCm environment as the previous post.
| Item | Details |
|---|---|
| Hardware | GMKtec EVO-X2 (Ryzen AI Max+ 395, Radeon 8060S / gfx1151, 48GB allocated VRAM) |
| Runtime | stable-diffusion.cpp master-929-3f8527a (win-rocm-7.14.0-x64) + llama.cpp HIP build DLLs |
| Base Weights | unsloth/MiniMax-H3-GGUF (fl2va_pruned-Q4_K, ref2va_pruned-Q4_K) |
| Text Encoder | qwen3vl_32b_minimax_h3-Q4_K_M |
| Common Settings | 8 steps, CFG 1.0, seed 42, --rng cpu, --diffusion-fa, GGML_CUDA_DISABLE_GRAPHS=1 |
fl2va conditions video on first and last frames, while ref2va conditions on reference visual and audio inputs.
CFG adjusts prompt adherence during generation; here it was fixed to 1.0.
Because reference image conversion stalled in the previous test, HIP graphs were disabled via GGML_CUDA_DISABLE_GRAPHS=1.
| Experiment | Model | Input | Resolution | Frames |
|---|---|---|---|---|
| Character turn | fl2va | First and last frames | 416x608 | 56 (~2.3s) |
| Jump | fl2va | First and last frames | 416x608 | 56 (~2.3s) |
| Duo dialogue | ref2va | 2 reference images | 640x384 | 107 (~4.5s) |
| 4-piece band | fl2va | First frame | 672x384 | 107 (~4.5s) |
| Without reference audio | ref2va | 1 reference image | 416x608 | 73 (~3s) |
| With reference audio | ref2va | 1 reference image + reference audio | 416x608 | 73 (~3s) |
| WAV only (ablation) | ref2va | 1 reference image + reference audio (no matching prompt) | 416x608 | 73 (~3s) |
First and Last Keyframe Conditioning
The fl2va model supports conditioning on both the first and last frames to interpolate video between them (FL2VA mode).
The initial frame is specified with -i, and the final frame with --end-img.
For both endpoints, we used back-view and jumping poses edited from the front-facing standing character art in our local Qwen-Image 2.1 article.
Character Turn

In the back-view illustration, the side ponytail and blue scrunchie sit on the left side of the frame (the character’s left).
In the front-facing illustration, they sit on the right side of the frame (also the character’s left), so the tied side matched at both endpoints.
The prompt instructed the character to turn slowly to her left in place until her back faced the camera, while keeping the ponytail on the same side:
Anime style animation. The girl with brown side ponytail and blue scrunchie stands facing the viewer, then slowly turns around in place to her left until her back faces the camera. Her side ponytail stays tied on the same side of her head while she turns. Smooth natural rotation, hair and skirt sway gently. Static camera, clean white background, consistent character design. Quiet room tone, soft footsteps.

By frame 11, the face turned toward the right side of the frame; by frame 22, it reached a direct profile; by frame 33, the back faced the camera.
Checking every 3 frames, the character began rotating rightward (toward her left) at 0.38s, completed the back view by 1.38s, adjusted her footing, and remained mostly stationary from 1.50s onward.
Across the six extracted frames, the side ponytail rotated behind the head during profile view and re-emerged on the left in the rear view.
The hair tie remained consistently on the character’s left side throughout this sequence.
In the previous duo video, turning sideways caused the ponytail to flip sides.
Here, that swap did not occur; however, because generation mode and framing also differed, the last-frame constraint was not the only variable.
Measured average audio volume was -50.1 dB.
Here, dB is relative to digital full scale (0 dBFS), where more negative values indicate quieter sound.
Although the prompt specified quiet room tone and soft footsteps, neither was audible on playback.
Jump

Anime style animation. The girl with brown side ponytail and blue scrunchie bends her knees, then jumps up energetically, thrusting her right fist into the air with an excited open-mouth smile. Her hair and skirt fly up with the motion. Static camera, clean white background, consistent character design. She shouts cheerfully in Japanese: "よーし、がんばるぞー!"

At frame 11, she bent her knees and crouched; at frame 22, she jumped upward; by frame 33, she reached the final pose.
Checking every 3 frames, she arrived at the final pose at 1.12s and remained frozen mid-air for about 1.1 seconds until 2.25s.
Because the final frame depicts an aerial pose, the clip ended before she touched the ground.
Transcription with Whisper large-v3-turbo produced: “よーし、がんばるぞー!” (Yoshi, ganbaruzo-!).
Listening to the audio matched the specified line.
As with previous videos, lip-sync did not align with speech.
We overlaid estimated fundamental frequency (F0) on the spectrogram.
While harmonics appeared between 0.3s and 0.8s, the pYIN algorithm did not classify most of that segment as voiced.

The median F0 of 580.1 Hz was calculated almost entirely from the voiced speech after 1.3s.
Extending Duo Dialogue to 4.5 Seconds
In the previous test, the dialogue between two characters exceeded the 56-frame limit (~2.3s) and cut off mid-sentence.
Keeping the prompt and seed identical, we increased the frame count to 107 (~4.5s).

| Time Range | Transcription | Median Pitch (F0) |
|---|---|---|
| 0.0–3.2s | ねえ今日の帰りクレープ食べに行こうよ (Hey, let’s go get crepes on the way home today) | 417 Hz |
| 3.2–4.3s | いいね行こ (Sounds good, let’s go) | 410 Hz |
The transcript captured both the first character’s line and the second character’s reply.
For the prompt’s “いいね、行こう!”, Whisper transcribed “いいね行こ”, and auditory check confirmed the missing trailing “u”.
Kana-chan’s mouth was open between frames 21 and 63, while the blonde character’s mouth opened at frame 84.

Across 4-frame inspection intervals, Kana-chan’s mouth remained open from 0.50s to 3.17s while the blonde character kept her mouth closed.
From 3.50s onward, the blonde character opened her mouth while Kana-chan smiled with eyes closed and mouth shut.
This alternation broadly corresponded to the analyzed audio segments.
However, actual lip movement did not match speech cadence upon playback; mouths often froze or shut mid-phrase.
Pitch was compared using fundamental frequency (F0) estimated with librosa’s pYIN algorithm on 16kHz mono audio (80–800 Hz range), restricted to voiced frames.
The median pitch for both speakers was nearly identical.
To compare vocal characteristics beyond pitch, we evaluated speaker embeddings using SpeechBrain ECAPA-TDNN.
ECAPA-TDNN extracts speaker vectors from 16kHz mono audio and compares them via cosine similarity (-1 to 1; higher indicates closer characteristics).
Cosine similarity between the first segment (0–3.2s) and the second segment (3.2–4.46s) was 0.532.
This is lower than the 0.885–0.980 similarity observed between outputs conditioned on the same reference WAV.
Because the lines differed and the second clip was short (~1.2s), whether the speakers were distinct cannot be determined from this metric alone.
On listening, however, the two voices sounded distinct.
The side ponytail shifted to the left side of the frame at frame 21 as she faced the blonde girl, consistent with the previous run.
Generating a 4-Piece Band Performance
Using the 4-piece band image generated via natural language role prompts from our early access article as the first frame:

The prompt specified each role, Japanese female vocal J-rock, and instruments synchronized with hand movements:
Anime style animation. A four-girl rock band performs live on a dark stage with purple spotlights. The blonde girl in the center sings into the microphone while strumming her white electric guitar, the pink-haired girl on the left plays the black bass guitar, the short-haired girl in the back plays the drums with sticks, and the brown-haired twin-tail girl on the right plays the red electric guitar. Energetic upbeat J-rock music with driving drums, distorted guitars, bass, and a bright female vocal singing in Japanese. Instruments and hand movements are synchronized with the music. Stage lights flicker, slow push-in camera.

Across the six extracted frames, hairstyles, instruments, and outfits remained stable without swapping.
Characters grew larger across successive frames, in line with the push-in camera prompt.
| Item | Result |
|---|---|
| Volume | Mean -11.0 dB, max -0.0 dB |
| Estimated tempo (librosa) | ~94 BPM |
| 20–150 Hz energy ratio | 47% |
| Whisper transcription (Japanese specified) | You choose a baby, tell your love to die, tell your love to die |
Audio was louder than dialogue clips (mean -15.7 to -19.8 dB), with low frequencies accounting for nearly half the energy.
Despite specifying Japanese, Whisper returned English-like lyrics.
This does not distinguish whether singing was non-Japanese or whether transcription failed on sung vocals.
Listening to the track revealed singing, but words were unintelligible.
Examining drummer crops every 2 frames alongside audio onset timestamps showed that hands and drumsticks were occluded by the foreground guitar neck and the right guitarist’s arm; visual synchronization could not be confirmed.
On playback, guitar strumming and drum movements generally followed the music, though fingering precision could not be confirmed.
The drummer struck cymbals, but hi-hat strikes were not visible.
Voice Pitch Control with Reference Audio
Between Kana-chan’s solo clip and the duo clip in the previous test, vocal pitch differed noticeably:
| Video | Median Pitch (F0) |
|---|---|
| Kana-chan solo (previous) | 618 Hz |
| Duo dialogue (previous) | 393 Hz |
The ref2va model supports specifying a WAV file as reference audio via --ref-audio.
We extracted a 1.9-second segment (“Yaho, Kana dayo!”) from the previous solo clip to use as reference audio, referenced in prompts as <Audio 1>.
We generated two clips with identical image, seed, resolution, step count, and dialogue line.
For the reference audio condition, we passed the WAV and appended a voice-matching instruction to the prompt:
Anime style animation. Use the girl from <Picture 1> (brown side ponytail with a blue scrunchie, red necktie, navy skirt, black socks). Keep her appearance and identity consistent with the reference. Her voice must match the voice in <Audio 1>: same speaker, same pitch and timbre. She stands in a sunny park, looks at the viewer, smiles and talks cheerfully with small hand gestures. Medium full shot, static camera, soft daylight, birds chirping. She says in Japanese: "今日はいい天気だね、お散歩に行こうよ!"
The condition without reference audio omitted Her voice must match ....
To isolate each component’s effect, we also tested passing the WAV file alone without the voice-matching sentence.
Without reference audio:
With reference audio:
| Audio | Transcription | Median Pitch (F0) | 10th–90th Percentile Range |
|---|---|---|---|
| Reference audio (previous) | ヤッホー!カナダよ! | 615 Hz | 506–739 Hz |
| Without reference audio | 今日はいい天気だねお散歩に行こうよ | 425 Hz | 315–521 Hz |
| With reference audio | 今日はいい天気だね。お散歩に行こうよ。 | 538 Hz | 429–633 Hz |
Adding reference audio and the prompt instruction raised median pitch from 425 Hz to 538 Hz, closer to the reference’s 615 Hz.
However, it did not reach the reference pitch itself.
Listening to both clips confirmed the spoken line “今日はいい天気だね、お散歩に行こうよ” (Nice weather today, let’s take a walk!).
| Audio | Listening Impression |
|---|---|
| Without reference audio | Slightly lower than reference, but natural |
| With reference audio | Closer to reference, but muffled as if underwater |


Both outputs maintained Kana-chan’s character features with the ponytail on her right side.
The reference audio version closed its eyes in a smile during the second half.
Ablation: Reference Audio File vs Prompt Instruction
We added a condition passing only the WAV file to the base prompt (without the voice-matching sentence).
Keeping image, resolution, steps, and line constant, we tested all three conditions across seeds 42, 43, and 44.
Seed 42 outputs used the clips above, with 7 additional clips generated for seeds 43 and 44.
Seed 42 output with WAV only:
| Condition | Median F0 (Seeds 42 / 43 / 44) | ECAPA Similarity to Reference WAV (42 / 43 / 44) |
|---|---|---|
| None | 425 / 507 / 462 Hz | 0.144 / 0.350 / 0.241 |
| WAV + prompt instruction | 538 / 540 / 523 Hz | 0.438 / 0.442 / 0.446 |
| WAV only | 538 / 541 / 520 Hz | 0.444 / 0.396 / 0.424 |
Across all three seeds, passing the WAV file alone yielded nearly identical pitch to adding the matching prompt clause.
Similarity to the reference was also comparable; the added prompt instruction showed no measurable effect.
| Comparison | Pairwise Similarity |
|---|---|
| 3 pairs without WAV across seeds | 0.726–0.887 |
| 15 pairs among 6 WAV-conditioned outputs | 0.885–0.980 |
| WAV + prompt vs WAV only (same seed) | 0.980 / 0.921 / 0.953 (seeds 42 / 43 / 44) |
Outputs conditioned on WAV had high speaker embedding similarity across different seeds.
However, similarity to the reference WAV itself remained between 0.396 and 0.446, compared to 0.898 between the source clip and the reference WAV.
ECAPA was trained on VoxCeleb human speech; its model card does not guarantee calibration on synthetic or short audio clips.
These figures are for reference only.
Listening to seeds 43 and 44 matched seed 42 trends: clips without WAV had slightly lower pitch, while WAV-conditioned clips moved closer to reference pitch regardless of prompt instructions, accompanied by a muffled quality.
| Condition | Seed 43 Video | Seed 44 Video |
|---|---|---|
| None | 43 | 44 |
| WAV + prompt | 43 | 44 |
| WAV only | 43 | 44 |
Generation Time and Memory
| Experiment | Frames | Resolution | Generation | Video Decode | Total |
|---|---|---|---|---|---|
| Turn (first and last) | 56 | 416x608 | 173.32s | 32.90s | 223.41s |
| Jump (first and last) | 56 | 416x608 | 173.76s | 32.97s | 223.56s |
| Duo dialogue | 107 | 640x384 | 319.60s | 66.60s | 402.38s |
| 4-piece band | 107 | 672x384 | 332.39s | 83.67s | 432.37s |
| Without reference audio | 73 | 416x608 | 198.91s | 43.14s | 255.70s |
| With reference audio | 73 | 416x608 | 203.85s | 43.70s | 261.53s |
Conditioning on both first and last frames added ~20 seconds to generation compared to first-frame only (152.68s for 416x608, 56 frames in the previous test).
For the duo dialogue, doubling frame count (~1.9x) doubled generation time (~2.0x).
During additional runs, we tracked Windows memory metrics across 13 completed generations:
| Metric | Initial Value | Recorded Min | Recorded Max |
|---|---|---|---|
| Committed Memory | 15.93–17.42 GiB | 14.64–17.42 GiB | 50.71–52.96 GiB |
| Commit Limit | — | — | 51.44–53.54 GiB |
| Available Physical RAM | — | 4.08–5.20 GiB | — |
| Dedicated GPU Memory | 1.20–2.10 GiB | — | 35.28–36.93 GiB |
| Shared GPU Memory | 3.04 GiB | 3.04 GiB | 3.04 GiB |
| sd-cli Private Bytes | — | — | 3.58–3.70 GiB |
Ranges summarize start, minimum, and maximum values across 13 runs.
The commit limit was 46.90 GiB at the start of the first run.
Metrics are reported in GiB (PowerShell 1GB = 2^30 bytes).
Committed memory and dedicated GPU memory each rose by ~35 GiB during generation.
The commit limit expanded dynamically rather than remaining fixed.