Tech11 min read

Anima 4-char LoRA dresses the wrong girl in its DiT half, not the Qwen encoder

IkesanContents

Update (2026-08-14): A follow-up tests how far prompts alone can assign expressions, poses, a two-girl high-five, and foreground/background placement with the same 4-girl LoRA → Anima’s 4-character LoRA high-fives whoever stands center, not who you name

In the previous 2×2 experiment that swapped standing positions and outfit assignment, Kana and Koharu swapped positions and actions just fine, but the black tights went to whoever stood foreground-right, 12/12. Restating the owner by name or by spatial role at the end of the prompt didn’t fix it, and shortening the prompt broke other parts of the scene.

Next, I wanted to pin down where this mis-dressing happens: in Qwen, T5, the LLMAdapter, or the DiT. I split the 4-char LoRA checkpoint by module and generated with the same prompt and noise.

Test environment

ItemSetting
EnvironmentComfyUI on M4 Mac mini
Modelanima-base-v1.0
Text encoderqwen_3_06b_base
4-char LoRAanima-4char-v1_epoch100
Turbonone
Samplerer_sde / simple
Steps / CFG25 / 4.0
Resolution1344×768

The main comparison is seed42, with the prompt, the T5 token ID sequence, the noise, and the sampler kept the same in every condition.

Which module produces the 4-char identities and the art style

Classifying the checkpoint keys by target module, there were zero LoRA entries for the Qwen model itself.

TargetLoRA modulesTensors
Qwen3-0.6B text encoder00
LLMAdapter36108
DiT280840

That’s 316 modules and 948 tensors in total. I split the file into LLMAdapter only and DiT only, and first compared them on a prompt with no outfit assignment. I used cell 110 from last time, where Kana’s wave and the Kurara–Kei conversation both worked, with the same seed42, prompt, and noise. For comparison I also generated full (the whole 4-char LoRA applied) and plain base Anima under the same conditions.

full LoRA. The reference image with the four identities, the conversation, and the wave at the back

full. The reference point, with the four identities and the art style intact.

base Anima. The scene layout comes out, but the four girls don't look like the trained characters

base Anima. The characters collapse into four strangers, and the front two are tangled at the arms and faces in a way that makes the picture genuinely unpleasant. Only the layout survives: two facing each other in front, one on the side, one waving at the back.

LLMAdapter-only. Close to base, without the 4-char LoRA's style or characters

LLMAdapter-only. Almost the same picture as base Anima, characters still collapsed.

DiT-only. Identities and art style close to the full LoRA

DiT-only. Identities and art style stay almost identical to full, down to the same lower-body outfit mistake.

This is still one prompt and one seed, but it suggests that base Anima alone can handle the layout, while the DiT-side LoRA is what keeps the characters from collapsing and carries the art style. For this image at least, the LLMAdapter LoRA wasn’t needed. Whether it’s unnecessary on other prompts is something this kind of split can’t answer.

Splitting the exact P1O0 that dressed the wrong girl

Cell 110 above had no outfit assignment in its prompt. The next split uses the prompt that actually got the outfit wrong.

I reused P1O0 seed42 from the previous 2×2, the cell where only the standing positions were swapped. In the P1O0 prompt, Koharu waves at the back, Kana crosses her arms foreground-right, and the skirt and black tights are assigned to Koharu at the back. The correct output has Koharu, waving at the back, wearing the black tights; in the previous full run, Kana at foreground-right wore them instead.

With this prompt and noise unchanged, I generated base, LLMAdapter-only, and DiT-only. Qwen3-0.6B, the T5 token ID sequence, and the sampler stayed the same. For full I reused the P1O0 image already generated in the 2×2.

Condition4-char identities & styleWho wears the black tightsMatches the P1O0 assignment
fullKana, foreground-right×
basen/athe girl waving at the back
LLMAdapter-onlyn/a, base-likethe girl waving at the back
DiT-only○, full-likeKana, foreground-right×
full 4-char LoRA on P1O0. The black tights are on Kana foreground-right, not on Koharu at the back as assigned

full. The four trained characters appear, and the black tights are on Kana crossing her arms foreground-right, not on Koharu at the back as assigned. Koharu, waving at the back, is bare-legged.

base Anima on P1O0. With the same Qwen and T5, the black tights are correctly on the girl waving at the back

base Anima. Same as with 110: the characters collapse, the front two are tangled, and the composition differs from full. But the black tights are on the girl waving at the back (Koharu’s position), exactly as the prompt assigns.

LLMAdapter-only on P1O0. A base-like frame, with the black tights on the girl waving at the back

LLMAdapter-only. Almost the same picture as base, and the black tights again on the girl waving at the back.

DiT-only on P1O0. Full-like identities and style, with the black tights on Kana foreground-right

DiT-only. Almost the same frame as full, with the four trained characters, and the black tights on Kana crossing her arms foreground-right.

The four girls in base and Adapter-only are still strangers to the trained characters. Even so, from the same 0.6B Qwen/T5 conditioning, the black tights went to the girl waving at the back and the navy skirt to the girl crossing her arms foreground-right, exactly as assigned.

Only with the DiT LoRA applied does the output match full’s identities and style, and Kana at foreground-right wears the black tights, same as full. On this P1O0 seed42, it’s less that the 0.6B encoder lacks the information and more that the 4-char LoRA’s DiT side has a strong habit of putting the black tights on whoever stands foreground-right; that habit overrides the mapping base Anima got right.

It’s a single-seed module split, so I can’t say whether this holds on other prompts. But the Qwen and T5 inputs were identical in every condition, and the presence of the DiT LoRA alone flipped the result between dressing the right girl and the wrong one.

Separating the wave difference by crossing Qwen and T5

010 is the cell with just the Kurara–Kei conversation; 110 adds Kana’s wave. In the seed42 010 image, Kurara and Koharu face each other while Kei stays turned toward the camera, looking only loosely part of the conversation. In 110, Kana waves, and Kei clearly turns toward the two. That difference is what I used for the separation.

The conversation text is identical in 010 and 110; the only difference is one sentence near the end of the prompt that has Kana waving. The 010 Qwen hidden state is 394 tokens and the T5 token ID sequence is 470. In 110 they become Qwen 404 and T5 481.

I generate with one cell’s Qwen output combined with the other cell’s T5 output. I’ll call this crossing below. With the 010 and 110 Qwen-side context and T5-side query sequences, I generated the following five combinations.

QwenT5Result
010 control010 controlreproduces the 010 failure
110 factor110 factorreproduces the 110 success
010 control110 factorfollows the factor side
110 factor010 controlfollows the control side
zero110 factorthree-girl hybrid in a broken space
Qwen=010, T5=010. No wave from Kana; Kurara and Koharu face each other and Kei stays toward the camera

Qwen=010, T5=010. No wave from Kana; Kurara and Koharu face each other and Kei stays turned toward the camera.

Qwen=110, T5=110. Kana waves and Kei turns toward the two

Qwen=110, T5=110. Kana waves and Kei clearly turns toward the two.

Qwen=010, T5=110. Same as the factor side: Kana waves and Kei turns toward the two

Qwen=010, T5=110. Same as the 110 side: Kana waves and Kei turns toward the two.

Qwen=110, T5=010. Same as the control side: Kei stays toward the camera and Kana doesn't wave

Qwen=110, T5=010. Same as the 010 side: Kurara and Koharu keep facing each other, Kana doesn’t wave, and Kei stays turned toward the camera.

Qwen=all zeros, T5=110. A three-girl hybrid in a broken, mirror-like space

Qwen=all zeros, T5=110. A three-girl hybrid in a broken, mirror-like space.

Kana’s wave and Kei’s orientation changed when the T5 side switched to 110, and didn’t change when only the Qwen side did. That means the 010-to-110 difference reached the image through the T5 query sequence.

Zeroing Qwen, on the other hand, destroyed the whole image; neither the scene’s meaning nor its space survived. Still, the 010/110 difference had nothing to do with the Qwen side.

The mean-pooled cosine similarity of the two Qwen hidden states was 0.999935, but comparing the 394 shared tokens position by position gave 0.9124. Averages alone make the two look nearly identical, so instead of judging from the hidden-state numbers, I generated an image for each crossed combination and judged by what changed.

I also started the same crossing on seed1234, but the directly executed 010 didn’t match the ComfyUI API version. The API version had four girls and the direct version a three-girl hybrid, so I stopped seed1234 partway. That’s why this crossing is only confirmed on seed42.

Trying the same crossing on the outfit assignment

With 010/110, swapping the T5 side alone flipped the image. Which of Qwen and T5, then, carries the outfit-assignment difference between P1O0 and P1O1?

P1O0 and P1O1 are the same scene after swapping Kana’s and Koharu’s positions. The only difference is whether the skirt and black tights are assigned to Koharu at the back (P1O0) or to Kana at foreground-right (P1O1). The correct output has Koharu wearing the tights in P1O0 and Kana wearing them in P1O1. Both cells have a 435×1024 Qwen hidden state and a 518-token T5 sequence.

I crossed the Qwen-side context and T5-side query sequences on seed42. The direct-execution output doesn’t match the API version pixel for pixel, but the four identities, Koharu waving at the back, Kana crossing her arms foreground-right, and the black tights on Kana all came out the same, so I went ahead with this method.

The black tights were on Kana at foreground-right in all four conditions.

QwenT5Matches the P1O0 assignment
P1O0P1O0×
P1O1P1O1-
P1O0P1O1-
P1O1P1O0×
Qwen=P1O0, T5=P1O0. Both assign the black tights to Koharu at the back, and Kana foreground-right wears them anyway

Qwen=P1O0, T5=P1O0. Both assign the black tights to Koharu at the back, and Kana at foreground-right wears them anyway.

Qwen=P1O1, T5=P1O1. The assignment to Kana foreground-right and the actual wearer match

Qwen=P1O1, T5=P1O1. The assignment (Kana at foreground-right) and the girl actually wearing them match.

Qwen=P1O0, T5=P1O1. The black tights stay on Kana foreground-right

Qwen=P1O0, T5=P1O1. The black tights stay on Kana at foreground-right.

Qwen=P1O1, T5=P1O0. Even with T5 assigning Koharu at the back, the black tights stay on Kana foreground-right

Qwen=P1O1, T5=P1O0. Even with T5 assigning Koharu at the back, the black tights stay on Kana at foreground-right.

Unlike 010/110, swapping only Qwen, only T5, or setting both to P1O0 changed nothing about who wears the outfit. The P1O0/P1O1 Qwen hidden states have a mean-pooled cosine similarity of 0.9999946, and 0.9806485 when the 435 positions are compared one by one. The text difference is present in both the Qwen hidden state and the T5 token ID sequence, yet on this checkpoint the girl wearing the black tights never changed.

As the earlier zero intervention showed, deleting Qwen breaks both the cast and the space, so Qwen itself is necessary for the scene. Even so, for who wears the black tights, it made no difference whether Qwen carried P1O0 or P1O1, or T5 either.

Checking whether a bigger Qwen difference changes who wears them

There’s still the possibility that the Qwen output difference between P1O0 and P1O1 is simply too small. So I fixed T5 to P1O0 and linearly extrapolated the Qwen hidden state toward the P1O0 side: x2 is 2 × Qwen(P1O0) - Qwen(P1O1) and x4 is 4 × Qwen(P1O0) - 3 × Qwen(P1O1).

This forces tensor values that never occur in training, so it doesn’t reproduce what a larger Qwen would output. The main question is whether the girl wearing the black tights changes, but if the image didn’t change at all, the extrapolation itself might simply not be reaching the image. That’s why I also tracked how much the whole image changed.

Qwen difference extrapolated x2 toward P1O0. The scene holds, and the black tights stay on Kana foreground-right

x2. The scene holds, and the black tights stay on Kana at foreground-right.

Qwen difference extrapolated x4 toward P1O0. The image changes enough that Koharu raises both hands, and the black tights stay on Kana foreground-right

x4. The image changes enough that Koharu at the back raises both hands, and the black tights are still on Kana at foreground-right.

Neither x2 nor x4 got Koharu into the black tights. At x4 the composition changed enough that Koharu raises both hands, so the intervention did reach the image. Still, even with the P1O0-to-P1O1 Qwen difference scaled up to 4×, Kana at foreground-right kept wearing the black tights.

The module split, meanwhile, showed base Anima dressing the right girl from the same 0.6B conditioning. If the 0.6B lacked the information, base would get it wrong too. So I think this one comes down to the DiT LoRA overriding the assignment.

Working on the DiT LoRA before reaching for a bigger Qwen

Qwen and T5 emit the same outputs in every condition, yet base Anima and LLMAdapter-only put the black tights on the girl waving at the back as assigned, and only DiT-only dressed Kana at foreground-right the way full did. A larger Qwen might make the mapping more reliable. But this time at least, the encoder wasn’t the one making the mistake.

I deliberately didn’t bind the outfits to trigger words, because I want to dress the characters at inference time. That dress-up failed to reach the right girl here. That leaves two things to try first: rebuilding the training data with swapped roles, or adjusting how the DiT LoRA is applied. Swapping in a larger encoder can wait.