Anima DiT Residual Steering on M1 Max Flips Flat Color, Not Hair Color Alone
Contents

This started from a conversation about the paper that swapped GPT-OSS hidden states for a symbolic-structure formula and got almost the same answers.
Rebuild every layer’s hidden state of an LLM with a closed-form tensor product representation and the model’s answers barely change, and role-rewriting interventions go through too.
The question was whether the same view works for an image-generation DiT.
What grammar and logic are to an LLM’s internal algorithms would map, I figured, to spatial layout, part structure, attribute binding, depth and occlusion, and the division of labor across denoise steps.
Earlier, in the write-up that looked inside the conditioning of a broken 4-character Anima LoRA, I intervened up to the conditioning right before the DiT, but never looked at how the DiT turns that conditioning into image space.
I had only probed the entrance and stopped at “it seems to understand the prompt”.
This time I open up the DiT itself.
The order goes from red and blue flat-color images, to find where color is decided and written, to one girl’s hair, and then to two girls side by side. Everything runs on an M1 Max 64GB with plain Anima-Base v1.0, no LoRA.
Test environment
| Item | Details |
|---|---|
| Machine | Apple M1 Max 64GB, macOS 26.5 |
| Runner | A homemade harness (.image-work/dit-probe/) that imports the comfy package from ComfyUI v0.30.1 (commit 0764232) directly |
| DiT | Anima-Base v1.0 (bf16) |
| Text encoder | Qwen3-0.6B-Base + T5 tokenizer, Anima’s own LLMAdapter |
| LoRA | None (Turbo removed too) |
| Common settings | 30 steps, CFG 4.0, euler, scheduler simple, 1216x832 (token grid 52x76) |
| Negative | worst quality, low quality, blurry, lowres, text, watermark, bad hands |
| Measured | Only the cond side of CFG |
The DiT lives in ComfyUI’s comfy/ldm/cosmos/predict2.py.
Each Block runs self-attention (image tokens to image tokens), then cross-attention (image tokens reading the conditioning), then MLP, with adaLN applying the timestep embedding at each stage.
Text enters the image side only through cross-attention, so in E1 (the cross-attention map experiment listed in the next section) I use the cross-attention weights to measure how strongly each image token reads each text position. Weights alone do not decide the actual contribution, so later sections also check by swapping the conditioning.
An MMDiT like Qwen-Image mixes image and text in the same attention, whereas Anima lets you handle the text-side weights separately.
The model layout follows the Anima-Base model card. The values below are what I measured by running it.
| Item | Measured |
|---|---|
| Blocks | 28 |
| hidden dim | 2048 |
| Heads | 16 (head_dim 128) |
| Token grid (rows x cols) | 52x76 = 3952 image tokens (patch 2, 1216x832 stays even) |
| conditioning shape | (1, 512, 1024). Actual prompts are around 38 tokens, the rest is zero padding |
| Time per image (30 steps) | 254 s (38 s at 4 steps) |
The cond and uncond sides are told apart with transformer_options["cond_or_uncond"].
At this resolution cond and uncond arrive as alternating separate calls, so I measured and modified only the rows marked 0.
The mapping from words to cross-attention key positions came from actually tokenizing with the T5 tokenizer.
The first snag in the harness check was the left/right self-attention block. Passing a (b,1,S,S) boolean mask to SDPA on MPS produced NaN a few Blocks in and a pitch-black image.
Forcing synchronization made it disappear. I never found the cause, but slicing the left and right tokens apart and running the model’s own attention on each side separately (equivalent to a block-diagonal mask) worked, so I went with that.
The other snag was the residual stream. ComfyUI keeps it in fp32 and the values are large (up to around 2.7e5), so saving in fp16 turned a handful of elements in the middle Blocks into inf. E2 (the residual difference experiment) saves in fp32.
To confirm that measuring does not change the image, I compared the output PNGs with and without the attn and resid hooks. Bit-identical. The conditioning swap also matched a plain generation of prompt B bit for bit when swapped across all Blocks and all steps.
Order of experiments
The actual order was the two-girl series (E0-E7) first; the flat-color series (S0-S5) was added afterwards.
| Stage | ID | Name | What is done | What is measured |
|---|---|---|---|---|
| Flat color | S0 | Does a flat color come out | Generate red, blue, green, yellow flat-color prompts at seed 42 | Mean RGB and standard deviation. Is it a flat color |
| Flat color | S1 | The step where color is decided | For red and blue, decode each step’s denoise prediction (x0) with the VAE and take mean RGB | At which step do red and blue separate |
| Flat color | S2 | Residual difference | Run red and blue on the same seed, save all Block, all step outputs, take the difference | Which Block has the largest L2 difference. Are the large-difference dims the same as the huge-value dims |
| Flat color | S3 | Direction transplant | Add the mean blue minus red direction to the red generation. 7 runs over Block groups, 2 runs over all Blocks with limited steps | Does mean RGB move toward blue. Which Block changes |
| Flat color | S3x | Single Block and coefficient | Split the Block group that changed color in S3 one by one, add at steps 0-1 only. One run at coefficient 0.7 | Is one Block enough |
| Flat color | S4 | Cut in space | Add the direction only to the left-half tokens of the Block group that changed color in S3 | Does only the left half turn blue |
| One girl | S5 | Transplant to hair | Add the same direction to the hair tokens of a single blonde girl | Does the hair turn blue. Do face and clothes survive |
| Two girls | E0 | Baseline | Generate two prompts with left/right hair colors swapped, 5 seeds each, no intervention | Are the left/right hair colors as instructed |
| Two girls | E1 | cross-attention map | Map the attention weights on blonde and black back to space for all Blocks and steps | Which Block and step the bias appears in |
| Two girls | E2 | Residual stream difference | Run both prompts on the same seed, take the difference of all Block outputs | Left/right asymmetry of the difference |
| Two girls | E3 | conditioning swap | Swap in the other prompt’s conditioning only for a given Block group or step range | Where the swap flips the hair colors |
| Two girls | E4 | Left/right self-attention block | Cut cross-references between left and right in image-token self-attention | What the block does |
| Two girls | E5, E6 | Token-level rewrite | Swap only the hair-color token positions. Narrow by Block, step, space, mix ratio | Does only the hair color change while the picture survives |
| Two girls | E7 | Rewrite the residual stream directly | Add the hair-color difference vector only to hair tokens | Does only the hair change |
S0 Does a flat color come out
The prompt is solid red background, plain background, no humans, flat color, empty with only the color word changed, four prompts. Negative stays the common one.
All four came out flat on the first try, so I did not iterate on the wording.
| Color | Mean RGB | Std dev (R / G / B) | Flatness |
|---|---|---|---|
| red | 213 / 18 / 30 | 0.61 / 0.55 / 0.46 | 100% |
| blue | 1 / 87 / 187 | 0.50 / 0.41 / 0.57 | 100% |
| green | 62 / 132 / 79 | 0.51 / 0.43 / 0.51 | 100% |
| yellow | 255 / 241 / 73 | 0.01 / 0.52 / 0.76 | 100% |
Flatness is the share of pixels whose three RGB channels all sit within plus or minus 20 of the image mean. Every image spans only 5-10 levels per channel.




T5 tokenization gives 15 tokens for all four, and the color word lands as a single token at position 1, right after solid. One row of the conditioning corresponded to one color word.
S1 The step where color is decided
For each sampling step I decoded the sampler’s denoise prediction (x0) with the VAE and took the mean RGB.
The capture uses ComfyUI’s sampler callback: save the denoised tensor handed over for previews as is, then decode them all after sampling.
The DiT is untouched, so the final image matched the normal gen output bit for bit (checked with SHA-256). Thirty decodes took 110 s, 3.7 s each.
Besides red and blue I also did G, the single-girl image used in S5 (1girl, upper body, looking at viewer, simple background, blonde hair, long hair).
| step | red | blue | G |
|---|---|---|---|
| 0 | 255 / 66 / 59 | 5 / 133 / 255 | 255 / 254 / 250 |
| 1 | 220 / 17 / 28 | 3 / 87 / 179 | 239 / 212 / 185 |
| 2 | 210 / 24 / 35 | 4 / 89 / 180 | 244 / 225 / 199 |
| 3 | 213 / 24 / 33 | 2 / 89 / 183 | 248 / 228 / 206 |
| 5 | 213 / 20 / 31 | 1 / 88 / 185 | 248 / 225 / 207 |
| 10 | 213 / 17 / 30 | 1 / 87 / 186 | 247 / 223 / 204 |
| 15 | 213 / 17 / 30 | 1 / 87 / 186 | 247 / 222 / 203 |
| 29 | 213 / 18 / 30 | 1 / 87 / 187 | 246 / 222 / 203 |



In all three, every channel is within plus or minus 30 of the final color by step 1.
The step 0 x0 has one channel pinned at 255 and misses, but by the second prediction the color is already set and the values only tighten from there. Red’s flatness is 51% at step 0, 96% at step 1, 100% at step 2.
G’s mean color was also settled by step 1. The shape, on the other hand, is a blurry blob at step 1, becomes recognizably a person at step 4, and details like eyes and collar only show up from around step 10.
Color seems to be decided in the first 2 steps. From S3 on, whenever I add a direction, there is a condition that adds it at steps 0-1 only, alongside the all-steps condition.
S2 Residual difference
I ran red and blue on the same seed 42 and saved the outputs of all 28 Blocks at all 30 steps in fp32 (25 GB per run).
For each step and Block I took the difference of the outputs and averaged over all tokens to get a vector d = blue minus red.
What I wanted to check was whether the difference is concentrated in a few dims, and if so whether those are the same huge-value dims that E2 later finds. So I computed the norm of d, its ratio to the norm of red’s residual, the share of the norm held by the top 20 dims of |d|, and how many of those 20 overlap with the 20 largest dims of red’s own residual.
The table shows steps 1 and 15, every fourth Block.
| step | Block | norm of d | residual norm | ratio | share of top 20 dims | overlap with huge-value dims |
|---|---|---|---|---|---|---|
| 1 | 0 | 5 | 30 | 0.18 | 0.93 | 0 |
| 1 | 4 | 1038 | 8309 | 0.13 | 0.88 | 0 |
| 1 | 8 | 1242 | 8703 | 0.14 | 0.87 | 2 |
| 1 | 12 | 1239 | 10630 | 0.12 | 0.88 | 0 |
| 1 | 16 | 1250 | 11960 | 0.11 | 0.88 | 0 |
| 1 | 20 | 1377 | 13700 | 0.10 | 0.86 | 0 |
| 1 | 24 | 3054 | 14220 | 0.22 | 0.89 | 2 |
| 1 | 27 | 12030 | 61820 | 0.19 | 0.58 | 1 |
| 15 | 0 | 27 | 57 | 0.48 | 0.90 | 1 |
| 15 | 4 | 3271 | 8614 | 0.38 | 0.94 | 1 |
| 15 | 8 | 3477 | 8793 | 0.40 | 0.93 | 1 |
| 15 | 12 | 3511 | 10600 | 0.33 | 0.92 | 2 |
| 15 | 16 | 3527 | 11930 | 0.30 | 0.92 | 2 |
| 15 | 20 | 3583 | 13610 | 0.26 | 0.92 | 1 |
| 15 | 24 | 4925 | 15020 | 0.33 | 0.91 | 4 |
| 15 | 27 | 11460 | 63520 | 0.18 | 0.65 | 4 |



The red-blue difference was 85-94% concentrated in 20 of the 2048 dims. Only Blocks 26-27 drop to 58-65%. Same shape as E7 later, where the hair-color difference is 70-80% concentrated in 20 dims too.
Those 20 dims barely overlap with the huge-value dims of the residual itself (0-2 in the middle Blocks, at most 4 in Blocks 24 and 27). Back in E7 of the two-girl series I had assumed they were the same dims. They were not. The color difference mostly lives in ordinary-sized dims.
Spatially the difference is uniform. In Blocks 0-11 and 22-27 the spread across tokens (coefficient of variation) is 0.02-0.17.
Only in Blocks 12-21 do 62 of the 3952 tokens (1.6%) show a difference more than 10 times the median, at the same positions across Blocks 12-21, and it vanished at Block 22.
Block 0’s difference is 5-27, essentially zero next to the 1000-plus from Block 1 onward. For these two prompts, Block 0’s output seems to barely reflect the prompt.
S3 Direction transplant
I added d from S2 to the red generation. Region: all tokens, mean direction, coefficient 1.0. Seven runs adding it at all steps to each of seven Block groups, plus two runs over all Blocks with a limited step range.
Projection is the mean RGB projected onto the red-to-blue line: 0 is red, 1 is blue.
| ID | Block | step | Mean RGB | Flatness | Projection | Appearance |
|---|---|---|---|---|---|---|
| S3-1 | 0-3 | all | 255 / 40 / 1 | 43% | -0.16 | Saturated orange-red with yellow grain |
| S3-2 | 4-7 | all | 254 / 21 / 1 | 34% | -0.18 | Saturated red with fine grain |
| S3-3 | 8-11 | all | 0 / 0 / 0 | - | - | All black (residual NaN) |
| S3-4 | 12-15 | all | 0 / 0 / 0 | - | - | All black (residual NaN) |
| S3-5 | 16-19 | all | 0 / 0 / 0 | - | - | All black (residual NaN) |
| S3-6 | 20-23 | all | 0 / 0 / 0 | - | - | All black (residual NaN) |
| S3-7 | 24-27 | all | 0 / 162 / 255 | 97% | 1.22 | Flat cyan-leaning blue. Faint token-grid weave |
| S3-8 | 0-27 | 0-1 | 0 / 0 / 0 | - | - | All black (residual NaN) |
| S3-9 | 0-27 | 2-29 | 0 / 0 / 0 | - | - | All black (residual NaN) |
For the two saturated runs I also tried a version with the top 20 dims zeroed.
| ID | Block | step | Mean RGB | Flatness | Projection | Appearance |
|---|---|---|---|---|---|---|
| S3-1c | 0-3 | all | 236 / 234 / 243 | 61% | 0.59 | Near-white pale purple, grainy |
| S3-2c | 4-7 | all | 167 / 223 / 255 | 30% | 0.80 | Light blue, strong diagonal weave |











It turned blue only when added to Blocks 24-27. And it overshot blue into cyan (projection 1.22).
Blocks 8-23 went NaN and all black when four Blocks were fed at once. Adding to all Blocks with a narrowed step range did the same. Blocks 0-7 stayed red and saturated.
The difference vector shifts the input of the next Block at every Block it touches, so adding to four at once stacks the effect. The NaN may be from that, but this measurement cannot tell.
With the top 20 dims zeroed, Blocks 0-7 also moved toward blue, but flatness fell to 30-60% and a weave pattern appeared.
S3 follow-up Single Block and coefficient
Since S1 showed color settled by step 1, I split Blocks 24-27 one by one and added at steps 0-1 only. One more run with all four at coefficient 0.7.
| ID | Block | step | Coefficient | Mean RGB | Flatness | Projection |
|---|---|---|---|---|---|---|
| S3x-1 | 24 | 0-1 | 1.0 | 0 / 142 / 255 | 100% | 1.20 |
| S3x-2 | 25 | 0-1 | 1.0 | 0 / 142 / 255 | 100% | 1.20 |
| S3x-3 | 26 | 0-1 | 1.0 | 0 / 136 / 255 | 100% | 1.19 |
| S3x-4 | 27 | 0-1 | 1.0 | 0 / 137 / 255 | 100% | 1.19 |
| S3x-5 | 24-27 | 0-1 | 0.7 | 0 / 175 / 255 | 100% | 1.23 |





One Block, steps 0-1 only, and red went fully blue. Any of 24-27 does it, and flatness stays at 100%.
The overshoot into cyan past projection 1 is the same whether all four or a single one, and dropping the coefficient to 0.7 still gave 1.23. Only two points, but the color does not look like it scales linearly with the size of the direction.
S4 Cut in space
I added d to Blocks 24-27 only on the left-half tokens (columns 0-37). One version at all steps, one at steps 0-1 only.
| ID | step | Left-half mean RGB | Right-half mean RGB | Boundary |
|---|---|---|---|---|
| S4-1 | all | 0 / 161 / 255 (projection 1.22) | 188 / 170 / 131 (projection 0.43) | Color gap 141 at x=631. Jagged along token borders |
| S4-2 | 0-1 | 0 / 162 / 255 (projection 1.22) | 171 / 6 / 9 (projection 0.07) | Color gap 327 at x=629. One pixel wide |


The steps 0-1 version has the left blue, the right still red, and the boundary is a one-pixel-wide straight line.
The all-steps version turns the left blue, but the right collapses into beige and the boundary becomes jagged along the token borders.
The boundary sits at pixels 629-631, 1.3-1.4 tokens to the right of pixel 608, where the column 38 token border should be. I assume that is the VAE’s receptive field (the surrounding area consulted to reconstruct one pixel).
Under these conditions, adding the flat-color difference direction to the last 4 Blocks at steps 0-1 changed the color on one side only.
Whether color is actually written in those Blocks during an unmodified generation is something I did not check. Adding to the middle Blocks (8-23) blows up the residual, so the same form cannot be added there.
S5 Transplant to hair
I added the direction obtained from flat colors to the hair tokens of G, a single blonde girl.
The hair is already yellowish, so the direction is blue minus yellow, and yellow’s residual was saved only for Blocks 24-27 at steps 0-1. I added one run with blue minus red for comparison.
G’s hair mask takes blonde pixels from plain G (R>150, G>110, B<160, R-B>40) and marks a token as hair at 25% or more coverage. 1054 of the 3952 tokens. The B ceiling is 160 instead of 140 to catch the highlights at the crown, which lets 2-3 tokens of neck shadow in.

Plain G’s mean RGB inside the hair is 235 / 190 / 118.
| ID | Direction | Block | step | Region | Pixel diff vs plain G (all / in hair / outside hair) | Mean RGB in hair |
|---|---|---|---|---|---|---|
| S5-1 | blue - yellow | 24-27 | 0-1 | hair | 94 / 185 / 61 | 0 / 3 / 252 |
| S5-2 | blue - red | 24-27 | 0-1 | hair | 68 / 137 / 43 | 0 / 177 / 256 |
| S5-3 | blue - yellow | 24-27 | 0-1 | all | 165 / 177 / 161 | 0 / 25 / 248 |
| S5-4 | blue - yellow | 24-27 | all | hair | 223 / 181 / 239 | 0 / 0 / 0 |




S5-1, added to the hair tokens, turned the hair mask’s exact shape into a flat blue slab.
The slab has a smooth outline with a black edge, the girl is redrawn small in the empty space below it, her clothes stay red, and the rendering flattens into line-art style.
S5-2 with blue minus red broke the same way, in light blue.
S5-3, added to all tokens, went entirely blue and the girl disappeared. S5-4, added at all steps, blew up the residual and went black.
The direction obtained from flat colors seems to be a direction that turns the region it is added to into a flat blue surface.
Added to the hair region, it drew that region as a blue object and the girl was redrawn in whatever space was left.
Steps 0-1 decide composition as well as color (G’s shape is not fixed until step 4 in S1), so I think adding a surface instruction there is what changed the composition along with it.
The two-girl hair-color series
Chronologically this is the experiment I did before the flat colors. Two girls side by side, blonde on the left and black hair on the right, and I tracked the binding of attribute to position from conditioning swaps through to editing the residual stream directly.
Why no LoRA
With a character LoRA applied, there is no telling whether a structure I find belongs to Anima or to that LoRA. In the 4-character LoRA write-up above I also confirmed that appearance gets absorbed into the trigger depending on how captions are written, so LoRA influence would get mixed in.
With plain Anima and ordinary attributes like blonde hair and black hair, each attribute maps to 1-2 tokens on the text side, which makes it easier to check which patches the word’s cross-attention weight landed on. For the same reason I removed the Turbo LoRA too and ran normal steps with CFG on.
Prompts
The base is two prompts of two girls side by side with only the left/right hair colors swapped.
| Label | Prompt |
|---|---|
| A | 2girls, standing side by side, upper body, looking at viewer, simple background. The girl on the left has blonde hair. The girl on the right has black hair. |
| B | 2girls, standing side by side, upper body, looking at viewer, simple background. The girl on the left has black hair. The girl on the right has blonde hair. |
A and B share the same set of words. Only the binding differs.
The encoder-side representations differ too, but mapping the cross-attention weights back to space per word should show where the DiT turns that difference into position.
E0 Baseline
A and B generated plain at seeds 42-46. All 10 images have two girls, with the left/right hair colors as instructed.
Judgment was done by opening each image and writing down just the head count and left/right hair colors as seen.
| Prompt | seed | Left hair | Right hair |
|---|---|---|---|
| A | 42 | blonde | black |
| B | 42 | black | blonde |
| A | 43 | blonde (highly saturated yellow) | black |
| B | 43 | black | blonde (highly saturated yellow) |
| A | 44 | blonde | black |
| B | 44 | black | blonde |
| A | 45 | blonde | black |
| B | 45 | black | blonde |
| A | 46 | blonde | black |
| B | 46 | black (bluish, near navy) | blonde |










Clothing color was never specified, but seed 42 A gives the left girl whitish clothes and the right girl black, and seed 46 B gives a navy jacket on the left and a red sweater on the right.
I had planned to use a seed with swapped hair colors and positions in E4, but all 10 followed the instruction and none appeared. Plain Anima apparently does not break on this prompt.
E4 adds clothing colors to the hair colors and first checks whether that departs from the instruction.
E1 cross-attention map
I regenerated A at seed 42, which followed the instruction in E0, under the same settings while recording the cross-attention weights of all Blocks and all steps.
The recording is computed separately from copies of q and k, so the output image is identical to the unrecorded E0 (bit-identity confirmed in the harness check).
For each Block and step I took the attention weight that each of the 3952 image tokens places on blonde (T5 token position 25) and black (position 34), averaged over heads, and put it back on the 52x76 grid.
Then I defined the left-half share as the sum of weight landing in the left half of the grid divided by the total. 0.5 means no left/right bias.


The numbers look like this.
| Block | blonde step 0 | blonde step 10 | blonde step 29 | black step 0 | black step 10 | black step 29 |
|---|---|---|---|---|---|---|
| 0 | 0.50 | 0.50 | 0.50 | 0.50 | 0.50 | 0.48 |
| 2 | 0.54 | 0.63 | 0.58 | 0.46 | 0.32 | 0.44 |
| 5 | 0.67 | 0.81 | 0.57 | 0.31 | 0.25 | 0.46 |
| 9 | 0.91 | 0.90 | 0.64 | 0.09 | 0.12 | 0.38 |
| 13 | 0.90 | 0.84 | 0.66 | 0.10 | 0.12 | 0.31 |
| 16 | 0.81 | 0.75 | 0.60 | 0.19 | 0.20 | 0.39 |
| 19 | 0.61 | 0.63 | 0.52 | 0.41 | 0.39 | 0.48 |
| 20 | 0.49 | 0.49 | 0.49 | 0.49 | 0.49 | 0.50 |
| 27 | 0.50 | 0.50 | 0.50 | 0.50 | 0.50 | 0.50 |
Blocks 0-1 and Blocks 20-27 sit around 0.50 for both words at every step, with no left/right bias.
The bias starts around Block 2 and is strongest in Blocks 8-16. At step 0, Block 9 has blonde at 0.91 and black at 0.09, almost mirror images.
The bias weakens as steps progress. Block 9’s blonde fell from 0.91 at step 0 to 0.64 at step 29.
The left-half share only tells direction, so I also computed the mean weight over the whole grid (per token). The conditioning has 512 tokens, so uniform attention would give 1/512, about 0.00195, per token. The table is scaled by 1000.
| Block | blonde step 0 | blonde step 10 | blonde step 29 | black step 0 | black step 10 | black step 29 |
|---|---|---|---|---|---|---|
| 0 | 1.0 | 0.9 | 1.0 | 0.9 | 0.9 | 0.9 |
| 5 | 2.3 | 1.6 | 0.3 | 1.7 | 0.9 | 0.3 |
| 9 | 6.2 | 2.3 | 0.4 | 7.7 | 2.2 | 0.3 |
| 13 | 3.6 | 1.1 | 0.2 | 2.9 | 1.2 | 0.2 |
| 19 | 5.6 | 2.1 | 0.8 | 5.9 | 2.0 | 0.7 |
| 20 | 0.9 | 0.9 | 1.0 | 0.9 | 0.9 | 1.0 |
| 27 | 0.9 | 0.9 | 1.0 | 1.0 | 1.0 | 1.1 |
At step 0, Block 9 puts 3-4 times the uniform weight on the hair-color words, and by step 29 that same Block has dropped below one fifth of uniform.
Blocks 0-1 and 20-27, on the other hand, stay at roughly half of uniform from start to finish and barely move.
The 0.50 in the later Blocks seems to be the value you get when these two words are hardly read at all. What they read instead is not something this measurement captures.
For the spatial maps I cut out Blocks 4, 8, 12, 16, 20 at steps 0, 5, 10, 20, 29 and normalized each panel by its own maximum (brightness is not comparable across panels; the figure is for shape only).



In the blonde figure, Block 4 brightens along the shape of the left girl’s hair from step 5 on, and Block 8 also shows the left girl’s outline from step 5. Blocks 12 and 16 concentrate the weight on a few points with no surface-like shape, and Block 20 is nearly flat across the grid.
Block 4’s left-half share is around 0.6, a small bias in the table, yet the spatial map shows the left girl’s hair shape. Weight also spreads faintly onto the right girl, which is what pulls the share down.
What token positions the point-like concentrations in Blocks 12 and 16 correspond to, this figure cannot tell.
E2 Residual stream difference
I generated A and B at the same seed 42 and saved every Block’s output (the residual stream, shape 1x52x76x2048) at steps 0, 1, 2, 5, 10, 15, 20, 29 in fp32. 6.8 GB per run.
I subtracted A’s and B’s outputs at the same step and Block, took the L2 norm per token, and put it back on the 52x76 grid.
At step 0 the initial latent is identical, so the difference comes purely from the conditioning. From step 1 on the latent trajectories also diverge and that gets mixed in.



The token-mean L2 at step 0 grew with Block depth.
| Block | L2 mean at step 0 |
|---|---|
| 0 | 0.08 |
| 1 | 18 |
| 4 | 27 |
| 8 | 56 |
| 9 | 116 |
| 12 | 330 |
| 16 | 542 |
| 20 | 546 |
| 24 | 613 |
| 27 | 1416 |
Block 0’s output is nearly the same for A and B, then it jumps to 18 at Block 1. It rose sharply from 170 to 330 between Blocks 11 and 12, and from 946 to 1416 between Blocks 26 and 27.
The size of the residual stream itself presumably differs per Block too, and here I only show the absolute difference.
In the spatial maps, the step 0 difference in Blocks 0-8 is grid-wide noise with no shape, while Blocks 24 and 27 already show two blobs at the two heads at that same step 0.
By steps 1-2, Blocks 0-8 also show blobs at the two heads, and from step 5 on the difference is simply the shape of the two girls’ hair and outlines.
At every Block and step the difference appears on both girls; no panel concentrates it on one side. A and B swap both hair colors, so the difference showed up on both sides exactly as the swap implies.
In Blocks 12, 16, 20, only a few tokens near the top edge of the grid carry values of 100,000-200,000 and the rest goes dark. These are the tokens that exceeded the fp16 range in the harness check, and presumably the reason ComfyUI keeps the residual stream in fp32.
Left/right asymmetry is the left-half mean minus the right-half mean, divided by the overall mean, and it stayed within plus or minus 0.4 at every Block and step.
Block 9 at step 0 is -0.38, right-leaning; Blocks 12-20 at steps 5-20 are -0.16 to -0.37, right-leaning; the same Blocks at step 2 are +0.11 to +0.22, left-leaning; and step 29 is near 0 everywhere.
The Blocks 12-20 values include the top-edge outlier tokens described above.
The difference alone cannot say which Block decides the binding, so E3 checks by swapping.
E3 conditioning swap
While generating A, I swapped in B’s conditioning for cross-attention only in the specified range. Seed 42, cond side only.
A swap across all Blocks and all steps matches plain B bit for bit, confirmed in the harness check.
Judgment was done by opening each image and writing down the hair-color arrangement as seen, with the mean absolute pixel difference (RGB, 0-255) against plain A (E0’s seed 42) and plain B as a supporting number. The pixel difference between plain A and plain B is 65.49.
E3a Block axis
Only a Block group is swapped, across all steps 0-29.
| Swapped range | Hair arrangement | Changes besides hair color | Pixel diff vs plain A | Pixel diff vs plain B |
|---|---|---|---|---|
| Blocks 0-3 | Still A | Both girls’ eyes turned black (A’s left is blue-green, right is brown) | 5.98 | 65.14 |
| Blocks 4-7 | Still A | Collar became a vertically ribbed stand collar with buttons, faces got longer | 11.70 | 62.52 |
| Blocks 8-11 | Still A | Tighter face crop with large black eyes, left top became a pale purple button shirt | 17.04 | 62.49 |
| Blocks 12-15 | Flipped to B | Both in black sleeveless turtlenecks with bare shoulders, right hair a short bob, tight upper-body crop | 59.04 | 38.06 |
| Blocks 16-19 | Still A | Collar changed from frills to a buckled high collar | 6.48 | 64.28 |
| Blocks 20-23 | Still A | Metal detail added at the collar | 4.23 | 65.87 |
| Blocks 24-27 | Still A | Almost the same | 3.11 | 65.63 |
| Blocks 0-13 | Still A | Both in sleeveless ribbed turtlenecks with bare shoulders, right hair a shoulder-length bob, left top gray-purple, tight crop | 30.76 | 59.40 |
| Blocks 14-27 | Flipped to B | Both in black sleeveless turtlenecks with bare shoulders, right hair a short bob, tight crop | 59.77 | 38.35 |









The hair arrangement flipped in two cases, Blocks 12-15 and Blocks 14-27. Both sit at a pixel difference of 59 from plain A and 38 from plain B, a different picture that is neither A nor B. The two flipped images also changed clothes, hairstyle, and composition the same way.
Swapping Blocks 0-13 did not flip, yet clothes and composition changed the same way as the two flipped ones (pixel difference 30.76).
That is the only run that includes Blocks 12-13 without flipping, so whether 14-15 are included out of Blocks 12-15 looks like the dividing line. But there are only two ways the Block groups were cut, and I did not try swapping Blocks 14-15 alone.
Swapping Blocks 0-3, 16-19, 20-23, 24-27 left the hair arrangement untouched with pixel differences of around 3-6. Only details like eye color and collar trim changed.
Blocks 4-7 and 8-11 have pixel differences of 12-17, with changes in face crop and tops, but the hair arrangement stays A.
In E1 the strongest left/right bias was in Blocks 8-16, and in E3a the flip came from swapping Blocks 12-15 out of those.
Swapping Blocks 8-11, where E1’s bias is also strong, did not flip, so in this experiment the strength of the cross-attention bias and whether swapping that Block flips the hair colors do not line up.
E3b step axis
Only a step range is swapped, across all Blocks.
| Swapped range | Hair arrangement | Changes besides hair color | Pixel diff vs plain A | Pixel diff vs plain B |
|---|---|---|---|---|
| steps 0-4 | Flipped to B | Same flat line-art as B, brown cardigan with pale purple frilled collar, right eye gray. Slightly wider shoulder framing than B | 65.34 | 10.11 |
| steps 5-9 | Still A | Collar changed to a buckled high collar | 5.77 | 66.15 |
| steps 10-19 | Still A | Almost the same | 2.23 | 65.56 |
| steps 20-29 | Still A | Almost the same | 0.75 | 65.46 |
| steps 0-9 | Flipped to B | Nearly identical to plain B | 65.60 | 1.90 |
| steps 10-29 | Still A | Almost the same | 2.29 | 65.55 |






Swapping only steps 0-4 to B pulled the hair arrangement, clothes, and line-art texture all toward plain B (pixel difference 10.11 from plain B). Steps 0-9 is nearly identical to plain B (1.90).
Conversely, swapping any of steps 5-9, 10-19, 20-29, or 10-29 stays close to plain A, with pixel differences of 3 or less apart from the collar change at steps 5-9.
Steps 5-9 is the same length as steps 0-4, yet the result stayed plain A.
This fits the trend measured in E1, where the cross-attention bias weakens with steps and the hair-color words are hardly read in later steps. Still, E1’s bias had not fully dropped to 0.5 even at the last step, so the later-step reading is not exactly zero either.
E4 Left/right self-attention block
E0 produced no seed where the hair-color binding broke, so I first added clothing colors to the hair colors as prompt C and generated 3 seeds plain.
| Label | Prompt |
|---|---|
| C | 2girls, standing side by side, upper body, looking at viewer, simple background. The girl on the left has blonde hair and wears a red dress. The girl on the right has black hair and wears a blue dress. |
| seed | Count | Left hair | Left clothes | Right hair | Right clothes | Matches |
|---|---|---|---|---|---|---|
| 42 | 2 | blonde | red | black | navy | 4/4 |
| 43 | 2 | blonde | red | black | blue | 4/4 |
| 44 | 2 | blonde | red | black | blue | 4/4 |



All three follow the instruction on all four items. Seed 42’s right dress is navy rather than blue, but I counted it as a match in the blue family.
Nothing broke under this condition either, so there is no way to compare whether the block improves binding.
Even so, the block would still show how the picture changes, so I put the self-attention mask in on the same 3 seeds.
The mask makes the image-token self-attention read only the same side, left or right. Cross-attention and MLP are unchanged. The implementation is the left/right split attention described in the test environment.
In all 6 images the leftmost girl is blonde in red and the rightmost has black hair; only the rightmost dress splits between navy and blue.
| Mask range | seed | Count | Rightmost dress | Difference from plain | Pixel diff vs plain |
|---|---|---|---|---|---|
| All Blocks | 42 | 4 | navy | Vertical seam down the center. Left half is flat line-art on white, right half is shaded painting, a different style. The two on the left have arms and torsos left white, arms cut off midway | 86.75 |
| All Blocks | 43 | 4 | navy | Center seam. Left half shaded, right half flat. The two on the right have no arms, just torso silhouettes, with thin necks not joined to the bodies | 99.88 |
| All Blocks | 44 | 4 | navy | Center seam. Left half white background, right half gray background with thick lines, a different style. The right girl on the right half has blue-green mixed into black hair | 88.94 |
| Blocks 8-16 | 42 | 4 | blue | No seam, continuous style. Two girls became four, the blonde-red and black-blue pair repeated twice | 78.25 |
| Blocks 8-16 | 43 | 4 | blue | No seam. Four girls, same pair twice. The second from left has a skin-colored chest not joined to her clothes | 79.19 |
| Blocks 8-16 | 44 | 3 | blue | No seam. Three girls; the middle one has blonde bangs and black at the back of the head, in blue | 71.60 |






Blocking on all Blocks put a vertical seam down the center of all three images, and the left and right halves became independent pictures.
Each half contains both the blonde-red and the black-blue girl, four in total. Style and background brightness differ between the halves.
When the halves cannot see each other, each half apparently draws the whole prompt on its own.
Blocking only Blocks 8-16 removes the seam and the style is continuous as one picture, but the head count rose to 4, 4, and 3. Seed 44 has one girl whose hair splits into blonde and black.
E5 Token-level rewrite
E3 swapped the whole conditioning, so more than hair color changed.
This time I swapped only the token vectors for blonde (position 25) and black (position 34) in the conditioning with those from another prompt. Besides B with the swap, I prepared R with red hair on the left.
| Label | Prompt | Source positions |
|---|---|---|
| B | A with left/right hair colors swapped | position 25 = black, position 34 = blonde |
| R | A with blonde replaced by red (The girl on the left has red hair.) | position 25 = red |
| ID | Source | Positions swapped | Range |
|---|---|---|---|
| E5-1 | B | 25 and 34 | all Blocks, all steps |
| E5-2 | B | 25 and 34 | Blocks 12-15, all steps |
| E5-3 | B | 25 and 34 | all Blocks, steps 0-4 |
| E5-4 | R | 25 | all Blocks, all steps |
| E5-5 | R | 25 | Blocks 12-15, all steps |
| E5-6 | R | 25 | all Blocks, steps 0-4 |
Seed 42, everything else as in E0. I added one plain generation of R for reference.
A, B, and R tokenize to 38 tokens with T5 and line up exactly: position 25 is blonde, black, red respectively, and position 34 is black, blonde, black.
The swap happens at the same stage as E3 (the 1x512x1024 conditioning reaching the DiT Blocks’ cross-attention), replacing only the specified rows.
Specifying all 512 positions matched E3’s whole swap bit for bit.
| ID | Source | Positions | Range | Left hair | Right hair | Closest reference | Changes besides hair color (vs plain A) |
|---|---|---|---|---|---|---|---|
| E5-1 | B | 25 and 34 | all Blocks, all steps | black | blonde | plain B | Pale purple frilled high collar, brown buttoned jacket, eye color and line weight all match plain B |
| E5-2 | B | 25 and 34 | Blocks 12-15 | black | blonde (bob) | plain A (style, composition, skin). Clothes differ | Black sleeveless ribbed turtleneck with shoulders and arms bare, no frilled collar, right in a bob |
| E5-3 | B | 25 and 34 | steps 0-4 | black | blonde | plain B | Nearly the same as plain B. Right hair slightly longer than plain B |
| E5-4 | R | 25 | all Blocks, all steps | red | black (bob) | plain R | Gray skin, high-contrast style in only red, black, and white, red and black high-neck tops. Pixel diff vs plain R 4.3 |
| E5-5 | R | 25 | Blocks 12-15 | orange-leaning light brown | black | plain A (style, composition, skin). Clothes differ | Dusty pink and black buttoned high necks, no frilled collar, eyes brown to black. Pixel diff vs plain R 39.7 |
| E5-6 | R | 25 | steps 0-4 | red | black (bob) | plain R | Nearly the same as plain R (pixel diff 5.4). Purple and brown patches at the left neckline |
| plain R | none | none | none | red | black (bob) | (reference) | Gray skin, high-contrast red, black, and white, red and black high-neck tops |
| ID | plain A overall | top left | top right | bottom left | bottom right | plain B overall |
|---|---|---|---|---|---|---|
| E5-1 | 64.3 | 64.7 | 65.6 | 69.0 | 58.0 | 9.5 |
| E5-2 | 57.7 | 48.9 | 55.5 | 77.6 | 48.8 | 39.5 |
| E5-3 | 64.2 | 64.3 | 65.7 | 68.8 | 57.9 | 12.3 |
| E5-4 | 48.3 | 48.9 | 25.1 | 76.5 | 42.5 | 65.6 |
| E5-5 | 18.7 | 17.3 | 12.3 | 32.7 | 12.4 | 57.7 |
| E5-6 | 47.6 | 48.7 | 25.0 | 73.7 | 43.2 | 66.0 |
| plain R | 47.7 | 49.1 | 25.7 | 76.0 | 39.9 | 65.9 |
Pixel differences are mean absolute differences vs plain A (RGB, 0-255), and the four quadrants split the frame in half both ways. Hair falls in the top row, clothes in the bottom row.







Swapping just two tokens to B across all Blocks and all steps turned the whole picture into plain B (pixel difference 9.5 from plain B). Steps 0-4 alone does the same (12.3). At seed 42, the difference between the A and B pictures was almost entirely decided by these two vectors out of 512.
Even with the positions narrowed, using the full range leaves nothing of the original A: clothes, collar, and line texture all switched to B’s along with the hair color.
Narrowing to Blocks 12-15, the hair color changed according to the source (flipped for B, orange-leaning for R) while face, composition, and style stayed plain A. The pixel difference vs plain A is 18.7 for R; by quadrant, the top row where the hair is gives 17 and 12, and the bottom left where the clothes are gives 33.
But those clothes are neither A’s, B’s, nor R’s. All three Blocks 12-15 swaps (E3a’s whole conditioning, E5-2, E5-5) lost the frilled collar and got a turtleneck or high neck instead.
Swapping R’s position 25 in all Blocks gave nearly plain R (4.3), and there the skin turned gray as well and the style became high-contrast red, black, and white. The single red token seems to change the color design of the whole frame.
With R limited to Blocks 12-15, the hair never reached a vivid red and stopped at an orange-leaning light brown.
E6 Single Block, mix ratio, spatial limit
E5 could not isolate hair color by cutting on either Block or step, so I tried the remaining conditions.
| ID | Axis | What | Runs |
|---|---|---|---|
| E6-a | Single Block | R’s position 25, and B’s positions 25 and 34, each put into one of Blocks 12, 13, 14, 15 | 8 |
| E6-b | Mix ratio | R’s position 25 in all Blocks at all steps, mixed with the original vector at ratios 0.25 and 0.5 | 2 |
| E6-c | Spatial limit | The swapped token affects cross-attention only for a specified image-token region. Outside the region, the output computed with the original conditioning is used | 4 |
E6-c regions are defined on the 52x76 token grid. “Left half” is columns 0-37; “top left” is columns 0-37 and rows 0-29 (the area holding the left girl’s hair in E1’s spatial map).
| ID | Source | Position and region | Range |
|---|---|---|---|
| E6-c1 | R | 25 in the left half | all Blocks, all steps |
| E6-c2 | R | 25 in the top left | all Blocks, all steps |
| E6-c3 | R | 25 in the left half | Blocks 12-15 |
| E6-c4 | B | 25 in the left half, 34 in the right half | all Blocks, all steps |
Seed 42, everything else as in E0. Judgment and pixel differences as in E5, with the overall difference vs plain R (E5’s reference) added.
The spatial limit is implemented by computing cross-attention twice at any Block and step that has region-tagged tokens, once with the unswapped conditioning and once with the swapped one, and picking per image token depending on whether it falls in the region.
Tokens outside the region use the unswapped output as is. Self-attention and MLP are unchanged.
Mix ratios 1.0 and 0.0 matched the ordinary swap and plain A bit for bit, respectively.
| ID | Swap | Range | Left hair | Right hair | Closest reference | Changes besides hair color (vs plain A) |
|---|---|---|---|---|---|---|
| E6-a | R 25 | Block 12 | blonde (slightly ochre) | black | plain A | Eyes turned brown, collar a vertically ribbed high neck with a button row (same on both) |
| E6-a | R 25 | Block 13 | blonde | black | plain A | Eyes unchanged. Collar a buttoned stand collar (same on both) |
| E6-a | R 25 | Block 14 | blonde (orange-leaning yellow) | black | plain A | Eyes brown, ribbed high neck with a button row, the right girl shifted left with shoulders touching |
| E6-a | R 25 | Block 15 | blonde | black | plain A | Eyes unchanged. Collar a buttoned stand collar |
| E6-a | B 25,34 | Block 12 | blonde | black | plain A | Eyes black, left top white-to-gray vertical stripes with a gray frilled collar |
| E6-a | B 25,34 | Block 13 | blonde | black | plain A | Almost the same |
| E6-a | B 25,34 | Block 14 | blonde (dull pale yellow) | black | plain A | Eyes dark brown, left top gray-purple vertical rib with a button row, the right girl shifted left with shoulders touching |
| E6-a | B 25,34 | Block 15 | blonde | black | plain A | Eyes black. Clothes almost the same |
| E6-b | R 25, ratio 0.25 | all Blocks, all steps | blonde (orange-leaning yellow) | black | plain A | Eyes brown, high neck with a button row, the right girl shifted left |
| E6-b | R 25, ratio 0.5 | all Blocks, all steps | orange (between blonde and red) | black | composition plain A, clothes and rendering toward plain R | Red on the left, black on the right, high necks, dull gray-brown skin, thick lines and flat fill |
| E6-c1 | R 25 in left half | all Blocks, all steps | red | black | plain R | Whole frame in a different style. Gray skin, high-contrast red, black, and white, right in a bob |
| E6-c2 | R 25 in top left | all Blocks, all steps | red | black | plain R | Whole frame in a different style. Left in a red V-neck, right in a thick ribbed turtleneck, right in a bob, rough lines |
| E6-c3 | R 25 in left half | Blocks 12-15 | orange (bright) | black | plain A | Eyes brown, left pale red-brown and right black high necks with a button row, the right girl shifted left |
| E6-c4 | B 25 in left half, 34 in right half | all Blocks, all steps | black | blonde | plain B | Whole frame in a different style. Flat fill with thick black lines, brown cardigan with pale purple frilled collar, right in a bob |
| ID | plain A overall | top left | top right | bottom left | bottom right | plain B overall | plain R overall |
|---|---|---|---|---|---|---|---|
| E6-a R 25 Block 12 | 12.3 | 8.7 | 11.0 | 18.3 | 11.2 | 63.1 | 44.4 |
| E6-a R 25 Block 13 | 5.1 | 4.5 | 2.9 | 9.5 | 3.7 | 65.7 | 47.3 |
| E6-a R 25 Block 14 | 14.9 | 13.8 | 11.0 | 23.3 | 11.7 | 59.8 | 41.3 |
| E6-a R 25 Block 15 | 4.0 | 2.9 | 2.0 | 7.9 | 3.4 | 66.0 | 47.4 |
| E6-a B 25,34 Block 12 | 9.4 | 6.1 | 9.1 | 11.8 | 10.5 | 65.2 | 44.9 |
| E6-a B 25,34 Block 13 | 3.2 | 1.7 | 3.3 | 5.6 | 2.4 | 65.7 | 47.2 |
| E6-a B 25,34 Block 14 | 17.0 | 12.4 | 13.6 | 29.4 | 12.6 | 58.9 | 45.2 |
| E6-a B 25,34 Block 15 | 6.0 | 4.2 | 5.5 | 8.5 | 6.0 | 65.3 | 45.9 |
| E6-b ratio 0.25 | 14.6 | 10.9 | 12.1 | 23.3 | 12.1 | 61.6 | 41.7 |
| E6-b ratio 0.5 | 26.9 | 26.5 | 15.2 | 49.6 | 16.1 | 57.8 | 33.2 |
| E6-c1 | 47.2 | 48.4 | 24.8 | 75.3 | 40.5 | 65.7 | 5.6 |
| E6-c2 | 39.7 | 46.6 | 23.0 | 60.3 | 29.0 | 59.7 | 25.2 |
| E6-c3 | 18.6 | 17.4 | 12.3 | 32.3 | 12.4 | 57.8 | 39.7 |
| E6-c4 | 70.1 | 67.8 | 71.7 | 71.6 | 69.3 | 18.1 | 69.7 |














With a single Block (E6-a), none of the four flipped the hair arrangement or turned it red. Blocks 13 and 15 stay nearly plain A at pixel differences of 3-6, while Blocks 12 and 14 change eye color and collar. The Blocks 12-15 range that flipped in E3a and E5 only flips when all four are swapped together.
The mix ratio (E6-b) changes color continuously: orange-leaning yellow at 0.25, a strongly reddish orange at 0.5, with face and composition still plain A. At 0.5, though, clothing color and skin rendering had also moved toward plain R, and there was no stage where only the hair changed.
The spatial limit (E6-c) could not isolate hair color either. Letting only the left-half image tokens read position 25 still turned the whole frame into plain R (pixel difference 5.6 from plain R). Narrowing to the top left still lands on the R side (25.2). The Blocks 12-15 version came out almost the same as E5-5 without a region, at 18.6 vs 18.7.
Even with the cross-attention output held in place outside the region, it spread over the whole frame through that Block’s self-attention and MLP and the Blocks after it. The left/right cross-referencing seen in E4 is what carries the swap’s influence across the frame here.
In the images with swapped hair-color tokens, the collar that was asymmetric in plain A (pale pink frills on the left, black frills on the right) was replaced every time by matching high necks with a button row on both. E5 did the same, and it is common to every swap that passes through Blocks 12-15.
Across Block, step, space, and mix ratio, no condition that rewrites hair color alone while keeping the original picture turned up within this harness.
Any swap passing through Blocks 12-15, where the flip happens, always changed the clothes too, and cutting in space did not stop self-attention from spreading it. I think the hair-color token is being handled together with the clothing information there.
To cut it, self-attention would have to be cut by region as well as cross-attention, and in E4 doing that split the picture into two halves.
E7 Rewrite the residual stream directly
Swapping the conditioning changed the clothes along with it, so I left the conditioning entirely alone and moved to rewriting the residual inside the DiT directly.
The GPT-OSS paper from the opening split the hidden state into role and content and rewrote only the content. On the image side I took the direction for the left girl’s hair color out of the Block output hidden state and added it to the hair tokens only.
The direction is built by running A and R on the same seed, saving each Block’s output at each step, and taking R minus A over the tokens of the hair mask.
The hair mask comes from plain A (seed 42) by thresholding blonde pixels by color and pooling to the 52x76 grid.
Two ways of adding: “mean direction”, which adds one averaged vector over the region, and “per token”, which adds each token’s own difference as is.
| ID | Direction | Blocks added | Region | Coefficient | Method |
|---|---|---|---|---|---|
| E7-1 | R - A | 12-15 | left hair | 1.0 | mean direction |
| E7-2 | R - A | 12-15 | top left (columns 0-37, rows 0-29) | 1.0 | mean direction |
| E7-3 | R - A | 8-19 | left hair | 1.0 | mean direction |
| E7-4 | R - A | 12-15 | left hair | 2.0 | mean direction |
| E7-5 | R - A | 12-15 | left hair | 1.0 | per token |
| E7-6 | R - A | 12-15 | right hair (left hair region mirrored) | 1.0 | mean direction |
| E7-7 | B - A | 12-15 | left hair | 1.0 | mean direction |
E7-6 is the control: does the right girl turn red when the direction meant for the left girl is added to her hair.
Seed 42, all steps, cond side only, everything else as in E0. Judgment and pixel differences as in E5, also split into inside and outside the hair mask.
The first seven runs saturated, so I added coefficients 0.1 and 0.3, a version with the 20 largest-magnitude dims of the difference vector zeroed, and combinations of both.
| ID | Direction | Blocks | Region | Coefficient | Dims removed |
|---|---|---|---|---|---|
| E7-8 | R - A | 12-15 | left hair | 0.1 | none |
| E7-9 | R - A | 12-15 | left hair | 0.3 | none |
| E7-10 | R - A | 12-15 | left hair | 1.0 | top 20 zeroed |
| E7-11 | R - A | 12-15 | left hair | 0.1 | top 20 zeroed |
| E7-12 | R - A | 12-15 | left hair | 0.3 | top 20 zeroed |
| E7-13 | R - A | 12-15 | left hair | 0.05 | top 20 zeroed |
| E7-14 | R - A | 12-15 | left hair | 0.1 | top 102 (5%) zeroed |
E7 hair mask and size of the difference vector
Blonde pixels from the left half of plain A (R>150, G>110, B<140, R-B>40), with a token counted as hair at 25% or more coverage per 16x16 pixel cell. 562 of the 3952 tokens, within rows 4-51 and columns 10-32. A few tokens also reach into the top of the frilled collar.

The residuals were re-saved for A, R, and B over Blocks 8-19 at all 30 steps in fp32 (11.7 GB each).
The table shows the norm of the mean vector inside the hair mask and the mean residual norm over the same region.
| Block | step | Difference vector norm | Residual norm | Ratio | Share of norm in top 20 dims |
|---|---|---|---|---|---|
| 12 | 0 | 80 | 9706 | 0.008 | 0.80 |
| 12 | 15 | 1494 | 8790 | 0.170 | 0.80 |
| 13 | 0 | 76 | 9781 | 0.008 | 0.70 |
| 13 | 15 | 1496 | 9125 | 0.164 | 0.79 |
| 14 | 0 | 103 | 10646 | 0.010 | 0.71 |
| 14 | 15 | 1619 | 9935 | 0.163 | 0.83 |
| 15 | 0 | 156 | 10892 | 0.014 | 0.84 |
| 15 | 15 | 1527 | 10800 | 0.141 | 0.79 |
The step 15 difference vector is about 19 times the step 0 one, at 14-17% of the residual.
Seventy to eighty percent of that norm sits in 20 of the 2048 dims, so the difference vector is decided less by a direction than by the size of a few dims. Whether these 20 dims are the same as the huge-value token dims found in E2 is something I did not check. In the flat-color S2 earlier, the 20 dims holding the color difference barely overlapped the huge-value dims.
Zeroing the top 20 dims brought the norm down to 0.55-0.71 times.
E7 results
| ID | Direction | Blocks | Region | Coefficient | Dims removed | Left girl | Right girl | Does the mask shape show |
|---|---|---|---|---|---|---|---|---|
| E7-1 | R-A mean | 12-15 | left hair | 1.0 | none | Where the hair was is filled with a white slab, red blocks at the edge and inside. Face gray and blurred, features unreadable | Bob length, gray skin, black high neck | yes |
| E7-2 | R-A mean | 12-15 | top left | 1.0 | none | Head covered by a flat orange rectangle, no hair visible. Clothes pink | Bob length, gray skin, black high neck | yes |
| E7-3 | R-A mean | 8-19 | left hair | 1.0 | none | Where the hair was is a light blue blocky mass. Face hidden. Whole frame in pink noise | Dark brown wavy medium-length hair, red spots on the face | yes |
| E7-4 | R-A mean | 12-15 | left hair | 2.0 | none | Where the hair was is a light blue blocky mass. Whole frame in pink noise | Gray-brown hair, red and pink spots on the face | yes |
| E7-5 | R-A per token | 12-15 | left hair | 1.0 | none | Where the hair was is light blue with only the hair strands left as lines. Face in light blue and red stripes | Bob length, gray skin, black high neck | yes |
| E7-6 | R-A mean | 12-15 | right hair | 1.0 | none | Hair still blonde, saturation up. Eyes brown, salmon-colored high neck | Where the hair was is white and gray grainy noise, face gray with no features | yes |
| E7-7 | B-A mean | 12-15 | left hair | 1.0 | none | All black (NaN in the residual) | Same | - |
| E7-8 | R-A mean | 12-15 | left hair | 0.1 | none | Hair a flat red-to-orange fill, bang strands gone. Face outline and expression remain, eyes dark red. Clothes a dark red ribbed high neck | Long straight hair and bangs kept, dark red at the collar, a green band at the right edge of the hair | slightly |
| E7-9 | R-A mean | 12-15 | left hair | 0.3 | none | Hair orange-to-red in blocky shading. Face a gray-white surface with features gone | Bob length, gray-white skin | yes |
| E7-10 | R-A mean | 12-15 | left hair | 1.0 | top 20 zeroed | Hair a beige halftone pattern with red patches. Face a gray surface | Bob length, gray-white skin, no frilled collar | yes |
| E7-11 | R-A mean | 12-15 | left hair | 0.1 | top 20 zeroed | Hair orange-to-red with strands kept. Face as in plain A, eyes from blue-green to brown. Pale pink frilled collar kept with a slight red tint | Long black hair, black jacket with frilled collar unchanged | top edge slightly squared |
| E7-12 | R-A mean | 12-15 | left hair | 0.3 | top 20 zeroed | Hair a flat red fill with vertical stripes, face reduced to feature lines with gray-white skin. Clothes also a red surface | Bob length, gray-white skin, black high neck | yes |
| E7-13 | R-A mean | 12-15 | left hair | 0.05 | top 20 zeroed | Hair orange with strands kept. Face as in plain A, eyes brown. Pale pink frilled collar kept | Long black hair, black jacket with frilled collar unchanged | no |
| E7-14 | R-A mean | 12-15 | left hair | 0.1 | top 102 zeroed | Hair still blonde with saturation barely up. Eyes black. Otherwise nearly plain A | Nearly plain A | no |
| ID | plain A overall | top left | top right | bottom left | bottom right | in hair | outside hair | plain R overall |
|---|---|---|---|---|---|---|---|---|
| E7-1 | 42.6 | 47.1 | 25.3 | 55.4 | 42.7 | 91.5 | 34.5 | 51.2 |
| E7-2 | 52.5 | 52.5 | 45.4 | 64.8 | 47.4 | 43.3 | 54.0 | 55.3 |
| E7-3 | 47.3 | 48.0 | 33.1 | 56.0 | 52.3 | 91.0 | 40.1 | 77.0 |
| E7-4 | 49.3 | 43.6 | 39.9 | 51.8 | 61.8 | 82.5 | 43.7 | 78.7 |
| E7-5 | 42.5 | 44.8 | 29.5 | 51.3 | 44.4 | 83.8 | 35.6 | 48.4 |
| E7-6 | 47.2 | 24.7 | 63.2 | 46.1 | 55.0 | 49.5 | 46.9 | 52.3 |
| E7-7 | 163.8 | 214.0 | 156.1 | 186.2 | 99.1 | 156.3 | 165.1 | 129.4 |
| E7-8 | 38.3 | 48.1 | 17.1 | 70.6 | 17.4 | 104.9 | 27.3 | 26.7 |
| E7-9 | 45.0 | 40.6 | 23.5 | 75.2 | 40.9 | 76.4 | 39.8 | 26.0 |
| E7-10 | 35.2 | 21.8 | 23.0 | 50.1 | 45.8 | 33.9 | 35.4 | 34.1 |
| E7-11 | 31.2 | 40.6 | 15.3 | 53.2 | 15.6 | 93.0 | 20.9 | 31.2 |
| E7-12 | 41.2 | 46.0 | 18.9 | 75.3 | 24.6 | 93.1 | 32.6 | 25.1 |
| E7-13 | 22.3 | 27.1 | 13.4 | 34.9 | 13.7 | 60.1 | 16.0 | 36.8 |
| E7-14 | 8.9 | 9.9 | 6.4 | 13.2 | 6.0 | 23.0 | 6.5 | 43.6 |
Pixel differences are mean absolute differences vs plain A (RGB, 0-255). In hair and outside hair refer to the hair mask mapped back to pixels.














The mean direction at coefficient 1.0 (E7-1) saturated. The hair mask region became a white slab with the token grid’s jaggedness showing at the outline. Adding the 20 dims that make up most of the difference vector to every token in the hair mask blew up the residual.
A rectangular region (E7-2), wider Blocks (E7-3), a larger coefficient (E7-4), and per-token (E7-5) each broke in a different way, and none of them drew anything that reads as hair.
The control on the right hair (E7-6) turned the right hair region into grainy noise while the left hair stayed blonde with higher saturation. A direction made for the left still wrecks the right when added there.
The B minus A direction (E7-7) gave NaN in the residual and an all-black image. The left/right swap difference vector could not be used the same way as the color-change one.
Lowering the coefficient to 0.1 (E7-8) made the hair red and kept the face outline, but the hair is a flat fill with the bang strands gone and the clothes turned into a dark red high neck.
Zeroing only the top 20 dims (E7-10) turned the saturation into a different kind of breakage (a halftone pattern).
With the top 20 dims zeroed and coefficient 0.1 (E7-11), only the left hair turned reddish while face, collar, and the right girl survived. The hair is orange-to-red with the strands kept, the pale pink frilled collar stays, and so do the right girl’s black hair and black jacket. What changed is the left girl’s eye color going brown and a slight red tint on the collar. Pixel difference is 31.2 from plain A, 20.9 outside the hair, and around 15 in the top-right and bottom-right quadrants.
Raising the coefficient to 0.3 under the same conditions (E7-12) saturated again.
Lowering it to 0.05 (E7-13) gives orange hair with the strands kept, and face, collar, and the right girl stay plain A. The outside-hair pixel difference is 16.0, the right-side quadrants 13-14, less change from plain A than E7-11, and the hair stops at orange short of red. The mask shape does not show.
Widening the removal to the top 5% (102 dims) at coefficient 0.1 (E7-14) left the hair blonde and brought the pixel difference vs plain A down to 8.9. The component that was changing the color also lived in dims outside the top 1% but inside the top 5%.
E7-11 and E7-13 gave what the conditioning route never did: rewriting the residual stream directly changed only the left hair’s color while face, collar, and the right girl survived.
E5-5 (conditioning side, limited to Blocks 12-15) had a pixel difference of 18.7 vs plain A with the clothes changed; E7-13 is 22.3 with the clothes intact. The difference lands in different places.
The window is very narrow, though. Coefficient 0.05-0.1, breakage at 0.3, the top 20 dims must be removed, and widening the removal to 102 dims stops the color from changing.
The direction itself is dominated by a few huge-value dims, and I think the remaining component only starts acting as color once those are removed, but whether that remainder represents hair color alone is something I did not check. The eyes turning brown is common to all four conditions, so hair and eyes may be changing through the same component.
Flat-color direction and the hair-color attribute
Bringing what the flat-color experiments showed back to the two-girl case.
Color is decided by denoise step 1. That held for red, blue, and the girl image. The red-blue residual difference is 85-94% concentrated in 20 of the 2048 dims, and those 20 are separate from the residual’s huge-value dims.
Adding the difference to Blocks 24-27 changed the color, and one of them at steps 0-1 alone was enough. Adding to the middle Blocks (8-23) blows up the residual. Adding to only the left half in the last Block group at steps 0-1 split the image into blue on the left and red on the right at a one-pixel-wide boundary.
Adding the same direction to a girl’s hair produced a blue surface in the exact shape of the hair, and the girl was redrawn. The direction I built from the flat-color difference seems to act by repainting a region as a surface.
Part of the reason hair color could not be isolated in the two-girl case is probably here: the color direction in the residual was “what color to make this region as a surface”, not “what color to make this hair”.
With this extraction method, no stable direction that changes hair color alone while keeping the object’s shape turned up. E7-11 and E7-13 are the two images that survived under the narrow condition of coefficient 0.05-0.1 with the top 20 dims removed.