Tech56 min read

Anima DiT Residual Steering on M1 Max Flips Flat Color, Not Hair Color Alone

IkesanContents

This started from a conversation about the paper that swapped GPT-OSS hidden states for a symbolic-structure formula and got almost the same answers.
Rebuild every layer’s hidden state of an LLM with a closed-form tensor product representation and the model’s answers barely change, and role-rewriting interventions go through too.
The question was whether the same view works for an image-generation DiT.
What grammar and logic are to an LLM’s internal algorithms would map, I figured, to spatial layout, part structure, attribute binding, depth and occlusion, and the division of labor across denoise steps.

Earlier, in the write-up that looked inside the conditioning of a broken 4-character Anima LoRA, I intervened up to the conditioning right before the DiT, but never looked at how the DiT turns that conditioning into image space.
I had only probed the entrance and stopped at “it seems to understand the prompt”.

This time I open up the DiT itself.
The order goes from red and blue flat-color images, to find where color is decided and written, to one girl’s hair, and then to two girls side by side. Everything runs on an M1 Max 64GB with plain Anima-Base v1.0, no LoRA.

Test environment

ItemDetails
MachineApple M1 Max 64GB, macOS 26.5
RunnerA homemade harness (.image-work/dit-probe/) that imports the comfy package from ComfyUI v0.30.1 (commit 0764232) directly
DiTAnima-Base v1.0 (bf16)
Text encoderQwen3-0.6B-Base + T5 tokenizer, Anima’s own LLMAdapter
LoRANone (Turbo removed too)
Common settings30 steps, CFG 4.0, euler, scheduler simple, 1216x832 (token grid 52x76)
Negativeworst quality, low quality, blurry, lowres, text, watermark, bad hands
MeasuredOnly the cond side of CFG

The DiT lives in ComfyUI’s comfy/ldm/cosmos/predict2.py.
Each Block runs self-attention (image tokens to image tokens), then cross-attention (image tokens reading the conditioning), then MLP, with adaLN applying the timestep embedding at each stage.
Text enters the image side only through cross-attention, so in E1 (the cross-attention map experiment listed in the next section) I use the cross-attention weights to measure how strongly each image token reads each text position. Weights alone do not decide the actual contribution, so later sections also check by swapping the conditioning.
An MMDiT like Qwen-Image mixes image and text in the same attention, whereas Anima lets you handle the text-side weights separately.
The model layout follows the Anima-Base model card. The values below are what I measured by running it.

ItemMeasured
Blocks28
hidden dim2048
Heads16 (head_dim 128)
Token grid (rows x cols)52x76 = 3952 image tokens (patch 2, 1216x832 stays even)
conditioning shape(1, 512, 1024). Actual prompts are around 38 tokens, the rest is zero padding
Time per image (30 steps)254 s (38 s at 4 steps)

The cond and uncond sides are told apart with transformer_options["cond_or_uncond"].
At this resolution cond and uncond arrive as alternating separate calls, so I measured and modified only the rows marked 0.
The mapping from words to cross-attention key positions came from actually tokenizing with the T5 tokenizer.

The first snag in the harness check was the left/right self-attention block. Passing a (b,1,S,S) boolean mask to SDPA on MPS produced NaN a few Blocks in and a pitch-black image.
Forcing synchronization made it disappear. I never found the cause, but slicing the left and right tokens apart and running the model’s own attention on each side separately (equivalent to a block-diagonal mask) worked, so I went with that.
The other snag was the residual stream. ComfyUI keeps it in fp32 and the values are large (up to around 2.7e5), so saving in fp16 turned a handful of elements in the middle Blocks into inf. E2 (the residual difference experiment) saves in fp32.

To confirm that measuring does not change the image, I compared the output PNGs with and without the attn and resid hooks. Bit-identical. The conditioning swap also matched a plain generation of prompt B bit for bit when swapped across all Blocks and all steps.

Order of experiments

The actual order was the two-girl series (E0-E7) first; the flat-color series (S0-S5) was added afterwards.

StageIDNameWhat is doneWhat is measured
Flat colorS0Does a flat color come outGenerate red, blue, green, yellow flat-color prompts at seed 42Mean RGB and standard deviation. Is it a flat color
Flat colorS1The step where color is decidedFor red and blue, decode each step’s denoise prediction (x0) with the VAE and take mean RGBAt which step do red and blue separate
Flat colorS2Residual differenceRun red and blue on the same seed, save all Block, all step outputs, take the differenceWhich Block has the largest L2 difference. Are the large-difference dims the same as the huge-value dims
Flat colorS3Direction transplantAdd the mean blue minus red direction to the red generation. 7 runs over Block groups, 2 runs over all Blocks with limited stepsDoes mean RGB move toward blue. Which Block changes
Flat colorS3xSingle Block and coefficientSplit the Block group that changed color in S3 one by one, add at steps 0-1 only. One run at coefficient 0.7Is one Block enough
Flat colorS4Cut in spaceAdd the direction only to the left-half tokens of the Block group that changed color in S3Does only the left half turn blue
One girlS5Transplant to hairAdd the same direction to the hair tokens of a single blonde girlDoes the hair turn blue. Do face and clothes survive
Two girlsE0BaselineGenerate two prompts with left/right hair colors swapped, 5 seeds each, no interventionAre the left/right hair colors as instructed
Two girlsE1cross-attention mapMap the attention weights on blonde and black back to space for all Blocks and stepsWhich Block and step the bias appears in
Two girlsE2Residual stream differenceRun both prompts on the same seed, take the difference of all Block outputsLeft/right asymmetry of the difference
Two girlsE3conditioning swapSwap in the other prompt’s conditioning only for a given Block group or step rangeWhere the swap flips the hair colors
Two girlsE4Left/right self-attention blockCut cross-references between left and right in image-token self-attentionWhat the block does
Two girlsE5, E6Token-level rewriteSwap only the hair-color token positions. Narrow by Block, step, space, mix ratioDoes only the hair color change while the picture survives
Two girlsE7Rewrite the residual stream directlyAdd the hair-color difference vector only to hair tokensDoes only the hair change

S0 Does a flat color come out

The prompt is solid red background, plain background, no humans, flat color, empty with only the color word changed, four prompts. Negative stays the common one.
All four came out flat on the first try, so I did not iterate on the wording.

ColorMean RGBStd dev (R / G / B)Flatness
red213 / 18 / 300.61 / 0.55 / 0.46100%
blue1 / 87 / 1870.50 / 0.41 / 0.57100%
green62 / 132 / 790.51 / 0.43 / 0.51100%
yellow255 / 241 / 730.01 / 0.52 / 0.76100%

Flatness is the share of pixels whose three RGB channels all sit within plus or minus 20 of the image mean. Every image spans only 5-10 levels per channel.

S0 generated result for red

S0 generated result for blue

S0 generated result for green

S0 generated result for yellow

T5 tokenization gives 15 tokens for all four, and the color word lands as a single token at position 1, right after solid. One row of the conditioning corresponded to one color word.

S1 The step where color is decided

For each sampling step I decoded the sampler’s denoise prediction (x0) with the VAE and took the mean RGB.
The capture uses ComfyUI’s sampler callback: save the denoised tensor handed over for previews as is, then decode them all after sampling.
The DiT is untouched, so the final image matched the normal gen output bit for bit (checked with SHA-256). Thirty decodes took 110 s, 3.7 s each.

Besides red and blue I also did G, the single-girl image used in S5 (1girl, upper body, looking at viewer, simple background, blonde hair, long hair).

stepredblueG
0255 / 66 / 595 / 133 / 255255 / 254 / 250
1220 / 17 / 283 / 87 / 179239 / 212 / 185
2210 / 24 / 354 / 89 / 180244 / 225 / 199
3213 / 24 / 332 / 89 / 183248 / 228 / 206
5213 / 20 / 311 / 88 / 185248 / 225 / 207
10213 / 17 / 301 / 87 / 186247 / 223 / 204
15213 / 17 / 301 / 87 / 186247 / 222 / 203
29213 / 18 / 301 / 87 / 187246 / 222 / 203

S1 x0 of each step for red, shrunk to 1/8 and laid out in a strip

S1 x0 of each step for blue, shrunk to 1/8 and laid out in a strip

S1 x0 of each step for G, shrunk to 1/8 and laid out in a strip

In all three, every channel is within plus or minus 30 of the final color by step 1.
The step 0 x0 has one channel pinned at 255 and misses, but by the second prediction the color is already set and the values only tighten from there. Red’s flatness is 51% at step 0, 96% at step 1, 100% at step 2.
G’s mean color was also settled by step 1. The shape, on the other hand, is a blurry blob at step 1, becomes recognizably a person at step 4, and details like eyes and collar only show up from around step 10.

Color seems to be decided in the first 2 steps. From S3 on, whenever I add a direction, there is a condition that adds it at steps 0-1 only, alongside the all-steps condition.

S2 Residual difference

I ran red and blue on the same seed 42 and saved the outputs of all 28 Blocks at all 30 steps in fp32 (25 GB per run).
For each step and Block I took the difference of the outputs and averaged over all tokens to get a vector d = blue minus red.
What I wanted to check was whether the difference is concentrated in a few dims, and if so whether those are the same huge-value dims that E2 later finds. So I computed the norm of d, its ratio to the norm of red’s residual, the share of the norm held by the top 20 dims of |d|, and how many of those 20 overlap with the 20 largest dims of red’s own residual.
The table shows steps 1 and 15, every fourth Block.

stepBlocknorm of dresidual normratioshare of top 20 dimsoverlap with huge-value dims
105300.180.930
14103883090.130.880
18124287030.140.872
1121239106300.120.880
1161250119600.110.880
1201377137000.100.860
1243054142200.220.892
12712030618200.190.581
15027570.480.901
154327186140.380.941
158347787930.400.931
15123511106000.330.922
15163527119300.300.922
15203583136100.260.921
15244925150200.330.914
152711460635200.180.654

S2 per-Block statistics of the blue minus red difference vector. From top: norm, ratio to residual norm, share of top 20 dims, overlap count with huge-value dims

S2 token-mean L2 of the difference. Rows are saved steps, columns are Blocks

S2 spatial map of the difference. Rows are steps, columns are Blocks

The red-blue difference was 85-94% concentrated in 20 of the 2048 dims. Only Blocks 26-27 drop to 58-65%. Same shape as E7 later, where the hair-color difference is 70-80% concentrated in 20 dims too.
Those 20 dims barely overlap with the huge-value dims of the residual itself (0-2 in the middle Blocks, at most 4 in Blocks 24 and 27). Back in E7 of the two-girl series I had assumed they were the same dims. They were not. The color difference mostly lives in ordinary-sized dims.

Spatially the difference is uniform. In Blocks 0-11 and 22-27 the spread across tokens (coefficient of variation) is 0.02-0.17.
Only in Blocks 12-21 do 62 of the 3952 tokens (1.6%) show a difference more than 10 times the median, at the same positions across Blocks 12-21, and it vanished at Block 22.
Block 0’s difference is 5-27, essentially zero next to the 1000-plus from Block 1 onward. For these two prompts, Block 0’s output seems to barely reflect the prompt.

S3 Direction transplant

I added d from S2 to the red generation. Region: all tokens, mean direction, coefficient 1.0. Seven runs adding it at all steps to each of seven Block groups, plus two runs over all Blocks with a limited step range.
Projection is the mean RGB projected onto the red-to-blue line: 0 is red, 1 is blue.

IDBlockstepMean RGBFlatnessProjectionAppearance
S3-10-3all255 / 40 / 143%-0.16Saturated orange-red with yellow grain
S3-24-7all254 / 21 / 134%-0.18Saturated red with fine grain
S3-38-11all0 / 0 / 0--All black (residual NaN)
S3-412-15all0 / 0 / 0--All black (residual NaN)
S3-516-19all0 / 0 / 0--All black (residual NaN)
S3-620-23all0 / 0 / 0--All black (residual NaN)
S3-724-27all0 / 162 / 25597%1.22Flat cyan-leaning blue. Faint token-grid weave
S3-80-270-10 / 0 / 0--All black (residual NaN)
S3-90-272-290 / 0 / 0--All black (residual NaN)

For the two saturated runs I also tried a version with the top 20 dims zeroed.

IDBlockstepMean RGBFlatnessProjectionAppearance
S3-1c0-3all236 / 234 / 24361%0.59Near-white pale purple, grainy
S3-2c4-7all167 / 223 / 25530%0.80Light blue, strong diagonal weave

S3 S3-1 Blocks 0-3

S3 S3-2 Blocks 4-7

S3 S3-3 Blocks 8-11

S3 S3-4 Blocks 12-15

S3 S3-5 Blocks 16-19

S3 S3-6 Blocks 20-23

S3 S3-7 Blocks 24-27

S3 S3-8 all Blocks, steps 0-1

S3 S3-9 all Blocks, steps 2-29

S3 S3-1c Blocks 0-3, top 20 dims zeroed

S3 S3-2c Blocks 4-7, top 20 dims zeroed

It turned blue only when added to Blocks 24-27. And it overshot blue into cyan (projection 1.22).
Blocks 8-23 went NaN and all black when four Blocks were fed at once. Adding to all Blocks with a narrowed step range did the same. Blocks 0-7 stayed red and saturated.
The difference vector shifts the input of the next Block at every Block it touches, so adding to four at once stacks the effect. The NaN may be from that, but this measurement cannot tell.
With the top 20 dims zeroed, Blocks 0-7 also moved toward blue, but flatness fell to 30-60% and a weave pattern appeared.

S3 follow-up Single Block and coefficient

Since S1 showed color settled by step 1, I split Blocks 24-27 one by one and added at steps 0-1 only. One more run with all four at coefficient 0.7.

IDBlockstepCoefficientMean RGBFlatnessProjection
S3x-1240-11.00 / 142 / 255100%1.20
S3x-2250-11.00 / 142 / 255100%1.20
S3x-3260-11.00 / 136 / 255100%1.19
S3x-4270-11.00 / 137 / 255100%1.19
S3x-524-270-10.70 / 175 / 255100%1.23

S3 follow-up S3x-1 Block 24 only, steps 0-1

S3 follow-up S3x-2 Block 25 only, steps 0-1

S3 follow-up S3x-3 Block 26 only, steps 0-1

S3 follow-up S3x-4 Block 27 only, steps 0-1

S3 follow-up S3x-5 Blocks 24-27, steps 0-1, coefficient 0.7

One Block, steps 0-1 only, and red went fully blue. Any of 24-27 does it, and flatness stays at 100%.
The overshoot into cyan past projection 1 is the same whether all four or a single one, and dropping the coefficient to 0.7 still gave 1.23. Only two points, but the color does not look like it scales linearly with the size of the direction.

S4 Cut in space

I added d to Blocks 24-27 only on the left-half tokens (columns 0-37). One version at all steps, one at steps 0-1 only.

IDstepLeft-half mean RGBRight-half mean RGBBoundary
S4-1all0 / 161 / 255 (projection 1.22)188 / 170 / 131 (projection 0.43)Color gap 141 at x=631. Jagged along token borders
S4-20-10 / 162 / 255 (projection 1.22)171 / 6 / 9 (projection 0.07)Color gap 327 at x=629. One pixel wide

S4 S4-1 left half of Blocks 24-27, all steps

S4 S4-2 left half of Blocks 24-27, steps 0-1

The steps 0-1 version has the left blue, the right still red, and the boundary is a one-pixel-wide straight line.
The all-steps version turns the left blue, but the right collapses into beige and the boundary becomes jagged along the token borders.
The boundary sits at pixels 629-631, 1.3-1.4 tokens to the right of pixel 608, where the column 38 token border should be. I assume that is the VAE’s receptive field (the surrounding area consulted to reconstruct one pixel).

Under these conditions, adding the flat-color difference direction to the last 4 Blocks at steps 0-1 changed the color on one side only.
Whether color is actually written in those Blocks during an unmodified generation is something I did not check. Adding to the middle Blocks (8-23) blows up the residual, so the same form cannot be added there.

S5 Transplant to hair

I added the direction obtained from flat colors to the hair tokens of G, a single blonde girl.
The hair is already yellowish, so the direction is blue minus yellow, and yellow’s residual was saved only for Blocks 24-27 at steps 0-1. I added one run with blue minus red for comparison.

G’s hair mask takes blonde pixels from plain G (R>150, G>110, B<160, R-B>40) and marks a token as hair at 25% or more coverage. 1054 of the 3952 tokens. The B ceiling is 160 instead of 140 to catch the highlights at the crown, which lets 2-3 tokens of neck shadow in.

S5 G's hair mask overlaid on plain G

Plain G’s mean RGB inside the hair is 235 / 190 / 118.

IDDirectionBlockstepRegionPixel diff vs plain G (all / in hair / outside hair)Mean RGB in hair
S5-1blue - yellow24-270-1hair94 / 185 / 610 / 3 / 252
S5-2blue - red24-270-1hair68 / 137 / 430 / 177 / 256
S5-3blue - yellow24-270-1all165 / 177 / 1610 / 25 / 248
S5-4blue - yellow24-27allhair223 / 181 / 2390 / 0 / 0

S5 S5-1 blue minus yellow added to hair tokens in Blocks 24-27 at steps 0-1

S5 S5-2 blue minus red added to hair tokens in Blocks 24-27 at steps 0-1

S5 S5-3 blue minus yellow added to all tokens in Blocks 24-27 at steps 0-1

S5 S5-4 blue minus yellow added to hair tokens in Blocks 24-27 at all steps

S5-1, added to the hair tokens, turned the hair mask’s exact shape into a flat blue slab.
The slab has a smooth outline with a black edge, the girl is redrawn small in the empty space below it, her clothes stay red, and the rendering flattens into line-art style.
S5-2 with blue minus red broke the same way, in light blue.
S5-3, added to all tokens, went entirely blue and the girl disappeared. S5-4, added at all steps, blew up the residual and went black.

The direction obtained from flat colors seems to be a direction that turns the region it is added to into a flat blue surface.
Added to the hair region, it drew that region as a blue object and the girl was redrawn in whatever space was left.
Steps 0-1 decide composition as well as color (G’s shape is not fixed until step 4 in S1), so I think adding a surface instruction there is what changed the composition along with it.

The two-girl hair-color series

Chronologically this is the experiment I did before the flat colors. Two girls side by side, blonde on the left and black hair on the right, and I tracked the binding of attribute to position from conditioning swaps through to editing the residual stream directly.

Why no LoRA

With a character LoRA applied, there is no telling whether a structure I find belongs to Anima or to that LoRA. In the 4-character LoRA write-up above I also confirmed that appearance gets absorbed into the trigger depending on how captions are written, so LoRA influence would get mixed in.
With plain Anima and ordinary attributes like blonde hair and black hair, each attribute maps to 1-2 tokens on the text side, which makes it easier to check which patches the word’s cross-attention weight landed on. For the same reason I removed the Turbo LoRA too and ran normal steps with CFG on.

Prompts

The base is two prompts of two girls side by side with only the left/right hair colors swapped.

LabelPrompt
A2girls, standing side by side, upper body, looking at viewer, simple background. The girl on the left has blonde hair. The girl on the right has black hair.
B2girls, standing side by side, upper body, looking at viewer, simple background. The girl on the left has black hair. The girl on the right has blonde hair.

A and B share the same set of words. Only the binding differs.
The encoder-side representations differ too, but mapping the cross-attention weights back to space per word should show where the DiT turns that difference into position.

E0 Baseline

A and B generated plain at seeds 42-46. All 10 images have two girls, with the left/right hair colors as instructed.
Judgment was done by opening each image and writing down just the head count and left/right hair colors as seen.

PromptseedLeft hairRight hair
A42blondeblack
B42blackblonde
A43blonde (highly saturated yellow)black
B43blackblonde (highly saturated yellow)
A44blondeblack
B44blackblonde
A45blondeblack
B45blackblonde
A46blondeblack
B46black (bluish, near navy)blonde

E0 prompt A seed 42 (instructed: left blonde, right black)

E0 prompt B seed 42 (instructed: left black, right blonde)

E0 prompt A seed 43 (instructed: left blonde, right black)

E0 prompt B seed 43 (instructed: left black, right blonde)

E0 prompt A seed 44 (instructed: left blonde, right black)

E0 prompt B seed 44 (instructed: left black, right blonde)

E0 prompt A seed 45 (instructed: left blonde, right black)

E0 prompt B seed 45 (instructed: left black, right blonde)

E0 prompt A seed 46 (instructed: left blonde, right black)

E0 prompt B seed 46 (instructed: left black, right blonde)

Clothing color was never specified, but seed 42 A gives the left girl whitish clothes and the right girl black, and seed 46 B gives a navy jacket on the left and a red sweater on the right.

I had planned to use a seed with swapped hair colors and positions in E4, but all 10 followed the instruction and none appeared. Plain Anima apparently does not break on this prompt.
E4 adds clothing colors to the hair colors and first checks whether that departs from the instruction.

E1 cross-attention map

I regenerated A at seed 42, which followed the instruction in E0, under the same settings while recording the cross-attention weights of all Blocks and all steps.
The recording is computed separately from copies of q and k, so the output image is identical to the unrecorded E0 (bit-identity confirmed in the harness check).
For each Block and step I took the attention weight that each of the 3952 image tokens places on blonde (T5 token position 25) and black (position 34), averaged over heads, and put it back on the 52x76 grid.
Then I defined the left-half share as the sum of weight landing in the left half of the grid divided by the total. 0.5 means no left/right bias.

E1 left-half share of cross-attention on the word blonde. Rows are sampling steps, columns are DiT Blocks

E1 left-half share of cross-attention on the word black. Rows are sampling steps, columns are DiT Blocks

The numbers look like this.

Blockblonde step 0blonde step 10blonde step 29black step 0black step 10black step 29
00.500.500.500.500.500.48
20.540.630.580.460.320.44
50.670.810.570.310.250.46
90.910.900.640.090.120.38
130.900.840.660.100.120.31
160.810.750.600.190.200.39
190.610.630.520.410.390.48
200.490.490.490.490.490.50
270.500.500.500.500.500.50

Blocks 0-1 and Blocks 20-27 sit around 0.50 for both words at every step, with no left/right bias.
The bias starts around Block 2 and is strongest in Blocks 8-16. At step 0, Block 9 has blonde at 0.91 and black at 0.09, almost mirror images.
The bias weakens as steps progress. Block 9’s blonde fell from 0.91 at step 0 to 0.64 at step 29.

The left-half share only tells direction, so I also computed the mean weight over the whole grid (per token). The conditioning has 512 tokens, so uniform attention would give 1/512, about 0.00195, per token. The table is scaled by 1000.

Blockblonde step 0blonde step 10blonde step 29black step 0black step 10black step 29
01.00.91.00.90.90.9
52.31.60.31.70.90.3
96.22.30.47.72.20.3
133.61.10.22.91.20.2
195.62.10.85.92.00.7
200.90.91.00.90.91.0
270.90.91.01.01.01.1

At step 0, Block 9 puts 3-4 times the uniform weight on the hair-color words, and by step 29 that same Block has dropped below one fifth of uniform.
Blocks 0-1 and 20-27, on the other hand, stay at roughly half of uniform from start to finish and barely move.
The 0.50 in the later Blocks seems to be the value you get when these two words are hardly read at all. What they read instead is not something this measurement captures.

For the spatial maps I cut out Blocks 4, 8, 12, 16, 20 at steps 0, 5, 10, 20, 29 and normalized each panel by its own maximum (brightness is not comparable across panels; the figure is for shape only).

E1 spatial map of cross-attention on the word blonde. Rows are steps, columns are Blocks, each panel normalized by its own maximum

E1 spatial map of cross-attention on the word black. Rows are steps, columns are Blocks, each panel normalized by its own maximum

E1 spatial maps of blonde and black at Block 12 laid out over steps 0-29

In the blonde figure, Block 4 brightens along the shape of the left girl’s hair from step 5 on, and Block 8 also shows the left girl’s outline from step 5. Blocks 12 and 16 concentrate the weight on a few points with no surface-like shape, and Block 20 is nearly flat across the grid.
Block 4’s left-half share is around 0.6, a small bias in the table, yet the spatial map shows the left girl’s hair shape. Weight also spreads faintly onto the right girl, which is what pulls the share down.
What token positions the point-like concentrations in Blocks 12 and 16 correspond to, this figure cannot tell.

E2 Residual stream difference

I generated A and B at the same seed 42 and saved every Block’s output (the residual stream, shape 1x52x76x2048) at steps 0, 1, 2, 5, 10, 15, 20, 29 in fp32. 6.8 GB per run.
I subtracted A’s and B’s outputs at the same step and Block, took the L2 norm per token, and put it back on the 52x76 grid.
At step 0 the initial latent is identical, so the difference comes purely from the conditioning. From step 1 on the latent trajectories also diverge and that gets mixed in.

E2 token-mean L2 of the A minus B residual stream difference. Rows are saved steps, columns are DiT Blocks

E2 spatial map of the A minus B residual stream difference. Rows are steps, columns are Blocks. Color bars differ per panel

E2 left/right asymmetry of the A minus B residual stream difference. Rows are saved steps, columns are DiT Blocks. Positive means larger difference in the left half, negative in the right half

The token-mean L2 at step 0 grew with Block depth.

BlockL2 mean at step 0
00.08
118
427
856
9116
12330
16542
20546
24613
271416

Block 0’s output is nearly the same for A and B, then it jumps to 18 at Block 1. It rose sharply from 170 to 330 between Blocks 11 and 12, and from 946 to 1416 between Blocks 26 and 27.
The size of the residual stream itself presumably differs per Block too, and here I only show the absolute difference.

In the spatial maps, the step 0 difference in Blocks 0-8 is grid-wide noise with no shape, while Blocks 24 and 27 already show two blobs at the two heads at that same step 0.
By steps 1-2, Blocks 0-8 also show blobs at the two heads, and from step 5 on the difference is simply the shape of the two girls’ hair and outlines.
At every Block and step the difference appears on both girls; no panel concentrates it on one side. A and B swap both hair colors, so the difference showed up on both sides exactly as the swap implies.
In Blocks 12, 16, 20, only a few tokens near the top edge of the grid carry values of 100,000-200,000 and the rest goes dark. These are the tokens that exceeded the fp16 range in the harness check, and presumably the reason ComfyUI keeps the residual stream in fp32.

Left/right asymmetry is the left-half mean minus the right-half mean, divided by the overall mean, and it stayed within plus or minus 0.4 at every Block and step.
Block 9 at step 0 is -0.38, right-leaning; Blocks 12-20 at steps 5-20 are -0.16 to -0.37, right-leaning; the same Blocks at step 2 are +0.11 to +0.22, left-leaning; and step 29 is near 0 everywhere.
The Blocks 12-20 values include the top-edge outlier tokens described above.

The difference alone cannot say which Block decides the binding, so E3 checks by swapping.

E3 conditioning swap

While generating A, I swapped in B’s conditioning for cross-attention only in the specified range. Seed 42, cond side only.
A swap across all Blocks and all steps matches plain B bit for bit, confirmed in the harness check.
Judgment was done by opening each image and writing down the hair-color arrangement as seen, with the mean absolute pixel difference (RGB, 0-255) against plain A (E0’s seed 42) and plain B as a supporting number. The pixel difference between plain A and plain B is 65.49.

E3a Block axis

Only a Block group is swapped, across all steps 0-29.

Swapped rangeHair arrangementChanges besides hair colorPixel diff vs plain APixel diff vs plain B
Blocks 0-3Still ABoth girls’ eyes turned black (A’s left is blue-green, right is brown)5.9865.14
Blocks 4-7Still ACollar became a vertically ribbed stand collar with buttons, faces got longer11.7062.52
Blocks 8-11Still ATighter face crop with large black eyes, left top became a pale purple button shirt17.0462.49
Blocks 12-15Flipped to BBoth in black sleeveless turtlenecks with bare shoulders, right hair a short bob, tight upper-body crop59.0438.06
Blocks 16-19Still ACollar changed from frills to a buckled high collar6.4864.28
Blocks 20-23Still AMetal detail added at the collar4.2365.87
Blocks 24-27Still AAlmost the same3.1165.63
Blocks 0-13Still ABoth in sleeveless ribbed turtlenecks with bare shoulders, right hair a shoulder-length bob, left top gray-purple, tight crop30.7659.40
Blocks 14-27Flipped to BBoth in black sleeveless turtlenecks with bare shoulders, right hair a short bob, tight crop59.7738.35

E3 result of swapping only Blocks 0-3 to B's conditioning. Hair arrangement still A

E3 result of swapping only Blocks 4-7 to B's conditioning. Hair arrangement still A

E3 result of swapping only Blocks 8-11 to B's conditioning. Hair arrangement still A

E3 result of swapping only Blocks 12-15 to B's conditioning. Hair arrangement flipped to B

E3 result of swapping only Blocks 16-19 to B's conditioning. Hair arrangement still A

E3 result of swapping only Blocks 20-23 to B's conditioning. Hair arrangement still A

E3 result of swapping only Blocks 24-27 to B's conditioning. Hair arrangement still A

E3 result of swapping only Blocks 0-13 to B's conditioning. Hair arrangement still A

E3 result of swapping only Blocks 14-27 to B's conditioning. Hair arrangement flipped to B

The hair arrangement flipped in two cases, Blocks 12-15 and Blocks 14-27. Both sit at a pixel difference of 59 from plain A and 38 from plain B, a different picture that is neither A nor B. The two flipped images also changed clothes, hairstyle, and composition the same way.
Swapping Blocks 0-13 did not flip, yet clothes and composition changed the same way as the two flipped ones (pixel difference 30.76).
That is the only run that includes Blocks 12-13 without flipping, so whether 14-15 are included out of Blocks 12-15 looks like the dividing line. But there are only two ways the Block groups were cut, and I did not try swapping Blocks 14-15 alone.

Swapping Blocks 0-3, 16-19, 20-23, 24-27 left the hair arrangement untouched with pixel differences of around 3-6. Only details like eye color and collar trim changed.
Blocks 4-7 and 8-11 have pixel differences of 12-17, with changes in face crop and tops, but the hair arrangement stays A.

In E1 the strongest left/right bias was in Blocks 8-16, and in E3a the flip came from swapping Blocks 12-15 out of those.
Swapping Blocks 8-11, where E1’s bias is also strong, did not flip, so in this experiment the strength of the cross-attention bias and whether swapping that Block flips the hair colors do not line up.

E3b step axis

Only a step range is swapped, across all Blocks.

Swapped rangeHair arrangementChanges besides hair colorPixel diff vs plain APixel diff vs plain B
steps 0-4Flipped to BSame flat line-art as B, brown cardigan with pale purple frilled collar, right eye gray. Slightly wider shoulder framing than B65.3410.11
steps 5-9Still ACollar changed to a buckled high collar5.7766.15
steps 10-19Still AAlmost the same2.2365.56
steps 20-29Still AAlmost the same0.7565.46
steps 0-9Flipped to BNearly identical to plain B65.601.90
steps 10-29Still AAlmost the same2.2965.55

E3 result of swapping only steps 0-4 to B's conditioning. Hair arrangement flipped to B

E3 result of swapping only steps 5-9 to B's conditioning. Hair arrangement still A

E3 result of swapping only steps 10-19 to B's conditioning. Hair arrangement still A

E3 result of swapping only steps 20-29 to B's conditioning. Hair arrangement still A

E3 result of swapping only steps 0-9 to B's conditioning. Hair arrangement flipped to B

E3 result of swapping only steps 10-29 to B's conditioning. Hair arrangement still A

Swapping only steps 0-4 to B pulled the hair arrangement, clothes, and line-art texture all toward plain B (pixel difference 10.11 from plain B). Steps 0-9 is nearly identical to plain B (1.90).
Conversely, swapping any of steps 5-9, 10-19, 20-29, or 10-29 stays close to plain A, with pixel differences of 3 or less apart from the collar change at steps 5-9.
Steps 5-9 is the same length as steps 0-4, yet the result stayed plain A.

This fits the trend measured in E1, where the cross-attention bias weakens with steps and the hair-color words are hardly read in later steps. Still, E1’s bias had not fully dropped to 0.5 even at the last step, so the later-step reading is not exactly zero either.

E4 Left/right self-attention block

E0 produced no seed where the hair-color binding broke, so I first added clothing colors to the hair colors as prompt C and generated 3 seeds plain.

LabelPrompt
C2girls, standing side by side, upper body, looking at viewer, simple background. The girl on the left has blonde hair and wears a red dress. The girl on the right has black hair and wears a blue dress.
seedCountLeft hairLeft clothesRight hairRight clothesMatches
422blonderedblacknavy4/4
432blonderedblackblue4/4
442blonderedblackblue4/4

E4 plain generation of prompt C at seed 42 (instructed: left blonde in red, right black hair in blue)

E4 plain generation of prompt C at seed 43 (instructed: left blonde in red, right black hair in blue)

E4 plain generation of prompt C at seed 44 (instructed: left blonde in red, right black hair in blue)

All three follow the instruction on all four items. Seed 42’s right dress is navy rather than blue, but I counted it as a match in the blue family.
Nothing broke under this condition either, so there is no way to compare whether the block improves binding.
Even so, the block would still show how the picture changes, so I put the self-attention mask in on the same 3 seeds.
The mask makes the image-token self-attention read only the same side, left or right. Cross-attention and MLP are unchanged. The implementation is the left/right split attention described in the test environment.

In all 6 images the leftmost girl is blonde in red and the rightmost has black hair; only the rightmost dress splits between navy and blue.

Mask rangeseedCountRightmost dressDifference from plainPixel diff vs plain
All Blocks424navyVertical seam down the center. Left half is flat line-art on white, right half is shaded painting, a different style. The two on the left have arms and torsos left white, arms cut off midway86.75
All Blocks434navyCenter seam. Left half shaded, right half flat. The two on the right have no arms, just torso silhouettes, with thin necks not joined to the bodies99.88
All Blocks444navyCenter seam. Left half white background, right half gray background with thick lines, a different style. The right girl on the right half has blue-green mixed into black hair88.94
Blocks 8-16424blueNo seam, continuous style. Two girls became four, the blonde-red and black-blue pair repeated twice78.25
Blocks 8-16434blueNo seam. Four girls, same pair twice. The second from left has a skin-colored chest not joined to her clothes79.19
Blocks 8-16443blueNo seam. Three girls; the middle one has blonde bangs and black at the back of the head, in blue71.60

E4 result of the left/right self-attention block on all Blocks, seed 42

E4 result of the left/right self-attention block on all Blocks, seed 43

E4 result of the left/right self-attention block on all Blocks, seed 44

E4 result of the left/right self-attention block on Blocks 8-16, seed 42

E4 result of the left/right self-attention block on Blocks 8-16, seed 43

E4 result of the left/right self-attention block on Blocks 8-16, seed 44

Blocking on all Blocks put a vertical seam down the center of all three images, and the left and right halves became independent pictures.
Each half contains both the blonde-red and the black-blue girl, four in total. Style and background brightness differ between the halves.
When the halves cannot see each other, each half apparently draws the whole prompt on its own.
Blocking only Blocks 8-16 removes the seam and the style is continuous as one picture, but the head count rose to 4, 4, and 3. Seed 44 has one girl whose hair splits into blonde and black.

E5 Token-level rewrite

E3 swapped the whole conditioning, so more than hair color changed.
This time I swapped only the token vectors for blonde (position 25) and black (position 34) in the conditioning with those from another prompt. Besides B with the swap, I prepared R with red hair on the left.

LabelPromptSource positions
BA with left/right hair colors swappedposition 25 = black, position 34 = blonde
RA with blonde replaced by red (The girl on the left has red hair.)position 25 = red
IDSourcePositions swappedRange
E5-1B25 and 34all Blocks, all steps
E5-2B25 and 34Blocks 12-15, all steps
E5-3B25 and 34all Blocks, steps 0-4
E5-4R25all Blocks, all steps
E5-5R25Blocks 12-15, all steps
E5-6R25all Blocks, steps 0-4

Seed 42, everything else as in E0. I added one plain generation of R for reference.

A, B, and R tokenize to 38 tokens with T5 and line up exactly: position 25 is blonde, black, red respectively, and position 34 is black, blonde, black.
The swap happens at the same stage as E3 (the 1x512x1024 conditioning reaching the DiT Blocks’ cross-attention), replacing only the specified rows.
Specifying all 512 positions matched E3’s whole swap bit for bit.

IDSourcePositionsRangeLeft hairRight hairClosest referenceChanges besides hair color (vs plain A)
E5-1B25 and 34all Blocks, all stepsblackblondeplain BPale purple frilled high collar, brown buttoned jacket, eye color and line weight all match plain B
E5-2B25 and 34Blocks 12-15blackblonde (bob)plain A (style, composition, skin). Clothes differBlack sleeveless ribbed turtleneck with shoulders and arms bare, no frilled collar, right in a bob
E5-3B25 and 34steps 0-4blackblondeplain BNearly the same as plain B. Right hair slightly longer than plain B
E5-4R25all Blocks, all stepsredblack (bob)plain RGray skin, high-contrast style in only red, black, and white, red and black high-neck tops. Pixel diff vs plain R 4.3
E5-5R25Blocks 12-15orange-leaning light brownblackplain A (style, composition, skin). Clothes differDusty pink and black buttoned high necks, no frilled collar, eyes brown to black. Pixel diff vs plain R 39.7
E5-6R25steps 0-4redblack (bob)plain RNearly the same as plain R (pixel diff 5.4). Purple and brown patches at the left neckline
plain Rnonenonenoneredblack (bob)(reference)Gray skin, high-contrast red, black, and white, red and black high-neck tops
IDplain A overalltop lefttop rightbottom leftbottom rightplain B overall
E5-164.364.765.669.058.09.5
E5-257.748.955.577.648.839.5
E5-364.264.365.768.857.912.3
E5-448.348.925.176.542.565.6
E5-518.717.312.332.712.457.7
E5-647.648.725.073.743.266.0
plain R47.749.125.776.039.965.9

Pixel differences are mean absolute differences vs plain A (RGB, 0-255), and the four quadrants split the frame in half both ways. Hair falls in the top row, clothes in the bottom row.

E5 E5-1. B's positions 25 and 34 swapped in all Blocks at all steps

E5 E5-2. B's positions 25 and 34 swapped in Blocks 12-15

E5 E5-3. B's positions 25 and 34 swapped at steps 0-4

E5 E5-4. R's position 25 swapped in all Blocks at all steps

E5 E5-5. R's position 25 swapped in Blocks 12-15

E5 E5-6. R's position 25 swapped at steps 0-4

E5 plain R. Prompt R generated plain at seed 42 as the reference

Swapping just two tokens to B across all Blocks and all steps turned the whole picture into plain B (pixel difference 9.5 from plain B). Steps 0-4 alone does the same (12.3). At seed 42, the difference between the A and B pictures was almost entirely decided by these two vectors out of 512.
Even with the positions narrowed, using the full range leaves nothing of the original A: clothes, collar, and line texture all switched to B’s along with the hair color.
Narrowing to Blocks 12-15, the hair color changed according to the source (flipped for B, orange-leaning for R) while face, composition, and style stayed plain A. The pixel difference vs plain A is 18.7 for R; by quadrant, the top row where the hair is gives 17 and 12, and the bottom left where the clothes are gives 33.
But those clothes are neither A’s, B’s, nor R’s. All three Blocks 12-15 swaps (E3a’s whole conditioning, E5-2, E5-5) lost the frilled collar and got a turtleneck or high neck instead.
Swapping R’s position 25 in all Blocks gave nearly plain R (4.3), and there the skin turned gray as well and the style became high-contrast red, black, and white. The single red token seems to change the color design of the whole frame.
With R limited to Blocks 12-15, the hair never reached a vivid red and stopped at an orange-leaning light brown.

E6 Single Block, mix ratio, spatial limit

E5 could not isolate hair color by cutting on either Block or step, so I tried the remaining conditions.

IDAxisWhatRuns
E6-aSingle BlockR’s position 25, and B’s positions 25 and 34, each put into one of Blocks 12, 13, 14, 158
E6-bMix ratioR’s position 25 in all Blocks at all steps, mixed with the original vector at ratios 0.25 and 0.52
E6-cSpatial limitThe swapped token affects cross-attention only for a specified image-token region. Outside the region, the output computed with the original conditioning is used4

E6-c regions are defined on the 52x76 token grid. “Left half” is columns 0-37; “top left” is columns 0-37 and rows 0-29 (the area holding the left girl’s hair in E1’s spatial map).

IDSourcePosition and regionRange
E6-c1R25 in the left halfall Blocks, all steps
E6-c2R25 in the top leftall Blocks, all steps
E6-c3R25 in the left halfBlocks 12-15
E6-c4B25 in the left half, 34 in the right halfall Blocks, all steps

Seed 42, everything else as in E0. Judgment and pixel differences as in E5, with the overall difference vs plain R (E5’s reference) added.

The spatial limit is implemented by computing cross-attention twice at any Block and step that has region-tagged tokens, once with the unswapped conditioning and once with the swapped one, and picking per image token depending on whether it falls in the region.
Tokens outside the region use the unswapped output as is. Self-attention and MLP are unchanged.
Mix ratios 1.0 and 0.0 matched the ordinary swap and plain A bit for bit, respectively.

IDSwapRangeLeft hairRight hairClosest referenceChanges besides hair color (vs plain A)
E6-aR 25Block 12blonde (slightly ochre)blackplain AEyes turned brown, collar a vertically ribbed high neck with a button row (same on both)
E6-aR 25Block 13blondeblackplain AEyes unchanged. Collar a buttoned stand collar (same on both)
E6-aR 25Block 14blonde (orange-leaning yellow)blackplain AEyes brown, ribbed high neck with a button row, the right girl shifted left with shoulders touching
E6-aR 25Block 15blondeblackplain AEyes unchanged. Collar a buttoned stand collar
E6-aB 25,34Block 12blondeblackplain AEyes black, left top white-to-gray vertical stripes with a gray frilled collar
E6-aB 25,34Block 13blondeblackplain AAlmost the same
E6-aB 25,34Block 14blonde (dull pale yellow)blackplain AEyes dark brown, left top gray-purple vertical rib with a button row, the right girl shifted left with shoulders touching
E6-aB 25,34Block 15blondeblackplain AEyes black. Clothes almost the same
E6-bR 25, ratio 0.25all Blocks, all stepsblonde (orange-leaning yellow)blackplain AEyes brown, high neck with a button row, the right girl shifted left
E6-bR 25, ratio 0.5all Blocks, all stepsorange (between blonde and red)blackcomposition plain A, clothes and rendering toward plain RRed on the left, black on the right, high necks, dull gray-brown skin, thick lines and flat fill
E6-c1R 25 in left halfall Blocks, all stepsredblackplain RWhole frame in a different style. Gray skin, high-contrast red, black, and white, right in a bob
E6-c2R 25 in top leftall Blocks, all stepsredblackplain RWhole frame in a different style. Left in a red V-neck, right in a thick ribbed turtleneck, right in a bob, rough lines
E6-c3R 25 in left halfBlocks 12-15orange (bright)blackplain AEyes brown, left pale red-brown and right black high necks with a button row, the right girl shifted left
E6-c4B 25 in left half, 34 in right halfall Blocks, all stepsblackblondeplain BWhole frame in a different style. Flat fill with thick black lines, brown cardigan with pale purple frilled collar, right in a bob
IDplain A overalltop lefttop rightbottom leftbottom rightplain B overallplain R overall
E6-a R 25 Block 1212.38.711.018.311.263.144.4
E6-a R 25 Block 135.14.52.99.53.765.747.3
E6-a R 25 Block 1414.913.811.023.311.759.841.3
E6-a R 25 Block 154.02.92.07.93.466.047.4
E6-a B 25,34 Block 129.46.19.111.810.565.244.9
E6-a B 25,34 Block 133.21.73.35.62.465.747.2
E6-a B 25,34 Block 1417.012.413.629.412.658.945.2
E6-a B 25,34 Block 156.04.25.58.56.065.345.9
E6-b ratio 0.2514.610.912.123.312.161.641.7
E6-b ratio 0.526.926.515.249.616.157.833.2
E6-c147.248.424.875.340.565.75.6
E6-c239.746.623.060.329.059.725.2
E6-c318.617.412.332.312.457.839.7
E6-c470.167.871.771.669.318.169.7

E6 E6-a. R 25 swapped in Block 12

E6 E6-a. R 25 swapped in Block 13

E6 E6-a. R 25 swapped in Block 14

E6 E6-a. R 25 swapped in Block 15

E6 E6-a. B 25,34 swapped in Block 12

E6 E6-a. B 25,34 swapped in Block 13

E6 E6-a. B 25,34 swapped in Block 14

E6 E6-a. B 25,34 swapped in Block 15

E6 E6-b. R 25 at mix ratio 0.25, all Blocks and all steps

E6 E6-b. R 25 at mix ratio 0.5, all Blocks and all steps

E6 E6-c1. R 25 in the left half, all Blocks and all steps

E6 E6-c2. R 25 in the top left, all Blocks and all steps

E6 E6-c3. R 25 in the left half, Blocks 12-15

E6 E6-c4. B 25 in the left half and 34 in the right half, all Blocks and all steps

With a single Block (E6-a), none of the four flipped the hair arrangement or turned it red. Blocks 13 and 15 stay nearly plain A at pixel differences of 3-6, while Blocks 12 and 14 change eye color and collar. The Blocks 12-15 range that flipped in E3a and E5 only flips when all four are swapped together.
The mix ratio (E6-b) changes color continuously: orange-leaning yellow at 0.25, a strongly reddish orange at 0.5, with face and composition still plain A. At 0.5, though, clothing color and skin rendering had also moved toward plain R, and there was no stage where only the hair changed.
The spatial limit (E6-c) could not isolate hair color either. Letting only the left-half image tokens read position 25 still turned the whole frame into plain R (pixel difference 5.6 from plain R). Narrowing to the top left still lands on the R side (25.2). The Blocks 12-15 version came out almost the same as E5-5 without a region, at 18.6 vs 18.7.
Even with the cross-attention output held in place outside the region, it spread over the whole frame through that Block’s self-attention and MLP and the Blocks after it. The left/right cross-referencing seen in E4 is what carries the swap’s influence across the frame here.
In the images with swapped hair-color tokens, the collar that was asymmetric in plain A (pale pink frills on the left, black frills on the right) was replaced every time by matching high necks with a button row on both. E5 did the same, and it is common to every swap that passes through Blocks 12-15.

Across Block, step, space, and mix ratio, no condition that rewrites hair color alone while keeping the original picture turned up within this harness.
Any swap passing through Blocks 12-15, where the flip happens, always changed the clothes too, and cutting in space did not stop self-attention from spreading it. I think the hair-color token is being handled together with the clothing information there.
To cut it, self-attention would have to be cut by region as well as cross-attention, and in E4 doing that split the picture into two halves.

E7 Rewrite the residual stream directly

Swapping the conditioning changed the clothes along with it, so I left the conditioning entirely alone and moved to rewriting the residual inside the DiT directly.
The GPT-OSS paper from the opening split the hidden state into role and content and rewrote only the content. On the image side I took the direction for the left girl’s hair color out of the Block output hidden state and added it to the hair tokens only.

The direction is built by running A and R on the same seed, saving each Block’s output at each step, and taking R minus A over the tokens of the hair mask.
The hair mask comes from plain A (seed 42) by thresholding blonde pixels by color and pooling to the 52x76 grid.
Two ways of adding: “mean direction”, which adds one averaged vector over the region, and “per token”, which adds each token’s own difference as is.

IDDirectionBlocks addedRegionCoefficientMethod
E7-1R - A12-15left hair1.0mean direction
E7-2R - A12-15top left (columns 0-37, rows 0-29)1.0mean direction
E7-3R - A8-19left hair1.0mean direction
E7-4R - A12-15left hair2.0mean direction
E7-5R - A12-15left hair1.0per token
E7-6R - A12-15right hair (left hair region mirrored)1.0mean direction
E7-7B - A12-15left hair1.0mean direction

E7-6 is the control: does the right girl turn red when the direction meant for the left girl is added to her hair.
Seed 42, all steps, cond side only, everything else as in E0. Judgment and pixel differences as in E5, also split into inside and outside the hair mask.

The first seven runs saturated, so I added coefficients 0.1 and 0.3, a version with the 20 largest-magnitude dims of the difference vector zeroed, and combinations of both.

IDDirectionBlocksRegionCoefficientDims removed
E7-8R - A12-15left hair0.1none
E7-9R - A12-15left hair0.3none
E7-10R - A12-15left hair1.0top 20 zeroed
E7-11R - A12-15left hair0.1top 20 zeroed
E7-12R - A12-15left hair0.3top 20 zeroed
E7-13R - A12-15left hair0.05top 20 zeroed
E7-14R - A12-15left hair0.1top 102 (5%) zeroed

E7 hair mask and size of the difference vector

Blonde pixels from the left half of plain A (R>150, G>110, B<140, R-B>40), with a token counted as hair at 25% or more coverage per 16x16 pixel cell. 562 of the 3952 tokens, within rows 4-51 and columns 10-32. A few tokens also reach into the top of the frilled collar.

E7 hair mask overlaid on plain A. Scaled up as is at token-grid resolution

The residuals were re-saved for A, R, and B over Blocks 8-19 at all 30 steps in fp32 (11.7 GB each).
The table shows the norm of the mean vector inside the hair mask and the mean residual norm over the same region.

BlockstepDifference vector normResidual normRatioShare of norm in top 20 dims
1208097060.0080.80
1215149487900.1700.80
1307697810.0080.70
1315149691250.1640.79
140103106460.0100.71
1415161999350.1630.83
150156108920.0140.84
15151527108000.1410.79

The step 15 difference vector is about 19 times the step 0 one, at 14-17% of the residual.
Seventy to eighty percent of that norm sits in 20 of the 2048 dims, so the difference vector is decided less by a direction than by the size of a few dims. Whether these 20 dims are the same as the huge-value token dims found in E2 is something I did not check. In the flat-color S2 earlier, the 20 dims holding the color difference barely overlapped the huge-value dims.
Zeroing the top 20 dims brought the norm down to 0.55-0.71 times.

E7 results

IDDirectionBlocksRegionCoefficientDims removedLeft girlRight girlDoes the mask shape show
E7-1R-A mean12-15left hair1.0noneWhere the hair was is filled with a white slab, red blocks at the edge and inside. Face gray and blurred, features unreadableBob length, gray skin, black high neckyes
E7-2R-A mean12-15top left1.0noneHead covered by a flat orange rectangle, no hair visible. Clothes pinkBob length, gray skin, black high neckyes
E7-3R-A mean8-19left hair1.0noneWhere the hair was is a light blue blocky mass. Face hidden. Whole frame in pink noiseDark brown wavy medium-length hair, red spots on the faceyes
E7-4R-A mean12-15left hair2.0noneWhere the hair was is a light blue blocky mass. Whole frame in pink noiseGray-brown hair, red and pink spots on the faceyes
E7-5R-A per token12-15left hair1.0noneWhere the hair was is light blue with only the hair strands left as lines. Face in light blue and red stripesBob length, gray skin, black high neckyes
E7-6R-A mean12-15right hair1.0noneHair still blonde, saturation up. Eyes brown, salmon-colored high neckWhere the hair was is white and gray grainy noise, face gray with no featuresyes
E7-7B-A mean12-15left hair1.0noneAll black (NaN in the residual)Same-
E7-8R-A mean12-15left hair0.1noneHair a flat red-to-orange fill, bang strands gone. Face outline and expression remain, eyes dark red. Clothes a dark red ribbed high neckLong straight hair and bangs kept, dark red at the collar, a green band at the right edge of the hairslightly
E7-9R-A mean12-15left hair0.3noneHair orange-to-red in blocky shading. Face a gray-white surface with features goneBob length, gray-white skinyes
E7-10R-A mean12-15left hair1.0top 20 zeroedHair a beige halftone pattern with red patches. Face a gray surfaceBob length, gray-white skin, no frilled collaryes
E7-11R-A mean12-15left hair0.1top 20 zeroedHair orange-to-red with strands kept. Face as in plain A, eyes from blue-green to brown. Pale pink frilled collar kept with a slight red tintLong black hair, black jacket with frilled collar unchangedtop edge slightly squared
E7-12R-A mean12-15left hair0.3top 20 zeroedHair a flat red fill with vertical stripes, face reduced to feature lines with gray-white skin. Clothes also a red surfaceBob length, gray-white skin, black high neckyes
E7-13R-A mean12-15left hair0.05top 20 zeroedHair orange with strands kept. Face as in plain A, eyes brown. Pale pink frilled collar keptLong black hair, black jacket with frilled collar unchangedno
E7-14R-A mean12-15left hair0.1top 102 zeroedHair still blonde with saturation barely up. Eyes black. Otherwise nearly plain ANearly plain Ano
IDplain A overalltop lefttop rightbottom leftbottom rightin hairoutside hairplain R overall
E7-142.647.125.355.442.791.534.551.2
E7-252.552.545.464.847.443.354.055.3
E7-347.348.033.156.052.391.040.177.0
E7-449.343.639.951.861.882.543.778.7
E7-542.544.829.551.344.483.835.648.4
E7-647.224.763.246.155.049.546.952.3
E7-7163.8214.0156.1186.299.1156.3165.1129.4
E7-838.348.117.170.617.4104.927.326.7
E7-945.040.623.575.240.976.439.826.0
E7-1035.221.823.050.145.833.935.434.1
E7-1131.240.615.353.215.693.020.931.2
E7-1241.246.018.975.324.693.132.625.1
E7-1322.327.113.434.913.760.116.036.8
E7-148.99.96.413.26.023.06.543.6

Pixel differences are mean absolute differences vs plain A (RGB, 0-255). In hair and outside hair refer to the hair mask mapped back to pixels.

E7 E7-1. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 1.0 (no dims removed)

E7 E7-2. R-A mean direction added to the top-left region in Blocks 12-15 at coefficient 1.0 (no dims removed)

E7 E7-3. R-A mean direction added to the left hair region in Blocks 8-19 at coefficient 1.0 (no dims removed)

E7 E7-4. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 2.0 (no dims removed)

E7 E7-5. R-A per-token direction added to the left hair region in Blocks 12-15 at coefficient 1.0 (no dims removed)

E7 E7-6. R-A mean direction added to the right hair region in Blocks 12-15 at coefficient 1.0 (no dims removed)

E7 E7-7. B-A mean direction added to the left hair region in Blocks 12-15 at coefficient 1.0 (no dims removed)

E7 E7-8. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 0.1 (no dims removed)

E7 E7-9. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 0.3 (no dims removed)

E7 E7-10. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 1.0 (top 20 dims zeroed)

E7 E7-11. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 0.1 (top 20 dims zeroed)

E7 E7-12. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 0.3 (top 20 dims zeroed)

E7 E7-13. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 0.05 (top 20 dims zeroed)

E7 E7-14. R-A mean direction added to the left hair region in Blocks 12-15 at coefficient 0.1 (top 102 dims zeroed)

The mean direction at coefficient 1.0 (E7-1) saturated. The hair mask region became a white slab with the token grid’s jaggedness showing at the outline. Adding the 20 dims that make up most of the difference vector to every token in the hair mask blew up the residual.
A rectangular region (E7-2), wider Blocks (E7-3), a larger coefficient (E7-4), and per-token (E7-5) each broke in a different way, and none of them drew anything that reads as hair.
The control on the right hair (E7-6) turned the right hair region into grainy noise while the left hair stayed blonde with higher saturation. A direction made for the left still wrecks the right when added there.
The B minus A direction (E7-7) gave NaN in the residual and an all-black image. The left/right swap difference vector could not be used the same way as the color-change one.

Lowering the coefficient to 0.1 (E7-8) made the hair red and kept the face outline, but the hair is a flat fill with the bang strands gone and the clothes turned into a dark red high neck.
Zeroing only the top 20 dims (E7-10) turned the saturation into a different kind of breakage (a halftone pattern).
With the top 20 dims zeroed and coefficient 0.1 (E7-11), only the left hair turned reddish while face, collar, and the right girl survived. The hair is orange-to-red with the strands kept, the pale pink frilled collar stays, and so do the right girl’s black hair and black jacket. What changed is the left girl’s eye color going brown and a slight red tint on the collar. Pixel difference is 31.2 from plain A, 20.9 outside the hair, and around 15 in the top-right and bottom-right quadrants.
Raising the coefficient to 0.3 under the same conditions (E7-12) saturated again.
Lowering it to 0.05 (E7-13) gives orange hair with the strands kept, and face, collar, and the right girl stay plain A. The outside-hair pixel difference is 16.0, the right-side quadrants 13-14, less change from plain A than E7-11, and the hair stops at orange short of red. The mask shape does not show.
Widening the removal to the top 5% (102 dims) at coefficient 0.1 (E7-14) left the hair blonde and brought the pixel difference vs plain A down to 8.9. The component that was changing the color also lived in dims outside the top 1% but inside the top 5%.

E7-11 and E7-13 gave what the conditioning route never did: rewriting the residual stream directly changed only the left hair’s color while face, collar, and the right girl survived.
E5-5 (conditioning side, limited to Blocks 12-15) had a pixel difference of 18.7 vs plain A with the clothes changed; E7-13 is 22.3 with the clothes intact. The difference lands in different places.
The window is very narrow, though. Coefficient 0.05-0.1, breakage at 0.3, the top 20 dims must be removed, and widening the removal to 102 dims stops the color from changing.
The direction itself is dominated by a few huge-value dims, and I think the remaining component only starts acting as color once those are removed, but whether that remainder represents hair color alone is something I did not check. The eyes turning brown is common to all four conditions, so hair and eyes may be changing through the same component.

Flat-color direction and the hair-color attribute

Bringing what the flat-color experiments showed back to the two-girl case.
Color is decided by denoise step 1. That held for red, blue, and the girl image. The red-blue residual difference is 85-94% concentrated in 20 of the 2048 dims, and those 20 are separate from the residual’s huge-value dims.
Adding the difference to Blocks 24-27 changed the color, and one of them at steps 0-1 alone was enough. Adding to the middle Blocks (8-23) blows up the residual. Adding to only the left half in the last Block group at steps 0-1 split the image into blue on the left and red on the right at a one-pixel-wide boundary.
Adding the same direction to a girl’s hair produced a blue surface in the exact shape of the hair, and the girl was redrawn. The direction I built from the flat-color difference seems to act by repainting a region as a surface.

Part of the reason hair color could not be isolated in the two-girl case is probably here: the color direction in the residual was “what color to make this region as a surface”, not “what color to make this hair”.
With this extraction method, no stable direction that changes hair color alone while keeping the object’s shape turned up. E7-11 and E7-13 are the two images that survived under the narrow condition of coefficient 0.05-0.1 with the top 20 dims removed.