Diffusers merged our fix for the Qwen-Image 2.1 TaylorSeer and KV cache collision in under 4 days. Notes on AI contribution guidelines, reproduction steps, and test design.
Tested on CoreS3 + StackChan Body: replaced 128x128 pixel art with 320x240 character art using Qwen-Image 2.1 diffs, layered in PSRAM for 5-emotion tap switching.
Tested on M1 Max: Diffusers stretches an 832×1216 reference to 832×1248, making figures 2.7% narrower. Setting output_resolution=1006 cut whole-image drift to about 0.1px.
Tested on M1 Max 64GB: the training-free TaylorSeer cache in Diffusers made Qwen-Image 2.1 about 2.7x faster. It crashes with the default KV cache; a small patch fixes it.
Tested Qwen-Image 2.1 open weights on M1 Max 64GB: overhead shots work locally, ~10 min per 832×1216 image (2x+ Anima), and i2i kept the face. Same prompts as my Anima/WAI tests.
Hands-on evaluation of Qwen-Image 2.1 early access on ModelScope Studio. Testing 10 camera composition presets, 4-member band role assignments, I2I orientation and outfit changes, and where multi-character consistency breaks down.
Tested on M1 Max 64GB with plain Anima-Base v1.0. A red-to-blue residual direction added to one of Blocks 24-27 flips a flat image, but on a hair mask it paints a blue slab. No hair-only direction found.
My ModelScope account switched from a monthly API-Inference quota to Magicube coins, so I measured what each call costs. Qwen-Image, Z-Image, Krea-2-Turbo and FLUX worked even though none of them appeared in the /v1/models list I got; of the embedding IDs I tried only Qwen3-Embedding went through; audio paths were 404 and Wan produced no video. In these tests 400s and 404s cost nothing, while a 200 with an empty body was charged.
Tested on M1 Max 64GB, ComfyUI v0.30.1: i2i and Qwen-Image-Edit lost the character, so I rebuilt the pose one prompt at a time, plus why negatives do nothing at cfg 1.0.
Tested on M1 Max 64GB, ComfyUI v0.33.3: from-behind and a back-row drummer now pass, out-of-frame crops get worse, and Anima-Base LoRAs need a 52-block key remap.
Tested on M1 Max 64GB, ComfyUI v0.33.3. Anima-2.9B inserts 12 DiT blocks, so Anima-Base LoRA keys hit the wrong layers with zero warnings, 0/3. Renumbering the keys gives 3/3.
Eight ComfyUI experiments on a 4-character Anima (Qwen-DiT) LoRA in plain terms: the conditioning right before the DiT decides who appears, and outfit mix-ups come from a biased DiT LoRA.