Fitted Anthropic's jacobian-lens on Qwen3-4B-Instruct-2507 (4090, 51 min, ~1 USD), then read a layer-swapped SFT corrector: outputs pass through while hidden states diverge to cos 0.88.
Tested arXiv 2607.01232's layer localization under SFT: Qwen3-4B depth 25/50/75% vs all-layer LoRA, trained on a RunPod 4090. The eval-loss U-shape is real; the rewrites disagree.
RL gains sit at 40-60% depth: on Qwen3-8B, training only layer 16 beats full-parameter RL (67.1 vs 66.5). Notes on arXiv 2607.01232 and what it doesn't claim about efficiency.
Measured on M1 Max ComfyUI: QIE 2511's pixel shift comes from the encode node's forced 1MP rescale. Stock ReferenceLatent fixes it; expression, outfit and pose diffs tested.
Fine-tuned Qwen3-4B on 799 of my own edit pairs, quantized to a 2.3GB GGUF at 38 tok/s on an M4 mini. Eval loss looked fine, but it barely removed slop — and the real fix was feeding paragraphs, not single sentences.
Three LLMs converted the same 10 Japanese scene briefs into Anima (Qwen-DiT) prompts, generated as 60 fixed-seed images on an M1 Max with a merged 3-character LoRA. The Qwen-to-Qwen affinity hypothesis did not survive; a strict formatter brief with character-count locks is what actually moved the results, and two failure modes survive any prompt.
Tested on an M1 Max, NumPy only: Qwen maps a prompt to a JSON of knobs, and a 2D Kuramoto oscillator field renders it. No objects, but composition, color, and motion change with the prompt.
Merged kei, kana, and koharu into a single Anima (Qwen-DiT) LoRA and ran my first training on Blackwell (RTX 5090, sm_120). Hands-on log: the cu128 / torch2.8 / SDPA stack swap from the 4090, why the weakest character gets absorbed (caption asymmetry, not rank), and how trigger-only prompts separate three close-packed characters at ep143 without ControlNet.
Tested Qwen3.7 Plus on ModelScope: native function calling and parallel tool calls work. I built a tool loop, skills, and error recovery with just the openai SDK, then had it ship a working Flask BBS.
Tested Qwen3.7 Max and Plus proofreading a Japanese novel: both barely fix, split on quote punctuation and names, and the one 'typo' was a character name.
After a US order pulled Claude Fable 5, which Chinese models drop into Claude Code? Kimi K2.7 Code, Qwen3.7 Max, DeepSeek V4 and GLM-5.1 — constraints, VRAM, benchmark caveats.
rank128 + 20 two-character images killed the v1 ahoge bleed and body fusion on this Anima dual-character LoRA. Lap-sit stays a Qwen3 text-encoder limit; sweet spot is ep140.