Tested on M1 Max 64GB: the training-free TaylorSeer cache in Diffusers made Qwen-Image 2.1 about 2.7x faster. It crashes with the default KV cache; a small patch fixes it.
Benchmarked on an M4 Mac mini: Ben Joffe's 2-instruction weekday hack beats plain %7 by 1.6-6.4x in clang, Rust, and V8, loses 3x in CPython. Plus a 25x V8 -0 deopt trap.
Tested Next.js 16.3, Astro 7 and Nuxt 4.5 on an M4 Mac mini with one identical blog. Pixel-identical output ships 172KB vs 21.5KB vs 0KB gzipped JS; build cache and dev-server memory measured too.
Tested on release day: drizzle-orm 3.9x, Effect 5.3x, lobe-chat 7.3x, the tsconfig hard errors, why --checkers 8 backfired on 16GB, and why Vue can't use TS7 yet
UC Berkeley's RDI team demonstrated that major benchmarks including SWE-bench and WebArena can be manipulated to near-perfect scores without completing any tasks. They identified 7 vulnerability patterns and released BenchJack, an automated benchmark attack tool.
Only 10 of 40 layers use KV cache, so raising llama-server ctx-size from 4096 to 65536 cost 800MB VRAM and no throughput. Measured on Ryzen AI Max+ 395.
François Chollet et al. publish new benchmark ARC-AGI-3. As of March 2026, all Frontier LLMs have achieved less than 1% of the interactive task of autonomously exploring an unknown environment with an unknown goal.
Setup notes for running WAI-Illustrious SDXL v16 on ComfyUI with an 8GB RTX 4060 Laptop. 1024x1024 generates in ~15 seconds without --lowvram, and a LoRA still loads. CUDA 12.8 portable build and path gotchas included.
Anthropic accused three Chinese AI companies of distilling Claude, and on the same day OpenAI retired SWE-bench Verified. Training fraud and evaluation flaws exposed simultaneously on February 23, 2026.
Using IBM and UC Berkeley's IT-Bench benchmark and the MAST failure taxonomy, this article examines why enterprise AI agents fail. It covers the reality of 11% SRE success and 0% FinOps success, plus the Replit production database deletion incident.