A three-instruction leap-year test hides the constants 1073750999, 3221352463, and 126976. I rebuilt an 18-bit version with Z3, then traced how the multiply, mask, and threshold encode divisibility by 4, 100, and 400.
Benchmarked on an M4 Mac mini: Ben Joffe's 2-instruction weekday hack beats plain %7 by 1.6-6.4x in clang, Rust, and V8, loses 3x in CPython. Plus a 25x V8 -0 deopt trap.
Tested Q_rsqrt on Apple M4 (Mac mini) and Zen 3 (Ryzen 5800HS / WSL2). M4's -O2 already rewrites 1/sqrtf to frsqrte and ties Q_rsqrt; x86 clang needs -ffast-math or hits a 12x gap. Hand-written NEON/SSE wrappers turn out slower. Newton 0/1/2 error and the Lomont constant covered too.