TechMar 23, 2026Flash-MoE: Running a 397B-parameter model on a 48GB MacBookFlash-MoE is a C/Metal inference engine that runs Qwen3.5-397B-A17B on a MacBook Pro M3 Max at 4.36 tokens/s. With expert streaming from SSD and hand-written Metal shaders, it fits the 209GB model into a 48GB memory budget.InferenceMPSLLMQwenMoELocal LLM