The Web Speech API + Gemini + VOICEVOX setup is complete — an AI character you can actually have a voice conversation with. Key implementation notes and impressions.
A comparison of major AI 3D generation tools such as TRELLIS, Hunyuan 3D, Tripo AI, and Hitem3D, with a focus on image requirements for better 3D output.
Setup notes for Qwen-Image-Edit-2511 on RunPod's RTX 4090 ($0.34/hr) using the ComfyUI template. Includes the fal Multiple-Angles LoRA (4 elevations × 8 azimuths × 3 distances) and a per-image cost breakdown that ends up cheaper than buying a 4090.
Technical prep for automating an implement → review → fix loop with Claude Code and OpenAI Codex via tmux. Can it build something overnight unattended?
VRAM per quantization for the 20B Qwen-Image-Edit-2511 in one table: FP8 from 6GB, NF4 16–20GB, BF16 24GB+, GGUF Q4 CPU-runnable. Minimum RTX 3060 12GB and recommended RTX 4070 Ti 16GB builds, plus the 96-angle Multi-Angle LoRA.
Emotion recognition used to mean fighting with old native libraries. Today there are cloud APIs and local libraries, but one major vendor has already left the field for ethical reasons.