Qwen3-ASR-0.6B STT on CPU, Qwen via ModelScope, streaming TTS on 4GB VRAM — one laptop, 11.9s voice-to-voice. Filler-audio job polling and measured timelines.
NII/LLMC released CC Audio and Archive.org Audio Dataset. URL lists, metadata, and a downloader covering 48,000+ hours of Japanese audio. What it actually contains and how it fits into TTS, ASR, and audio model training.