Tested Qwen3-Embedding and Qdrant memory on a StackChan voice chat server. Ramen and guitar were recalled, but weekend plans hovered right at the 0.45 threshold.
Tested on CoreS3 + StackChan Body: head taps caused ghost recordings and phantom Qwen3-ASR replies. Fixed across 3 tiers with VAD, emotion tags, and lip sync.
Tested on CoreS3 + StackChan Body: the 3-zone head pad works as a talk button, strokes need a 30–400 ms gap between zones, and servo motion leaves touch stuck until recalibrate().
Tested on an RTX 3050 Ti Laptop: Qwen3.8-Omni-Flash handles audio input directly, shrinking our voice chat server RSS from 1,327MB to 118MB with zero idle warmup lag.
Qwen3-Embedding-0.6B on CPU vs ModelScope's API on a Ryzen 7 5800HS voice server. Local embedding adds up to +2.6s per voice turn by competing for CPU; the API adds none.
A pre-implementation design for adding Qwen3-Embedding and Qdrant memory to an RTX 3050 Ti and CoreS3 voice-chat stack without depending on StackChanWorld's API. It keeps raw logs, searchable memory, persona, and body separate.
Qwen3.8 restored 625k chars of mangled PDF text; 5 Irodori-TTS workers on one RTX 4090 gave 38h of audio at 13x realtime. Plus an 18x reasoning_effort win and the cuInit 999 check for dead RunPod hosts.
Qwen3-ASR-0.6B STT on CPU, Qwen via ModelScope, streaming TTS on 4GB VRAM — one laptop, 11.9s voice-to-voice. Filler-audio job polling and measured timelines.
Tested Irodori-TTS 500M-v3 on a 4GB RTX 3050 Ti Laptop (Windows): default-voice and zero-shot clone samples, real timings, and the FFmpeg DLL fix for MP3 references.
Tested ZONOS2 on an 8GB RTX 4060 Laptop (WSL2): the 15.3GB bf16 weights run via Windows system-memory fallback, a KV-cache override, and the CUDA toolkit at ~1/20 realtime. Plus a Japanese name-accent gotcha with A/B audio.
VoxCPM2 sits in the tokenizer-free corner. Mapped vs F5-TTS, CosyVoice2, Irodori-TTS, Style-Bert-VITS2; plus why Japanese TTS still leans on OpenJTalk.
NII/LLMC released CC Audio and Archive.org Audio Dataset. URL lists, metadata, and a downloader covering 48,000+ hours of Japanese audio. What it actually contains and how it fits into TTS, ASR, and audio model training.