Tested Qwen3-Embedding and Qdrant memory on a StackChan voice chat server. Ramen and guitar were recalled, but weekend plans hovered right at the 0.45 threshold.
Tested on CoreS3 + StackChan Body: head taps caused ghost recordings and phantom Qwen3-ASR replies. Fixed across 3 tiers with VAD, emotion tags, and lip sync.
Tested on an RTX 3050 Ti Laptop: Qwen3.8-Omni-Flash handles audio input directly, shrinking our voice chat server RSS from 1,327MB to 118MB with zero idle warmup lag.
Qwen3-Embedding-0.6B on CPU vs ModelScope's API on a Ryzen 7 5800HS voice server. Local embedding adds up to +2.6s per voice turn by competing for CPU; the API adds none.
A pre-implementation design for adding Qwen3-Embedding and Qdrant memory to an RTX 3050 Ti and CoreS3 voice-chat stack without depending on StackChanWorld's API. It keeps raw logs, searchable memory, persona, and body separate.
Qwen3-ASR-0.6B STT on CPU, Qwen via ModelScope, streaming TTS on 4GB VRAM — one laptop, 11.9s voice-to-voice. Filler-audio job polling and measured timelines.
NII/LLMC released CC Audio and Archive.org Audio Dataset. URL lists, metadata, and a downloader covering 48,000+ hours of Japanese audio. What it actually contains and how it fits into TTS, ASR, and audio model training.