Tested Qwen3-Embedding and Qdrant memory on a StackChan voice chat server. Ramen and guitar were recalled, but weekend plans hovered right at the 0.45 threshold.
Tested with 16 memories and 10 questions: a date bonus on cosine got the 9/9 memory and a 0.441 café-au-lait memory over the 0.45 cutoff, with 0.016 headroom left on hit-free questions.
Ryzen 7 5800HS test before StackChan: subject-less memories got claimed by the character, a 0.45 cosine threshold flipped on a comma, and ModelScope's embedding API vanished mid-test.
Qwen3-Embedding-0.6B on CPU vs ModelScope's API on a Ryzen 7 5800HS voice server. Local embedding adds up to +2.6s per voice turn by competing for CPU; the API adds none.
Tested Qwen3-Embedding-0.6B on CPU and Qdrant's local mode, then checked the 1.5-1.7GB RAM footprint against a Ryzen 7 5800HS voice server with only 4GB VRAM.
A pre-implementation design for adding Qwen3-Embedding and Qdrant memory to an RTX 3050 Ti and CoreS3 voice-chat stack without depending on StackChanWorld's API. It keeps raw logs, searchable memory, persona, and body separate.
Gemini API File Search now indexes images alongside text in the same store. Metadata filters can isolate NPC memories by chapter and character, and a single-character prototype costs under $1/month on Flash-Lite. Notes on tier limits, pricing breakdown, and what to test first.
158K lines of AI-generated C# for a Cities: Skylines II total conversion mod. CivicRAG for codebase indexing, 300+ custom Roslyn analyzers as compile-time design rules, and manual visual debugging for render bugs AI couldn't see.
Vektor Memory v1.5.4 supersession chains positioned against YourMemory decay, Cloudflare key-overwrite, and CTX, with a BM25 vs cosine threshold trap and a 5-field minimum schema for agent memory.
The paper argues that RAG, vector stores, and scratchpads are retrieval, not learning. Read alongside CTX and OCR-Memory, the gap between 'better search' and 'weight-level learning' becomes concrete.
A read of CTX, which auto-injects context into Claude Code via the UserPromptSubmit hook. Compared with auto-memory, YourMemory, WUPHF, and Cloudflare Agent Memory on persistence and storage. Also looked at why 1M context still isn't enough and how each agent architecture uses its window differently.
Hands-on log of building the DEV article's PDF RAG on M1 Max 64GB, extending it with images via CLIP, and pushing through Japanese with bge-m3 + Qwen3.6 35B. Documents the modality gap, the dual inference server crash, and LLM-jp 4-8B's empty chat template silently dropping the system role.