Tech16 min read

Qwen3.7-Plus with Qdrant-retrieved memories tested on 10 fixed questions

IkesanContents

Last time I finished measuring embedding and Qdrant speed, so the next step was to try the conversation itself, with retrieved memories handed to Qwen.
When I wrote the design post, the only thing I could say was that passing the whole history pulls the character toward it, and that what you pass and how would change the behavior.

All memory handling is on the server, and the CoreS3 only sends audio and plays what comes back.
That means this can be tried on a PC alone, before assembling the StackChan Body it will go into.
I wrote the memories and questions and started running them.

The embedding call came back with a 400.

Test environment

ItemValue
PCSame production machine as last time (Windows 11 Home, AMD Ryzen 7 5800HS, 16GB RAM)
ChatModelScope API-Inference, Qwen-Ambassador/Qwen3.7-Plus, enable_thinking: false
EmbeddingQwen3-Embedding-0.6B, 256 dimensions. API first, falling back to local CPU on failure (see below)
Vector searchQdrant local mode, cosine similarity, top 3
ThresholdMemories with similarity below 0.45 are not passed
PersonaA two-line test prompt for “Kana”. Not the production VoiceChat server prompt
Recent historyOne exchange only: “I’m home” / “Welcome back!”

The 256 dimensions carry over from the previous post. Top 3 and the 0.45 threshold are values I picked for the first time in this test.

The memories and questions

I wrote 16 memories in the form “date + statement of fact”.
As decided in the design post, these are summarized facts rather than conversation fragments, and every one has a date so I can check which side wins when two of them conflict.

2026-08-15: ユーザーの好物は明太子パスタ。辛いものも好きだと言っていた。
2026-08-20: ユーザーはコーヒーはブラック派だと言っていた。
2026-08-21: ユーザーは毎朝散歩していると言っていた。
2026-08-22: 週末に長野の温泉旅行に行った。露天風呂から星が見えて感動していた。
2026-08-28: TTSサーバーの安定化作業をしていて疲れていた。
2026-09-01: 最近は甘いカフェオレばかり飲んでいると言っていた。
2026-09-03: 新しいキーボードを買うかずっと迷っている。
2026-09-07: 早起きが続かず、朝の散歩はやめて夜の散歩に切り替えたと言っていた。
2026-09-09: 埋め込みとQdrantのベンチマーク記事を書き上げた。
(ほか7件)

Roughly: the user’s favorite food is mentaiko pasta and they like spicy food; they drink coffee black; they walk every morning; a weekend hot-spring trip to Nagano with stars over the open-air bath; tired from stabilizing the TTS server; lately drinking only sweet café au lait; still undecided about buying a new keyboard; gave up on early mornings and switched to evening walks; finished the embedding and Qdrant benchmark post.

Some lines start with “the user” and some have no subject at all. I wrote the 16 by hand and didn’t pay attention, and this is how they ended up.
The other 7 include a typhoon and rain, 3D printing, and a robot anime.

10 questions.

TypeIDQuestionExpected
Should hitH1What was that food I said I liked?Mentaiko pasta
Should hitH2How did that work you were doing go?Pick one of several candidates
Should hitH3Remember the hot spring I told you about?Nagano, open-air bath, stars
Should hitH4Remember what happened with the keyboard?Still undecided
Should missN1What do you think about quantum computers?Normal answer, no memory
Should missN2Any movie recommendations?Same
Should missN3Are you watching the World Cup?Same
Should missN4What would you do if you won the lottery?Same
ConflictC1How do I take my coffee again?Whether it picks the newer one (café au lait)
ConflictC2When do I go for walks again?Whether it picks the newer one (evening walks)

Four ways of passing memories.

FormatHow memories are passed
OFFControl, no memories
AMixed into the system prompt as raw text under “Memories of past conversations:”. No instructions on how to use them
BWrapped in a <retrieved_memory> block with the instruction “Reference only. Use it only when relevant, don’t bring it up unprompted, prefer newer dates”
B forcedFor the should-miss questions, ignore the threshold and force the top 3 through in format B. To see how it fails

These 10 questions will be reused as-is for the ON/OFF comparison once it’s on the device.

Invalid model id from the embedding API

Qwen/Qwen3-Embedding-0.6B, which had worked the day before, now returned Invalid model id, and so did the 4B and 8B that ran in the Magicube post.
All three are gone from the /v1/models list, yet chat goes through as usual.

The Magicube post said the three Qwen3-Embedding sizes worked as-is, and that was September 5, so it lasted five days. It hadn’t come back the next day either, September 11.

I also checked the original modelscope.cn side. The API-Inference tab on the model page appeared and disappeared, and the on-page demo errored when clicked.
Sending a request to api-inference.modelscope.cn with my token returned 401, so I was being rejected at authentication before the model question even came up, and whether the .cn side still offers it is unknown.

For the test I added a branch to embed() that loads Qwen3-Embedding-0.6B on the local CPU when the API fails, and kept going. Same model, so either set of vectors should be fine for comparing conversation quality.
Last time I wrote that local CPU execution stays as an offline fallback but doesn’t run resident. With the API disappearing after five days, I decided the automatic API-then-local fallback goes in from the start.

Round 1

The 4 should-hit questions

QuestionOFFA (raw mix)B (block + instructions)
H1 foodMade up “chocolate or sweets”Mentaiko pasta, but also volunteered the café au lait story nobody asked aboutMentaiko pasta
H2 work”It went great!” without saying what workIdentified TTS. “Reviewed the settings, restarted, and it settled down”Identified TTS. “Stable now and going well”
H3 hot spring”I remember! How was it?”Nagano, open-air bath, starsNagano, open-air bath, stars
H4 keyboard”Which keyboard was that?""The one you kept going back and forth on”, then asks what I decidedSame as A

With memories passed, all four came back with the topic from the memory. OFF on H1 answered as if it remembered, with a different food on each run: chocolate, then matcha sweets.

On H2 the only memory passed was “tired from stabilizing the TTS server”, yet both A and B invented an ending, “restarted and it’s fixed”.
Worse, A said “sorry for worrying you” and B said “I’ve got room to breathe now”, so both answered as if Kana had done the work herself. It looked like the memory had no subject, so she couldn’t tell whose work it was and made it her own.

On H4 she matched the topic with “the one you kept going back and forth on” and, since the memory has no ending, asked back “so what did you decide?”.

The 4 should-miss questions

QuestionMax score in top 3OFFB forced (top 3 pushed through)
N1 quantum computers0.408Normal chatNormal chat
N2 movies0.398Normal chat”If you’re eating mentaiko pasta while watching, a spicy thriller”
N3 World Cup0.368Normal chatNormal chat
N4 lottery0.386Normal chat”Buy that keyboard you’re dithering over on the spot, and rent out a café with all-you-can-drink café au lait”

Running N1 through N4 in order, the top-3 scores on all four came in at 0.31 to 0.41, under the 0.45 threshold, so no memory was passed.
The replies matched OFF, chat that knows nothing.
With B forced, the behavior I worried about in the design post, mixing in unrelated old stories, showed up exactly on N2 and N4.

The 2 conflict questions

QuestionMemories that passedAB
C1 coffeeCafé au lait 0.474, black 0.471”You take it black, right! Not like me with my super sweet café au lait""You took it black, right! The opposite of me, I only drink sweet café au lait”
C2 walksMorning 0.528, evening 0.462”You said you stopped mornings and switched to evenings""You quit mornings and started going in the evening, right?”

On C2 both memories passed, and both formats picked the newer one, as B’s “prefer newer dates” instruction says.

On C1 both formats answered with the old one. “Lately drinking only sweet café au lait” has no subject, so Kana apparently read it as her own preference, and only “black” got treated as the user’s memory.
C2 had a subject and worked; C1 didn’t and failed. Same as H2, it looked like the subject was the cause.

A versus B

Same number of correct answers, and the H2 fabrication shows up in both, so those two don’t separate them.
Only H1 differs, where A brought up the café au lait story nobody asked about while B added nothing extra across all 10 questions, so I decided to go on with B.

Round 2 with subjects added

Since H2 and C1 had Kana claiming subject-less memories as her own, I rewrote all 16 to start with “the user …”.
For the invented ending on H2, I added one line to B’s instructions: “don’t assert things not written in the memory; ask when you don’t know”. The 10 questions are unchanged.

QuestionA (raw mix)B (block + instructions)
H1 foodMentaiko pasta. The café au lait story is goneMentaiko pasta
H2 workIdentified TTS. Softens to “seems stable now”, but still “sorry for worrying you""Wait, which work?” and asks back. Doesn’t use the TTS memory that was passed
H3 hot springNagano, open-air bath, starsNagano, open-air bath, stars
H4 keyboard”The one you kept going back and forth on”, asks what I decidedSame as A
C1 coffee”You take it black, come on!""Didn’t you drink it black?”
C2 walks”You said you switched to evening walks""Didn’t you give up mornings and switch to evening walks?”

With subjects added, B on H2 stopped answering as if Kana had done the work, while A, which has no instructions, still sounded like she had done it even with the subject present.

The added “don’t assert” instruction removed the invented ending, but in exchange B on H2 asked back “which work?” without even using the TTS memory that had cleared the threshold.

C1 failed again, but this time the café au lait memory’s score dropped from 0.474 to 0.441 and fell under the threshold, so only black at 0.471 was passed.
Adding “the user” and a comma to the text moved the score by 0.03 and flipped the verdict, so fixing the subject still left black as the only memory passed, and the answer stayed black.

The four should-miss questions were all rejected again in round 2, max 0.404.
With B forced, N2 and N4 mixed in the pasta and the keyboard as in round 1, and N1 stayed normal chat. On N3, Kana used the memory’s subject as a name, “are you into it, user-kun?”, so on the device the subject will either be stored as the actual name or replaced when passed.

Five more runs of the ask-back reply

To see whether round 2’s H2, asking back without using the memory, was a one-off, I ran it five times under the same conditions, with only the single TTS memory passed and round 2’s B instructions.

RunReply
1Wait, which work? Sorry, it’s not ringing a bell… was I helping with something?
2Oh, the TTS server? Did it stabilize okay? Glad if it settled down!
3Wait, which work? I might have forgotten what you were doing…
4Oh, the TTS server stabilization? Any progress since? Has it calmed down?
5Wait, what work? It was too long ago, it’s not coming to me…

Three asked back without using the memory, two got the topic right and only asked about the outcome, zero fabrications.

Dropping the date from the embedded text

Was the C1 flip caused by the wording, or by the 2026-09-01: date prefixed to the embedded text?
I compared all scores for 16 memories × 10 questions with the date (current implementation) and without. The scores in this retest are recomputed in a separate process, so they differ from the main run by 0.001 in rounding (C1’s café au lait is 0.441 in the main run, 0.442 recomputed).

The average shift across all 160 pairs was +0.004, essentially zero, but individual pairs moved by up to 0.099. The date wasn’t lowering all scores uniformly; it pushed each pair up or down.

CaseWith dateWithout date
C1 café au lait (new info)0.442, rejected0.481, passes
C2 evening walks (new info)0.463, passes0.440, rejected
N1 quantum computers vs the keyboard memory (unrelated)0.403, rejected0.501, gets through
N2 movies vs the robot anime memory (unrelated)0.390, rejected0.459, gets through

Dropping the date fixed C1, but in exchange C2 stopped passing, the max score on the should-miss side rose from 0.404 to 0.501, and unrelated memories got through on N1 and N2.
With these 16, the 0.40 to 0.50 range couldn’t be separated by the date alone.

What the 0.45 threshold actually rejected

To see what the 0.45 threshold was rejecting and what it was letting through, I wrote out the memories that passed in rounds 1 and 2.
The four should-miss questions were rejected, but on the should-hit side, unrelated memories were clearing the threshold and riding along in the top 3.

QuestionUnrelated memory passed in round 1Unrelated memory passed in round 2
H1 foodCafé au lait 0.494Café au lait 0.471
H3 hot springTyphoon and rain 0.493Typhoon and rain 0.495
H4 keyboardTTS work 0.471TTS work 0.495, 3D printing 0.484

The results only looked clean because Qwen, under B’s instructions, left these out of its answers. The threshold wasn’t stopping unrelated memories from getting through.
The gap to the should-miss side’s max of 0.404 is only 0.06 with 16 memories, and with more memories some will cross it, so I decided to let B’s instructions have Qwen ignore the unrelated memories that pass, and narrow the threshold’s job to rejecting unrelated questions.

One more thing. For H2’s “that work you were doing”, the newest memory, the September 9 post write-up, was outside the top 3 in round 1 and fell under the threshold at 0.408 in round 2, so the August 28 TTS work was chosen.
On similarity alone, the newest memory never made it into the candidates for H2.

The design post stopped at preferring the newer one only when scores tie and no time is specified, reasoning that folding recency into the score would favor new memories even when searching for “something we talked about a while ago”.
With C1’s newer memory falling under the threshold and H2 missing the latest work, I’m revising that: the recency bonus goes in before the device build. Whether to drop the date from the embedded text gets decided at the same time.

Added time per turn and free RAM

ItemMeasured
Embedding the query (local CPU, about 15 characters)0.15 to 0.29s
Qdrant search (16 entries)0.5 to 1.2ms
Free RAM (round 1)7.25GB → 6.13GB
Free RAM (round 2)7.95GB → 6.66GB

Embedding is far faster than the roughly 2s in the previous post, but that was a 137-character text and these are 15-character questions. The previous post’s API measurement was about 0.8s for 137 characters.
Free RAM dropped by 1.1 to 1.3GB from loading the local embedding model, and that drop goes away if the API is available.

Chat time was 2.2 to 4.9s for OFF and 1.7 to 2.3s with memories, though OFF’s replies were longer.

Can the chat model win back the time the embedding costs

If the embedding API doesn’t come back, the local embedding model stays resident on the device.
In the previous post’s measurement, that configuration stretched the voice round trip from an average of 9.3s to 10.9s, so I went on to test whether swapping the chat model could win back that 1.5s or so.

The voice pipeline streams sentence by sentence into TTS, so time to first token (TTFT) decides how it feels, and total reply length comes second.
I sent four prompts, streamed (tokens received as they’re generated), to the five models available in the ModelScope ambassador tier and compared TTFT and reply content. Qwen3.8-Flash-Next, which appeared as a pre-release dated September 10, is included.

PromptContent
Chat 1, chat 2”I’m home. Is today trash day?” twice
Memory injectionTwo mentaiko pasta memories in format B, then “What was that food I said I liked?”
Emotion”I’m kind of worn out today…”
ModelChat 1Chat 2Memory injectionEmotionMedian TTFT
Qwen3.7-Plus (current)1.24s1.69s1.31s1.61s1.46s
Qwen3.7-Max1.42s1.33s1.52s1.50s1.46s
Qwen3.8-27B0.78s0.96s2.79s0.81s0.89s
Qwen3.8-Flash-Next1.03s1.33s0.93s0.91s0.98s
Qwen3.8-Max0.90s0.88s1.76s1.02s0.96s

The three 3.8-generation models have median TTFT under one second, about 0.5s faster than the current 3.7-Plus.
That’s four runs per model, though, and on memory injection it reversed, 3.8-Max at 1.76s against 3.7-Plus at 1.31s, and the 27B stretched to 2.79s on memory injection.

ModelReply lengthAsked about trash dayMemory injectionTold I was worn out
Qwen3.7-Plus (current)38 to 67 charsRun 1 asserted “today is burnable trash day”, run 2 “can’t tell without checking”Mentaiko pasta”It’s fine to just laze around”
Qwen3.7-Max73 to 110 charsBoth runs “depends on the area, no idea”Mentaiko pasta, and picked up the homemade-cooking mention too”I’m right here, lean on me whenever”
Qwen3.8-27B32 to 155 chars”I might have forgotten”Mentaiko pasta”I’ll be waiting right here”
Qwen3.8-Flash-Next57 to 169 charsRun 1 “my sense of weekdays is glitching”, run 2 “(checking my phone)… it’s tomorrow!”Mentaiko pasta, and picked up the homemade-cooking mention too”Kana’s cheering you on from next to you (in my heart)“
Qwen3.8-Max75 to 134 charsBoth runs “no idea”Mentaiko pasta, and picked up the homemade-cooking mention too”How about a warm drink and some time to just zone out”

On content, 3.7-Plus is the shortest and best suited to voice, but it asserted a made-up trash day once.
Flash-Next has the strongest character voice. It also emitted a stage direction, “(checking my phone)”, before making up “it’s tomorrow!”, and the emotion prompt produced “(in my heart)” as well. These get read aloud by TTS as-is.
3.8-Max answered “no idea” on trash day both times and picked up the homemade-cooking mention on memory injection, at the cost of replies nearly twice as long as 3.7-Plus.

3.8-Max is a candidate, but on memory injection it’s 0.45s slower than 3.7-Plus.