Tech10 min read

StackChan AI Agent Notifications: Claude Code Hooks and I2S Buffer Fix

IkesanContents

In the previous post, I connected a CO2 sensor to StackChan so that indoor air quality readings could be referenced in voice chat responses.
When delegating development tasks to AI agents such as Claude Code, Codex, or Antigravity while working on something else, you cannot tell whether an agent has finished its task or is waiting for user approval without checking the terminal.

To solve this, I set up a notification pipeline using the agent hooks—which run external commands upon events—to have StackChan announce completed tasks aloud.
On the voice server side, I added endpoints that receive notifications from these hooks, summarize the message into 1–2 sentences using Qwen3.8-Max, synthesize Japanese voice via TTS, and hold the output.
StackChan polls for new notifications during its idle state and speaks them in Kana’s voice.

Test Environment

ItemEnvironment Details
DeviceM5Stack CoreS3 (ESP32-S3), StackChan Body
FirmwareESP-IDF 5.5.4, arduino-esp32 3.3.10, M5Unified 0.2.17, StackChan-BSP 1.1.0
Server ConnectionDirect connection via Tailscale (no relay)
Voice ServerSeparate PC: LLM Qwen3.8-Max (ModelScope), TTS speech synthesis
Notification AgentsClaude Code, Codex, Antigravity (all running on a Windows PC)

Modifications to the voice server were handled by another Claude instance running on that PC.
What I built here was the StackChan firmware logic and the hook configuration for each AI agent.

Notification Flow

flowchart LR
  A["AI Agent<br/>Hook"] -->|POST /notify| B["Voice Server<br/>Qwen Summary & TTS"]
  B -->|GET /notifications| C["StackChan<br/>Fetch while idle"]

Two endpoints were added on the server side: one to accept hook payloads, and one for the device to poll notifications.

EndpointDescription
POST /notifyAccepts any payload format—whether it is raw JSON from Claude Code’s hook or plain text. The payload is passed directly to Qwen3.8-Max to summarize who did what project and what finished (or what confirmation is pending) into 1–2 sentences in Kana’s tone, assigned an emotion parameter, and sent to TTS.
GET /notifications?after=<id>Returns completed voice notifications in chronological order, including text, emotion, and audio URL. Passing the highest returned ID into the subsequent after parameter enables sequential retrieval without duplicates.

The POST /notify request returns immediately, and subsequent summarization and speech synthesis take 7–15 seconds.
Notifications are stored in server memory (retaining up to the last 50 entries); restarting the server resets the ID counter to 1.

StackChan Firmware Implementation

StackChan polls for notifications every 7 seconds only while idle—meaning when it is neither recording, conversing, nor playing audio.
When a new notification arrives, it downloads the 16 kHz audio file, sets its facial expression and servo movement to match the notification emotion, and begins playback.
The head motion logic from the previous voice chat implementation was reused directly.

ScenarioBehavior
Immediately after bootRecords the latest ID only; backlog notifications accumulated while powered off are skipped
Multiple queued itemsPlays them sequentially. The next item is retrieved only after the first finishes playback
Server restartIf the latest ID is smaller than the device’s tracked value, polling resets from ID 0
Pressing the Key Unit during playbackAborts current playback and transitions directly into speech recording mode
Server unreachableAborts connection attempts after 1.5 seconds

The 1.5-second connection timeout was added after hardware testing.
During server restarts, polling had previously stalled for 5 seconds per attempt, freezing facial eye blinks because the default socket connection timeout remained at 5 seconds despite configuring an HTTP response timeout of 3 seconds.

Hook Configurations for Each AI

All three AI agents support execution hooks that run commands upon task completion. These hooks transmit completion context to the voice server, which passes the text to the Qwen API.

AITrigger EventMethod
Claude CodeTurn completed, waiting for command permissionCalls curl from hook, sending JSON payload directly
CodexGoal completion onlyCalls a lightweight Python script from hook
AntigravityTask completion (excluding subagents)Calls a lightweight Python script from hook

Claude Code

Two hook events, Stop (response finished) and Notification (notification emitted), were registered in settings.json:

"hooks": {
  "Stop": [
    {"hooks": [{"type": "command",
      "command": "curl -s -m 5 -X POST \"http://<voice-server>/notify?agent=claude-code\" --data-binary @- >/dev/null 2>&1 || true"}]}
  ],
  "Notification": [
    {"matcher": "permission_prompt",
      "hooks": [{"type": "command",
      "command": "curl -s -m 5 -X POST \"http://<voice-server>/notify?agent=claude-code\" --data-binary @- >/dev/null 2>&1 || true"}]}
  ]
}

Initially, Notification had no matcher.
As a result, leaving the machine unattended after a task finished caused Claude Code to emit an idle reminder approximately 60 seconds later, prompting StackChan to speak again. Adding permission_prompt restricted notifications strictly to permission prompts.

When having the running Claude Code edit this configuration, its automatic safety check blocked the modification as “attempting to exfiltrate conversation data to an external server.” The settings were instead populated using Codex.

Codex

Initially, notifications were sent after every turn using Codex’s notify configuration.
However, writing blog articles involves launching Codex in read-only mode repeatedly for source verification.
Having StackChan speak on every turn proved disruptive, so this was changed to a hook that triggers strictly when a Goal completes.

The script flags turns containing successful goal-completion tool calls and fires a notification only if the flag is active when the final response ends. Running normal inquiry turns confirmed that notifications are suppressed properly.

Antigravity

Antigravity’s hook does not pass response JSON via stdin like Claude Code.
Instead, a helper script reads the final assistant message directly from the transcript.jsonl log file provided in the hook arguments.

Because Antigravity spawns multiple concurrent subagents during tasks such as style checks, firing on every subagent would trigger continuous speech. The script filters out notifications under the following conditions:

Suppressed ConditionDetection Method
SubagentsInitial conversation input is not a direct user instruction
Style-check scansInitial prompt contains style-check definition filenames or response matches review templates
Rapid sequential completionsCompleted within 20 seconds of the previous transmission
Manual suppressionEnvironment variable or mute sentinel file exists

Sending the raw agent string antigravity caused the TTS engine to struggle with English pronunciation in Japanese summaries. Changing the identifier to katakana アンチグラビティ resolved the issue.

Notification Latency

Before wiring up hooks, I sent synthetic JSON payloads matching Claude Code’s schema from my PC to verify StackChan’s behavior.

Payload ContentTime to Device Pick-upSpoken NotificationEmotionAudio Length
Turn completion (firmware flash finished)10.5 sクロードコードがCO2Monitorのプロジェクトで、スタックチャンに通知を読み上げる処理を追加して書き込みまで終わらせたよ。smile9.0 s
Permission prompt9.1 sクロードコードがリルティングチャンネルウェブでコマンド実行の許可を待ってるよfun5.4 s
Turn completion (sent concurrently with above)15.6 sクロードコードがLiltingChannelWebで、スタックチャンの通知記事に実機テスト結果を追記して一区切りついたよ。smile8.5 s

Latency from transmission to pickup combines server summarization/TTS processing time with StackChan’s 7-second polling interval.
When two notifications were sent simultaneously, StackChan fetched the second item 0.2 seconds after finishing the first, playing both in sequence.
Single polling requests took 68–354 ms, while downloading the 16 kHz audio file (~280 KB) took around 1.2 seconds.

Audio Clipping and Stutter Troubleshooting

While speech worked, initial feedback noted that words were intelligible but the voice sounded harsh and clipped. Stutter had been observed slightly the day before, but the first notification on this day was particularly degraded.

Audio Peak Analysis

Measuring the notification audio recordings on a PC revealed that average volume (RMS, Root Mean Square) was comparable to conversational filler phrases.
The difference was peak amplitude: notification audio consistently peaked near the digital maximum of 0 dBFS (where waveform clipping occurs).

Audio SamplePeakAverage (RMS)
Sample 1 (Audibly clipped)0.0 dBFS-16.8 dBFS
Sample 2 (Mild distortion)-0.6 dBFS-16.5 dBFS
Sample 3 (Mild distortion)0.0 dBFS-16.6 dBFS
15 filler audio clips-6.2 to 0.0 dBFS-15.2 to -16.5 dBFS

Because the neck servos move during playback, I tested whether servo current noise affected audio quality.

ConditionAudio Quality
Disable neck motionNoticeably clipped
Enable neck motionSlightly improved, but still clipped
Cut servo torqueClipped in the first half; clear in the second half

Mechanical motion was not the primary factor.
Inspecting 0.5-second slices showed that the first 6 seconds peaked between -0.1 and -1.5 dBFS, whereas the latter half dropped to -4 to -9 dBFS, precisely aligning with where distortion was heard.

To prevent clipping, the firmware was modified to attenuate audio data if downloaded peaks exceed -6 dBFS.
Because normal voice chat replies share the same download path, this attenuation also prevents distortion during conversations.
Testing the same clip with attenuation confirmed that distortion vanished, though overall volume dropped slightly, which was balanced by raising CoreS3 speaker volume from 3 to 5.

Inconsistent Audio Quality Across Runs

Repeating playback after raising the volume showed noticeable variations even with identical audio files:

ConditionObserved Sound
Volume 4Stuttering
Volume 5 (immediately after)Clean
Volume 5, new notificationSeverely clipped
Volume 5, same notification repeatedModerately clean
Volume 5, repeated once moreModerately clean

The first playback suffered severe distortion, whereas subsequent replays of the same file sounded acceptable.
This indicated that degradation depended on runtime state and resource contention on the device rather than the audio stream itself.

Additional tests—filtering frequencies below 250 Hz before normalizing to -6 dBFS, and disabling display rendering during playback—produced negligible difference. Peaks remained at 0 dBFS after low-cut filtering, confirming that energy was not concentrated in low frequencies.

Fixing I2S Buffer Underruns

Complaints that longer utterances suffered crackling and stutter prompted me to reconsider the root cause: the issue was not purely analog clipping, but digital playback underrun.

M5Unified handles speaker output using a dedicated FreeRTOS task that streams audio chunks over I2S (the digital audio bus connecting the MCU to the audio DAC/amplifier).
The default and revised configurations are compared below:

ParameterDefaultRevised
Output Buffer256 samples × 8 (~43 ms at 48 kHz output)1024 samples × 8 (~170 ms)
Task Priority210
Target CPU CoreUnpinnedCore 1 (separate from Wi-Fi stack)

StackChan’s firmware concurrently executes Wi-Fi operations, Tailscale encryption via WireGuard, neck servo kinematics, and top-touch debouncing in the background.
If any background task monopolizes the CPU for more than ~43 ms, the audio task starves, producing audible gaps and popping sounds.
Updating the speaker settings to the revised values (expanding buffer capacity, increasing priority, and pinning to Core 1) eliminated playback dropouts, noticeably improving speech clarity.

The video below shows StackChan reading out a task-completion notification with these settings applied.

Work Logging and History

Voice notifications are easily missed if the user is away from their desk.
To prevent missed updates, the voice server parses incoming notifications and archives completed work items into persistent history.

MechanismDescription
Filtering completed tasksQwen determines whether the agent actually finished work, discarding pure question answers and permission prompts
Long-term memory aggregationWhen work from the same agent and project pauses for 30 minutes, it is consolidated into a single summary entry in external vector storage
7-day recent context injectionUp to 15 recent log lines are injected into the system prompt on every voice chat interaction

Injecting recent items directly into prompt context addresses an issue where vague queries like “What were you working on recently?” scored ~0.40 in vector similarity search, falling below the retrieval threshold of 0.45.
Injecting dummy work records on the server and inquiring via voice chat yielded accurate answers:

QueryKana’s Response
What were you working on recently?Recently, I added debounce handling for touch detection and implemented the notification endpoint.
What did you fix on StackChan?I added touch debouncing so the build and tests pass cleanly.
What should I have for lunch today?It’s almost 3 PM, so how about some ramen?

Unrelated questions did not trigger work log discussions, and Qwen correctly classified task completion across all six test cases.

The video below shows a voice chat where StackChan is asked what was worked on today.