Introduction
When a maker decides to put a voice-first AI sidekick on a keychain the usual path means relying on cloud endpoints and a phone as the brain Bruce Li took a different route building Kai a pocket-sized AI charm that listens thinks and speaks from an M5StickS3 dev board the result is a device that balances a fast conversational path with a slower research-capable backend all while staying under 21 in parts and fitting on a lanyard
What Happened
The build started with a simple question how much of a cloud-native agent architecture can run on a 21 dev board without soldering custom PCB or dedicated firmware Kai runs on an M5StickS3 a compact module packing an ESP32-S3 a tiny display a MEMS mic and a modest speaker The device streams Opus-encoded audio over Wi-Fi to a Python relay on a home server which bridges to Gemini Live for sub-second answers and to a locked-down Hermes Agent for slower memory-backed tasks Every design decision from audio encoding to power domains stemmed from the constraint of a 250 mAh battery and the need for a device that can deep-sleep between interactions
- Audio pipeline Opus encoding decoding WebSocket streaming and a 2 MB PSRAM buffer smooth bursty playback
- Display 1.14 ST7789 panel showing pixel faces and five-line caption cards
- Power architecture Two core domains on the ESP32-S3 active cores at 240 MHz for encoding and an always-on RTC domain keeping time and Wi-Fi state at microamp levels
Why This Matters
Real-time voice on limited hardware forces trade-offs that matter to anyone building embedded AI The fast path Gemini Live answers in about two seconds but lacks built-in research capability The slow path Hermes Agent handles memory scheduling and multi-source lookups in 15–50 seconds Kai architecture keeps the device thin by offloading all heavy logic to a relay while the Stick itself stays deliberately dumb streaming audio rendering UI and sleeping deeply when idle
Power budgeting revealed that the idle screen-redraw loop and post-answer Wi-Fi hold time consumed most energy Simple fixes redrawing only on change letting the CPU idle at 80 MHz and gating the LCD audio rails in deep sleep projected battery life from 1.3 days to over six days at 20 questions daily Switching from raw PCM to Opus cut Wi-Fi airtime by more than 10× making the audio pipeline viable on a coin-cell budget
Key Takeaways
- Measure the tail The work is cheap the time after the answer often costs more Post-processing screen idle and Wi-Fi wind-down dominate power use
- Keep the brain in one place Offloading all tools memory and scheduling to a single Python relay means prompt tweaks ship without device reflashing
- Let the model route but verify Gemini Live decides fast vs slow but an eval suite and code-level guards prevent the model from promising work it wont do
- Async by default Acknowledge immediately deliver later into the conversation or an inbox and never block the user while background work completes
- Guard invariants in code Model behavior drifts relay-side checks ensure reminders watches and tool calls happen exactly as intended
Conclusion
Kai proves that a capable livable AI pocket device doesnt need a flagship budget or a custom ASIC With a 21 dev board a home server and two clearly defined brains it answers the questions a photographer and angler actually ask while hands are full and it does so with transparent routing measurable power use and fully open source code The projects real value lies in the lessons know your tail keep logic centralized and let the model route then code the invariants that make it reliable




Discussion
Join the conversation
Thoughtful reactions, questions, and follow-up ideas help shape the next story.