kit is a voice AI bot that runs entirely on a single M1 Max, for free, with zero external APIs. This article is the technical record of how I built kit as a parallel system. Five lines (listening, thinking, backchanneling, deep reasoning, speaking) run at once, and a single Arbiter collapses them into one verdict. It does not use the usual serial pipeline that listens, then thinks, then replies.
Why build it? Because I wanted a partner to talk with for a podcast. In (Podcast Log #1) I Want to Try a Podcast After Turning 40, I wrote that "talking alone is hard, so I'll build the tool, a voice to talk with, first." This is entry #2, the continuation, and kit is what's inside that tool. If #1 is "why I'm building it," #2 is "how I built it."
The starting point was simple: I wanted a bot I could talk to by voice, entirely on my own machine, for free. No external APIs. No monthly subscription. Nothing sent to the cloud. I speak into the mic, it gets transcribed, a local LLM thinks, and a synthesized voice answers. When the answer ends, it automatically goes back to listening, so it works as a partner for recording a one-person show. That is the "fully local, free, owned" voice bot I wanted.
In this article I give the real function names, the real config values, and the numbers I actually measured, without rounding. I also write plainly about what a single GPU cannot yet deliver. That honesty is where a technical article earns its value.
How is kit different from a normal voice bot?
The biggest difference is that kit does several things at once while the person is still speaking. A normal voice chat bot runs one straight line: the person finishes speaking, then transcribe, then think, then reply.
I could never make that straight line feel like talking to a person. When you talk with someone, they don't sit silently until you finish. They nod while listening, they backchannel, and they nudge you along when you trail off. In their head, they're already preparing a reply. It all runs at the same time.
So I dropped the single line and built kit as a parallel system. During one human utterance, these five things run in parallel:
- The ears (whisper) keep transcribing the in-progress speech with a sliding window
- The DSP keeps measuring the voice's volume (RMS) and pitch (true F0) every 80 ms
- When it finds a "phrase valley," it lays a colorless backchannel ("uh-huh," "ahh") over the voice
- As silence deepens, it gradually offers questions or new topics
- In the background, a large model is slowly simmering a deeper answer
That creates a problem. With five "eager talkers," several of them try to vocalize at the same instant and collide. The voice filling the silence, the backchannel laid over you, and the interrupting utterance all fight over the same beat.
So kit is designed so that the moment of vocalizing always passes through one point: the Arbiter (control thread). Even when all five systems say "I want to speak," only the Arbiter decides what happens at this exact moment. It returns exactly one of {STAY_SILENT | BACKCHANNEL | FACILITATE | SPEAK}. The details are under "Why is the decision to speak narrowed to a single Arbiter?"
What does the whole of kit look like?
kit is five lines (perception, control, timing, generation, output) spinning in parallel, converging on a single Arbiter only at the moment a sound goes out. One diagram shows it fastest.
Saying "it runs in parallel" was easy. Once I started writing code, I got lost myself until I drew, once and properly, where things branch and where they merge.
┌─────────────────────────────────────────────────────┐
│ Human (mic input) │
└───────────────────────┬─────────────────────────────┘
│ getUserMedia (1 stream・no double tap)
▼
┌──────────── single AnalyserNode (fftSize=2048) ───────────────┐
│ │
┌────┴─────┐ ┌──────────┐ ┌──────────────┐ ┌──────────┐ ┌────┴──────┐
│ ① Percept │ │ ② Control │ │ ③ Timing │ │ ④ Gen │ │ ⑤ Output │
│ P layer │ │ FSM │ │ reflex/bc │ │ 2-tier │ │ O record │
│ │ │ │ │ │ │ │ │ │
│ readRms │ │ IDLE │ │ bcTick 80ms │ │ Swallow │ │ raw mux │
│ readF0 │ │ ACQUIRING │ │ overlap bc │ │ fast ~8s │ │ AI voice │
│ partial │ │ LISTENING │ │ silence nudge│ │ + │ │ bc mix │
│ STT │ │ THINKING │ │ stage1/2/3 │ │ Gemma-4 │ │ → wav │
│ (whisper) │ │ NUDGING │ │ thinking fill│ │ deep │ │ → mp4 │
│ │ │ SPEAKING │ │ │ │ ~110-185s│ │ │
└────┬─────┘ └────┬─────┘ └──────┬───────┘ └────┬─────┘ └───────────┘
│ │ │ │
▼ ▼ ▼ ▼
┌────────────────────────────────────────────────────────────┐
│ ★ single Arbiter (control thread) arbitrate(kind, stage) │
│ look at muteMode → floor → governor, return 1 action │
│ {STAY_SILENT | BACKCHANNEL | FACILITATE | SPEAK} │
└────────────────────────────────────────────────────────────┘
│
▼
Server (body: HTTP I/O・serialize the single GPU)
The point is that ① through ⑤ all spin in parallel, but the moment a sound goes out, it always passes through the Arbiter. That structure prevents the three-way collision of "fill the silence," "lay a backchannel over you," and "interrupt and speak" fighting over the same instant. If only one place makes the decision, a collision can't happen in the first place.
Here is the hardware and the models. Everything runs locally, so this is the physical ceiling.
| Role | Implementation |
|---|---|
| Machine | M1 Max 64GB / 1 Metal GPU |
| Ears (speech→text) | whisper.cpp (whisper-cli) + ggml-small |
| Brain (fast voice LLM) | Llama-3.1-Swallow-8B-Instruct Q4_K_M |
| Brain (deep / text LLM) | Gemma-4-26B Q4_K |
| Mouth (text→speech) | AivisSpeech (VOICEVOX-compatible local HTTP) |
All numbers are real values confirmed from the actual config file (data/chat_config.json, about 400 keys), not guesses.
Why is the decision to speak narrowed to a single Arbiter?
Because letting each system decide on its own made the voices collide over and over. I gathered the decision into one function, arbitrate(), and focused on getting that one place right.
arbitrate() is kit's one and only decision point. It takes the "I want to speak" requests from each parallel system and returns exactly one verdict. Scatter "may I speak?" checks across the code with if-statements, and a gap always opens somewhere. Here is the actual priority order.
arbitrate(request):
1. Arbiter disabled → legacy path (kept to avoid regression)
2. muteMode (the human has muted it) → SILENT (first class)
3. non-overlap and floor not open to AI → SILENT
(a backchannel doesn't take the floor, so it's exempt from the floor gate)
4. non-overlap and stage >= FACILITATE → legacy path
(staged escalation isn't cut by budget = separation of duties)
5. overlap and governor-exempt → legacy path
6. governor on and this valley's AI utterances over budget → SILENT
7. legacy path (map temperature to 1 action)
Two "separations of duties of the same shape" are at work here.
The first is floor. The conversational floor (the right to speak) is held by the human by default. Only when the human has been silent for a while (floor_open_silence_ms=2200 or more) does the floor open to the AI. If the floor is closed, the AI can't speak. The backchannel is the one exception: it's exempt from the floor gate. A backchannel is a signal that may sound even while the human holds the floor, saying "I'm listening." Utterances that take the floor and backchannels that don't are treated as clearly separate things.
The second is the governor (valley budget). The "valley" is the gap between when the human finishes a stretch of talking and starts again. To keep the AI from rattling off fillers there, there's a budget on how many times the AI can speak in one valley (governor_valley_budget=2). Staged escalation into deep silence (stage2/3) is not counted against that budget. If it were, you'd get a secondary bug: one light reflexive backchannel uses up the budget, and the genuinely needed deep prompt never fires.
And crucially, choosing silence (SILENT) is an active decision, not a non-action. For kit, going quiet is a breath the Arbiter deliberately chose because staying quiet is the right call right now.
How does kit backchannel while listening?
Backchannels come from a "spinal reflex" layer on the front end, without waiting for the LLM. A function called bcTick() runs every 80 ms (bc_tick_ms=80) and lays a colorless backchannel over the "phrase valley" while the human is mid-speech.
Waiting on the LLM would be too slow. There's no time to call it if you want to lay a backchannel over someone mid-sentence.
How do you find the "phrase valley"? The trick is the true F0 (pitch). At first I hunted for valleys using only volume (RMS), but that wasn't enough. The moment to backchannel isn't a mere dip in volume. It's the moment the voice drops toward a clause boundary (low-pitch settling). When people head into closing a sentence, their pitch naturally falls, and that comes at a different timing than the volume dip.
So I measure pitch in Hz space and check whether it dropped below a recent baseline (f0_settle_ratio=0.92). When the pitch is uncertain (clarity below f0_clarity_min=0.45), kit doesn't fabricate a sloppy value. It treats it as 0. It doesn't pretend to measure what it can't.
The interesting part is that the backchannel fires probabilistically. Even when the conditions line up, it fires only with a probability of bc_fire_prob=0.12, and otherwise lets the moment pass. This avoids a fixed metronome. "Uh-huh, uh-huh, uh-huh" at perfectly regular intervals feels mechanical and, ironically, like it's not listening. Backchannels come alive precisely because they arrive at unpredictable timing.
On top of that, I layered a refractory window of bc_suppress_window_ms=4500 and a rate limit of at most 8 per 60 seconds. That keeps the effective rate sparse, about 7–8 per minute. Turning the firepower down creates the posture of someone who has stepped back to listen.
What supports this reflex layer is an FSM (finite state machine) that holds only one state. kit's states form this single road, and transitions happen only through setState().
IDLE → ACQUIRING → LISTENING → THINKING → NUDGING → SPEAKING
Two states are never held at once (I forbid flag combinations like isCharging && isStunned). If the state is unique, "what happens in this state" is fully predictable, and no gaps or dead zones appear. For example, the overlap backchannel spins only in LISTENING.
The playback path also has a triple guard of "generation counter + watchdog." It physically guarantees that playback happens exactly once, even if external async work (mic acquisition, playback end, fetch resolution) cuts in.
How does kit turn a 100-second lag into an asset? (two-tier thinking)
kit runs a fast shallow answer (Swallow 8B, about 8 seconds) and a slow deep answer (Gemma-4 26B, about 110–185 seconds) in parallel as two tiers. The light reply keeps the tempo alive first; the deep answer slips in when a moment to speak arrives. This is kit's most intellectual gimmick.
A large local model takes time to return a deep answer. In the production environment, Gemma-4 26B takes over 100 seconds for a single turn. By any normal reasoning, that's fatal. Stay silent for 100 seconds and the conversation dies.
So when the human finishes speaking, Swallow 8B first returns a "receiving beat" in about 8 seconds: an immediate "I'm listening." The tempo survives. In the background, Gemma-4 26B simmers a deeper answer for about 110–185 seconds. When the simmering is done and a moment it's OK to speak arrives, the deep answer slips in. That moment means three conditions hold: the human isn't talking, isn't trailing off, and the answer is about the latest topic.
This is also an information-theoretic gimmick. You place a light, low-expectation reply up front as a stepping stone. When the deep answer arrives later, the upside ("oh, it was actually thinking"), the reward prediction error, is maximized. Set expectations low, then exceed them.
The most important part: when a new utterance comes in, kit immediately cancels the deep computation and yields the GPU. Even mid-simmer, the instant the human starts saying the next thing, deep_cancel mercilessly stops it. One of kit's North Stars is "the human always wins." No matter how good an answer the AI is forming, the human's voice, interruption, or new utterance takes priority over any AI utterance, instantly. I built an AI that steps back the instant the human speaks, not one that grabs the GPU and won't let go.
The worker managing the deep tier has a single slot. When a new request comes in, it bumps a generation counter and discards the old computation, adopting only the latest generation's result. That prevents an answer to an old topic showing up late.
Where is the limit of running everything locally?
The M1 Max has only one Metal GPU. The ears (whisper), the brain (LLM), and the mouth (AivisSpeech) all use that one GPU, so they are mutually exclusive in time. True full parallelism, "listening and thinking deeply at the same time," is not realized, in principle. This is the section I should write about most honestly.
What looks parallel is the single GPU being switched finely. It isn't truly running simultaneously. Time-slicing is the only option.
To make this work, GPU access is serialized through one lock (_GPU_SPAWN_LOCK). The main voice conversation waits until the lock is free (it waits if the deep tier is running). The silence-filling nudge is best-effort: if the GPU is busy, it immediately returns a "GPU busy" error and gracefully gives up with an empty voice (again, the human wins). The deep tier always releases the lock in a finally block no matter which exit it takes, so the GPU never gets permanently jammed.
With that said, here are the honest limits.
| Limit | Detail |
|---|---|
| The deep tier gets canceled every turn and rarely lands | In long turns where the human keeps talking, the deep tier is discarded each time a new utterance comes in to yield the GPU, so the deep answer often doesn't complete. "Turning the 100-second lag into an asset" is a concept; in real use with frequent intermittent cancels, it still rarely lands. Tuning the content-reply (how the deep answer is delivered) is a work in progress. |
| Memory pressure makes the deep tier slow | As I'll cover below, a model that's fast in isolation slows down a lot with co-resident processes. |
| Overlap interruption only during recording | While the AI is mid-speech (SPEAKING), the analyzer for measuring the voice is torn down, so voice interruption doesn't work — only button interruption. You can interrupt by voice only within the recording window. |
The one I most want to be honest about is how measurement disproved my own hypothesis.
At first I assumed "the deep LLM is slow because of Gemma-4's MoE architecture." If the architecture is the problem, swapping to a simpler, smaller model should speed it up. But measured in isolation, Gemma-4 puts out about 28 tok/s. Plenty fast.
Re-measured in the production environment (with Swallow and AivisSpeech co-resident), it was about 4.7 tok/s. A 6x slowdown. The bottleneck wasn't the architecture. It was memory pressure from co-resident processes.
Before I acted on the architecture suspicion and downgraded the model, the measurement disproved it. Killing a hypothesis with a measurement saved me a pointless model swap. Don't act on a guess; measure first. That's a lesson I want on the record as a technical article.
What did real-hardware testing fix?
kit was not built by writing a spec and implementing it in one go. I live-tested it on real hardware, pointed out "this feels off" by feel, diagnosed it, and fixed it. It grew by stacking up that loop. Here are some representative grubby fixes.
| Symptom (what I felt on real hardware) | Diagnosis | Fix |
|---|---|---|
| The show's mp4 was missing the person's raw-voice turns | During high latency, interruptions caused recording events to be dropped | On the server side, persist the raw voice to disk before the turn response, and later pick it back up using file mtime as a fallback. The human's voice always remains in the mp4. |
| Backchannels chained and became absurdly long (the owner was furious) | The thinking-filler chain was unbounded | Limited the chain to one shot, placed a silence window, and changed it to a breathing pattern of "filler → silence → one survival signal → merge with the response." |
| A 224-second monologue wouldn't end the turn | Crossing a filler mid-recording reset the turn's cap-timer origin, rewinding it | Pinned the cap origin to a turn-level origin so it stays put even across fillers. |
| The voice stops completely (permanent GPU block) | A lock contention between nudge and turn left the lock dangling when the process was killed | Always reap the child process in finally before releasing the lock. No path leaks even once. |
I also fixed whisper producing hallucinations like "memememe." A single-character check missed multi-character loops, so I switched to n-gram repeat detection. To protect legitimate two-repeat backchannels like "sou-sou," it treats something as a hallucination only from three repeats up. Verification runs independently from the production build, and the results are recorded in a roughly 70KB verification report.
Summary: kit's technical essentials
kit does not implement "talking to an AI by voice" as a single-line request/response. It is a system where five lines (perception, understanding, timing, generation, output) spin in parallel and a single Arbiter collapses them into one verdict. Boiled down to three points:
- FSM exclusive state: holding only one state drives gaps and dead zones to zero
- The single Arbiter: passing through one point only at the moment of vocalizing structurally prevents three-way collisions
- Two-tier thinking: running a fast shallow answer and a slow deep answer in parallel, trying to turn the inherently unfavorable 100-second lag into an asset
And honestly: because of the physical constraint of a single GPU, "listening and thinking deeply in full parallel" is not realized yet. The deep tier tends to get canceled every turn, it's slow under memory pressure, and tuning the content-reply is a work in progress. Even so, the two North Stars, "never let the conversation die" and "the human always wins," hold. A design that completes the turn even while dropping quality best-effort, plus the serialized GPU lock, keeps them standing.
Fully local, free, owned: a single M1 Max time-slices ears, brain, and mouth while backchanneling at a near-human tempo, thinking deeply, and recording it all as a show. This was the record of building that whole thing as a parallel system, not a single line. A lot is still in progress, but I've written up the implementation so far, end to end, honestly.
FAQ
Q. Does kit use any external API or cloud service? A. No. Everything runs locally on a single M1 Max 64GB: the ears (whisper.cpp + ggml-small), the brains (Llama-3.1-Swallow-8B-Instruct and Gemma-4-26B), and the mouth (AivisSpeech). There's no monthly subscription either.
Q. What happens when the five systems all try to speak at once?
A. Only one place, arbitrate(), decides whether to speak, and it returns exactly one of {STAY_SILENT | BACKCHANNEL | FACILITATE | SPEAK}. Because there's a single decision point, voice collisions can't happen structurally.
Q. The deep answer takes over 100 seconds. Doesn't the conversation stall?
A. Swallow 8B returns a light reply first in about 8 seconds, while Gemma-4 26B builds the deep answer in the background over about 110–185 seconds. But when the human starts speaking again, deep_cancel stops the deep tier, so in real use with long turns, the deep answer often doesn't land yet.
Q. Can I interrupt by voice while the AI is speaking? A. No. During SPEAKING the voice analyzer is torn down, so interruption is by button only. Voice interruption works only within the recording window.


