Making a local Home Assistant voice assistant answer in under three seconds
Home Assistant’s voice pipeline has three stages: speech-to-text (STT) turns what I said into words, a conversation agent (here an LLM) works out what to do and what to say back, and text-to-speech (TTS) says it. I wanted all three running on the new Strix Halo box, with nothing leaving the house.
By Saturday morning it all worked end to end. Asking “what’s the weather tomorrow?” took about 2 seconds to transcribe and about 5 to think, which is long enough to wonder whether it heard you. Two days of fiddling later, STT takes a quarter of a second, and the LLM takes about three.
The bit I talk to is a Home Assistant Voice Preview Edition, in a 3D-printed case that sticks magnetically to the fridge.
Speech-to-text: NPU, then CPU, then GPU
The plan was to run Whisper on the NPU, via Lemonade’s FastFlowLM backend, to keep the GPU free for the LLM. After a bit of wrangling over model names, it plugged into Home Assistant through the openai_stt custom integration. But it was slow: about 2 seconds for a 2-second phrase. Timing the NPU on its own showed a flat ~2.1 s whatever I said, because it pads every clip out to Whisper’s full 30-second window.
So I tried the standard Wyoming faster-whisper container on the CPU. The first try was slower, about 3 s with a bigger model. There was a clue in the log: the model was stored as float16, the CPU couldn’t do float16 efficiently, so it had quietly converted everything to float32. It was also only using 4 threads, on a 16-core CPU. Forcing int8 and raising the thread count took a smaller model from about 2 s to 1.05 s, and the bigger distil-large-v3 to 1.34 s.
faster-whisper can only use a GPU through CUDA, so that was as far as it could go on this box. whisper.cpp, which Lemonade also ships, can use an AMD GPU. My first benchmark of it gave the same ~0.77 s for every model size, which should have been a hint. None of them had actually loaded. whisper.cpp’s default GPU backend on this box is ROCm, and that build crashed on start-up (exit 134). The runs I was timing were just failed load attempts. Forcing the Vulkan backend fixed it, and the difference was night and day:
Through Home Assistant, including the HTTP round trip, it’s 0.26 s, and that’s with the full large-v3-turbo model, which should cope better with the voice satellite’s noisy audio than the distilled one. The price is that Lemonade has to be told to use Vulkan every time: an idle eviction or a reboot would otherwise reload it with the crashing ROCm backend. So the model’s options are saved as Vulkan and pinned (never unloaded), and the Lemonade quadlet reloads it that way at boot.
As for the NPU, it’s idle again. So much for keeping the GPU free, but the stages run one after another anyway, so the GPU isn’t doing anything else while it’s transcribing.
Text-to-speech: finding a British voice
This was the easy stage speed-wise, and the fiddly one otherwise. Lemonade can also serve Kokoro, a small and very natural-sounding TTS model, but every British voice I tried came back as the same 0.32-second blip. It turned out that Lemonade’s copy only had the American voices, and asking for any other voice got you a stub rather than an error. Standalone Kokoro-FastAPI ships the lot, and bf_alice did the job.
Today I also set up Piper, the standard Wyoming TTS for Home Assistant, with the en_GB-alba-medium voice, and by this evening I’d switched to it, simply because I prefer the voice. Either way, TTS takes milliseconds, and it starts speaking from the first sentence.
The LLM: two passes and a prompt cache
The conversation agent is Qwen3.5-35B-A3B in llama.cpp, connected through Home Assistant’s OpenAI-compatible integration. Home Assistant’s event log showed where the ~5 seconds went. Because the agent calls a tool (here, the weather forecast), every turn is two LLM passes. The first reads the whole prompt and decides which tool to call. The tool runs. Then the second reads the result and writes the answer. So the prompt gets processed twice.
The model was at full BF16 precision, and generation ran at about 24 tokens a second. Generation speed on this box is mostly limited by memory bandwidth, and 8-bit weights are half the bytes of BF16, so I switched to the Q8_0 quantisation, which is near-lossless. A quick benchmark said 46.4 tokens a second: almost exactly twice as fast.
Home Assistant still said 6 seconds. That turned out to be on the Home Assistant side, where the model is set in the integration, and once that was sorted, the thinking stage came down to 3.13 s:
Generation halved, as expected. The second prefill (reading the prompt) nearly halved too, and that’s the prompt cache at work. I’d added llama.cpp’s --cache-reuse flag for this, but the log said this model doesn’t support it and had quietly turned it off. It didn’t matter: llama.cpp’s ordinary prefix caching still works, and the second pass only had to read about 500 new tokens out of roughly 3,900.
Those 3,900 tokens are mostly Home Assistant listing every entity I’ve exposed to the agent, along with its current state. Trimming that list is the obvious next win, and I’ve decided against it. Half the point of an LLM agent is that it can control the house when my phrasing doesn’t match Home Assistant’s built-in sentences, and for that it needs to see the house. The states change constantly, so part of the prompt has to be read again every time, and that costs roughly a second on the first pass. I’ll take it.
It also feels faster than 3.13 s, because the speech starts streaming as soon as the first sentence is ready: about 2.55 s from the end of my sentence to the start of the answer.
One last hiccough: sharing
This afternoon I moved my news app off the old box and onto the same llama-server, giving it a slot of its own with --parallel 2. Straight away, the next voice request took 8 seconds. The model hadn’t been unloaded; the log showed that the prompt cache had filled up. llama.cpp keeps 8 GiB of cached prompts by default, and the news app’s constant stream of small requests had pushed out the oldest entry: the 2.2 GB voice prompt. So the next voice turn had to read all 3,900 tokens from scratch, which took 4.6 s.
--cache-ram 16384 doubles the cache, which leaves room for both. The box has plenty of memory to spare.
Where it’s at
Fully local: Whisper large-v3-turbo on the GPU via Vulkan in 0.26 s, Qwen3.5 at Q8_0 in about 3 s, and Piper in milliseconds, with a British voice at the end of it. The LLM is still most of the wait. The remaining levers are a smaller model or a tighter quantisation, and both would trade some accuracy for the speed.
That said, about 3 seconds is for pretty much the longest answer it will ever be asked for: tomorrow’s weather. The weather intent calls a script of mine, which turns Home Assistant’s hourly forecast into a line per hour (“Lunch time cloudy, 18 degrees (feels like 16), 20% chance of 0.2 millimeters rain”) and hands the LLM some instructions for what to do with it:
instructions: >-
Briefly present tomorrow's weather in a concise conversational style,
summarising the weather for the morning, lunch time, afternoon and
evening, if there are significant changes e.g. 'It will be nice and warm
and sunny tomorrow, warming up to 22 degrees at lunch time, before
clouding over in the afternoon and cooling to 15 degrees in the evening.'
Output as if you were speaking out loud, so do not use abbreviations
or emojis.
That last line matters more than you’d think. A text-to-speech engine reading out “22°C, 60% ☔” is not a pretty sound. Most questions get a much shorter answer than this, and those come back quicker still.
Project assisted by Claude Code.