Linux

Does local Whisper dictation need a GPU on Linux?

7 October 2026 5 min read

No, local Whisper dictation does not need a GPU on Linux. faster-whisper runs Whisper models through CTranslate2, an inference engine optimised for CPUs, and with int8 quantization the small and base models transcribe a sentence of speech faster than you spoke it on an ordinary modern laptop. A GPU helps with the largest models and with batch jobs, but dictation is short bursts of speech, which is a much lighter workload.

The catch is that "modest hardware" has a floor. An old dual-core chip, or a laptop that throttles after two minutes under load, will still make dictation feel sluggish. Both halves of that are worth understanding before you install anything.

Why the GPU assumption is so common

Most people meet Whisper through the original PyTorch reference implementation, and that one is slow on CPU. It loads float32 weights, uses a general-purpose runtime, and was written for research rather than for sitting in the background of a desktop. The reasonable conclusion is "I need CUDA." On Linux that also means proprietary drivers, a few gigabytes of toolkit, and a card that may not exist in your machine.

That assumption matters because of the alternative. If local transcription requires a GPU, many people fall back to a cloud speech-to-text API. The failure modes there are concrete: every utterance adds a network round trip, an API outage means you cannot type by voice at all, and your recorded audio sits on someone else's server. Removing the GPU requirement removes the main practical reason to accept those tradeoffs.

What makes CPU inference practical

Two things do most of the work. The first is CTranslate2, the C++ inference engine underneath faster-whisper. It uses optimised CPU kernels, fuses operations, and avoids the overhead of a full deep-learning framework. The same Whisper weights run noticeably faster and in less memory than in the reference implementation, without changing the model itself.

The second is quantization. Whisper weights are normally stored as 32-bit or 16-bit floats. Converting them to 8-bit integers (int8 in faster-whisper's compute_type setting) roughly halves or quarters the memory the model needs and lets the CPU do much cheaper arithmetic. The cost is a small drop in accuracy, which tends to show up as an occasional wrong word on hard audio rather than broad degradation. I won't quote a percentage here because it varies by model, language and audio.

The model size is the other lever, and it matters more than quantization. Roughly in order: tiny, base, small, medium, large. Each step up is more accurate and noticeably slower. For dictation on CPU, small with int8 is a common sweet spot, and base is the fallback when the machine is struggling.

Voice activity detection helps too. A VAD stage figures out when you're actually speaking, so the model only processes speech and not the silence between sentences. For a tool that listens for a hotkey and then transcribes, that keeps CPU usage near zero while idle.

Voxtty is built this way. You press Alt+D, speak, and the text is typed into whichever app has focus, using faster-whisper on CPU by default. Nothing is uploaded, and the basic cleanup of filler words is rule-based and offline. It runs as a systemd user service, so if you want to know how that stays alive across logins and crashes, we've covered how to keep a background tool running with systemd --user.

Where CPU-only still struggles

Old hardware is the obvious one. A decade-old low-power laptop CPU may take longer to transcribe an utterance than the utterance lasted, and the delay is noticeable when you're waiting for a sentence to appear. Dropping to the tiny or base model helps, but those models make more mistakes, so you trade speed for cleanup work.

Thermal throttling is the less obvious one. A laptop that benchmarks fine for ten seconds can slow down sharply once the fans can't keep up, especially if you're also compiling something or running a browser with forty tabs. Dictation is bursty, so it usually copes better than a long batch job, but a thin-and-light machine under sustained load will show the lag.

Plain accuracy limits apply regardless of hardware. Whisper models handle clear speech in a quiet room well, and they get worse with strong accents, background noise and unusual technical vocabulary. Larger models help, but a larger model on CPU is exactly where the speed problem returns. On-device transcription is also not as fast as the best cloud services, and anyone who tells you otherwise is rounding in their own favour.

Two smaller practical notes. Typing into other apps relies on ydotool, which needs its daemon running and permission to use /dev/uinput, so check that your distro packages it and that your user is set up for it. And hands-free wake word activation is experimental, because always-on listening is harder to get right than a hotkey.

One thing to do today

Check what you actually have before assuming you need new hardware. Run lscpu to see your core count, then try dictating a single paragraph with the small model on int8. If it keeps up, you're done; if it lags, drop to base and see whether the accuracy is tolerable. Try Voxtty free if you'd like a ready-made way to run that test with a hotkey and no cloud account.

Try Voxtty free

Local-first voice dictation for Linux. Press Alt+D, speak, and your words land in whatever app has focus โ€” nothing leaves your machine.

Try Voxtty free โ†’

Related Articles

Linux
How do I keep a background tool running with systemd --user?
โ† Back to blog