Dictation feels awkward at first because typing lets you edit as you compose โ backspace is one keystroke away โ while speaking forces you to commit to a sentence before you know how it ends. That's not a skill you lack; it's a different cognitive mode, and it takes deliberate practice to build, not just more attempts at dictating your usual prose.
Most people try dictation once, produce three rambling, half-finished sentences, and conclude it doesn't work for them. What actually happened is they tried to think and edit in real time while also speaking linearly, which is a much harder task than typing the same thought, because typing gives you a visual buffer to revise before you commit.
When you type, you're composing in short, revisable bursts. You write a clause, glance at it, adjust a word, keep going. Your eyes and your editing brain are in a tight loop with your hands. Most people never notice how much micro-editing they do until it's gone.
Speaking removes that loop. You have to hold the shape of a sentence in your head before you say it, because once it's out, it's out โ there's no cursor to go back and fix a word mid-clause without stopping and restarting. This is the same skill involved in public speaking or dictating a letter to someone taking notes: composing linearly, in complete thoughts, rather than assembling text incrementally.
The result is that your first few dictation sessions will produce rougher text than your typing does, even though you're a competent writer. That's not a sign dictation doesn't suit you โ it's a sign you're using a muscle that typing let you neglect for years.
Vague advice like "just get used to it" doesn't move the needle. What does is treating dictation as a skill with specific drills, the same way you'd approach touch typing or a new instrument.
Dictate an outline first, then the prose. Before dictating a paragraph, say three or four short phrases capturing the structure โ "problem, why it matters, the fix, one caveat." Having the shape decided removes the hardest part of linear composition: figuring out where the sentence is going while you're saying it.
Say punctuation explicitly, at least at first. Say "comma," "full stop," "new paragraph" out loud rather than relying on the model to infer pacing from pauses. It feels stilted for the first few sessions, but it forces you to think in sentence units, which is exactly the skill dictation requires. You can drop it once the habit sets in.
Accept a rougher first pass. Don't try to dictate publish-ready prose. Dictate the idea, then edit the transcript the way you'd edit a first draft from someone else. Separating "get the thought down" from "make it read well" is the same discipline good writers already use for typed drafts โ dictation just makes the two stages more obviously separate.
Start with low-stakes text. Commit messages, Slack replies, code comments, quick notes to yourself. These have low structural demands, so you can focus purely on the mechanics of speaking in complete thoughts without also worrying about getting an email exactly right.
Re-read out loud before you speak a sentence you're unsure of. If you're not sure how a sentence should end, say it in your head first, the way you'd silently draft a tricky line before typing it. This closes some of the gap between the composition loop typing gives you and the one-shot nature of speech.
On Linux, the mechanics matter less than the habit, but it helps if the tool gets out of your way. Voxtty runs as a systemd user service โ press Alt+D, speak, and the transcript gets typed into whatever has focus, transcribed on-device with faster-whisper so there's no round trip to a cloud API adding latency while you're mid-thought. Try Voxtty free if you want to run these drills without setting up anything beyond a hotkey.
Some of the awkwardness won't fully disappear, and it's worth knowing which parts are permanent rather than a skill issue. Technical jargon, acronyms, and code identifiers are genuinely harder for any speech model โ Whisper included โ because they're low-frequency in training data and easy to mishear against similar-sounding words. Expect to correct these by hand regardless of practice.
Long, structurally complex sentences with multiple subordinate clauses are also harder to dictate cleanly than to type, because you're holding more grammatical state in your head with no visual scaffold. If you catch yourself losing the thread mid-sentence, that's a sign to dictate shorter, then combine clauses in the edit pass rather than trying to nail the long sentence in one go.
Background noise and strong accents remain real accuracy hits for on-device models โ a busy office or a video call in the background will degrade transcription quality more than it degrades your own comprehension. None of this is specific to Voxtty; it's a property of current speech recognition, cloud or local.
Pick one paragraph you'd normally type โ a Slack message, a commit message, a short email โ and dictate it instead, saying punctuation out loud and accepting whatever comes out. Don't edit until you're done speaking. That single rep, repeated a handful of times this week, does more than reading another article about why dictation feels strange.
Local-first voice dictation for Linux. Press Alt+D, speak, and your words land in whatever app has focus โ nothing leaves your machine.
Try Voxtty free โ