Skip to main content

Check which model you are on first

Changing model is the single largest accuracy lever.
The Speech models page listing cloud providers and local models with their sizes

The Speech models page. The active model is the dark card; local models show their download size.

Local models live on the Speech models page in the left sidebar, not in the Settings modal.
If you use a custom Dictionary, moving off Whisper Tiny or Base matters twice over. Those two models are also the only local Whisper sizes that receive no vocabulary at all. See Custom vocabulary below.

What custom vocabulary actually does

It depends entirely on your engine. Some use your Dictionary well, some use it weakly, and some ignore it completely. Each has its own ceiling on how many terms it will take.
If your Dictionary “does nothing”, check your model before anything else. Several of the most commonly used local models receive no vocabulary at all:
  • Whisper Tiny and Base, including their .en variants
  • Parakeet, every variant, on macOS and Windows
  • Nemotron EN and Nemotron Multilingual
On these, terms you add to the Dictionary never reach the speech engine. Nothing is broken and nothing is misconfigured. The words simply are not sent.Whisper Tiny and Base are the sharp edge here. Base was Vowen’s recommended default for a long time, so long-time users are disproportionately sitting on a model that ignores every term they add. Vocabulary starts at Whisper Small. Below that, the model cannot act on your terms reliably, and tends to either ignore them or echo them into the transcript.Parakeet is the current default on both macOS and Windows, so new users land on an engine that takes no vocabulary too.See What to do instead below.

How each engine receives vocabulary

These steer recognition towards your terms directly, so an unusual name comes back spelled right rather than guessed at:You do not have to manage these caps. Vowen trims and tidies your list to fit each one, so a long Dictionary never turns into a failed transcription.
These cannot be steered. Your terms go in as a hint about what is likely to be said, which the model may or may not act on:
  • Local Whisper Small and larger. Small, Medium, Large v3 and Large v3 Turbo, plus the English-only variants of Small and Medium. Tiny and Base are excluded, and get nothing.
  • Whisper Large v3 and Large v3 Turbo (Groq)
  • Custom speech servers running an OpenAI-compatible Whisper endpoint
  • Sarvam Saaras v3, live dictation only
  • Google Gemini, file transcription only
Vowen keeps the hint short enough that the model uses all of it, and ranks your terms so the ones you say most often are the ones that fit.The Small floor applies to dictation, file transcription and meeting notes alike.These hints are real, but much softer than the engines above, and will not always overcome a confident misrecognition.
Adding terms has no effect on:
  • Whisper Tiny and Base, including tiny.en and base.en
  • Parakeet, every variant, on macOS and Windows
  • Nemotron EN and Nemotron Multilingual
  • Sarvam Saaras v3, file transcription. Sarvam takes vocabulary during live dictation only.
  • Cartesia, file transcription. Cartesia’s batch model takes no vocabulary.
  • Gemini Live, used for live dictation. Only Gemini file transcription applies your terms.
  • OpenAI’s diarization model (gpt-4o-transcribe-diarize), which takes no vocabulary of any kind
Sarvam and Cartesia are split. Both use your terms when streaming, and neither does when transcribing a file.

Which terms get sent

Your Dictionary is listed alphabetically, so simply taking the first terms would give a user with 200 entries a vocabulary that effectively stops at “C”. Vowen prioritises by how often you have actually said each term, then how recently, then when you added it, and fits as many as the provider’s ceiling allows.

What to do instead of vocabulary

If your engine takes no vocabulary, or it is not enough on its own, these all still work:
1

Add a misspelling alias to the Dictionary entry

Each Dictionary entry can carry the specific wrong spellings you actually see. That is a text substitution applied after transcription, so it works on every engine including Parakeet and Nemotron.
2

Use a Thread for consistent replacements

For a phrase that is always transcribed the same wrong way, a Thread replaces it every time. Triggers are matched as whole phrases, so a sig thread will not rewrite the middle of “design”.
3

Turn on AI enhancement

Enhancement receives your Dictionary terms as context and fixes misrecognitions the speech model got wrong. This is the most effective route for Parakeet and Nemotron users. Turn it on at Configure your AI > Configuration > Enhance transcription with AI.
4

Switch to an engine that uses your Dictionary

If your work depends on a specific set of proper nouns, this is the real fix.Moving from Whisper Tiny or Base up to Whisper Small is the smallest possible change and turns vocabulary on at all, since Small is the floor.For the strongest result, use one of the engines that steers recognition directly: Deepgram Nova 3, AssemblyAI, ElevenLabs Scribe v2 or Speechmatics. The difference on proper nouns is large.
See Dictionary and Threads for the full detail.

Set your language explicitly

Auto-detect occasionally picks the wrong language, and a wrong-language decode is far worse than a slightly weaker model.
Settings > Language showing UI Language, Dictation Language, Notes Summary Language and English Spelling

Settings > Language. Dictation Language is the one that reaches the speech model.

  1. Go to Settings > Language
  2. Set Dictation Language to your primary language
The available languages change with your model. Deepgram Nova 2 offers 33, Nova 3 offers 50, xAI 25, Mistral 13, Sarvam 23, Speechmatics 39, AssemblyAI streaming only 6. If you switch to a model that does not support your stored language, the selector silently resets to Auto. English-only models such as Parakeet V2 and Nemotron EN lock the selector entirely.

Turn on AI enhancement

AI enhancement catches and fixes most remaining transcription errors, and it receives your Dictionary terms as context. Turn it on at Configure your AI > Configuration > Enhance transcription with AI.
The Configure your AI page with the Enhance transcription with AI row highlighted

Enhance transcription with AI, under the provider grid in Configure your AI.

If no AI provider is connected, that toggle is replaced by a locked card rather than being greyed out. Connect a provider on the same tab first.

Specific issues

  • Slow down slightly. Rapid speech is harder for every speech model.
  • Switch to a cloud model. Nova 3 (Deepgram) and Scribe v2 (ElevenLabs) handle fast speech well.
  • On a local Whisper model, move up a size.
Check which engine you are on first, because the answer differs completely. See Custom vocabulary above.
  • Whisper Tiny or Base, Parakeet, or Nemotron: vocabulary does nothing. Use a Dictionary misspelling alias, a Thread, or AI enhancement, or move to an engine that takes vocabulary.
  • An engine that uses your Dictionary well: add the terms to your Dictionary (the top-level sidebar page, not a settings section) and recognition is steered towards them directly.
  • Local Whisper Small or larger, or Groq: add them to the Dictionary. They go in as a hint, which helps but is weaker.
  • Use a multilingual model. Never a .en Whisper variant, and never Parakeet V2 or Nemotron EN, all of which are English-only.
  • Set the language explicitly rather than relying on auto-detect.
  • Parakeet V3 covers 25 European languages and is fast.
  • Nemotron Multilingual is macOS only and streams with punctuation.
  • For Japanese, Parakeet Japanese is a dedicated macOS-only model. For Mandarin, Parakeet Mandarin, also macOS only.
  • For Indian languages, Saaras v3 (Sarvam AI).
  • Whisper Large v3 covers 99 languages and remains the broadest option.
Repetition is a known speech-model behavior, more common on smaller models.
  • Re-record the phrase, or use Retry on the Voice Log entry to re-run the saved audio through a different model
  • Move up a model size, or to a cloud model
  • AI enhancement detects and removes repetitions
Whisper-family models invent text over silence, and “thank you” is the most common output.
  • Start speaking as soon as you press the shortcut rather than holding it in silence
  • Vowen already discards local results that are exactly “thank you”, “thank you.”, ”.” or ”.\n”. If something is reaching your app, it was not an exact match.
  • Turn on Settings > Experimental > Enhanced silence detection for cloud models. A Sensitivity slider from 1 to 5 appears once it is on, defaulting to 2.
  • Whisper-family audio already has its trailing silence trimmed before upload, precisely to reduce this.
The recording was probably too short.
  • Hold the shortcut for the full duration of your speech, or switch that shortcut to Tap mode in Settings > Shortcuts so you are not fighting the key
  • Parakeet requires at least 300 ms of audio. Anything shorter is silently discarded, with no error and no Voice Log entry, on the assumption you tapped the key without speaking.
  • Check that your shortcut is not being consumed by the target app
See First or last word cut off.Use Retry on the Voice Log entry to confirm. Retry always re-reads your original recording.
Vowen applies smart spacing and capitalisation on device, using the text already around your cursor, so a dictation dropped mid-sentence does not arrive with a stray capital or a doubled space. It runs after Threads are applied.It does not apply to text expander injections.
There is a dedicated setting for this. Settings > Language > English spelling switches between American and British using an on-device converter. You do not need to write a custom AI instruction for it.
macOS, fast and accurate, fully local:
  • Speech: Parakeet V3, or Nemotron EN if you want punctuated live preview
  • AI: Gemini 2.5 Flash-Lite or Claude Haiku 4.5
Windows, fast and accurate:
  • Speech: Parakeet V3 locally, or Whisper Large v3 Turbo (Groq) in the cloud
  • AI: Gemini 2.5 Flash-Lite
Maximum accuracy, any platform:
  • Speech: Scribe v2 (ElevenLabs), Universal-3.5 Pro (AssemblyAI), or Nova 3 (Deepgram), all of which use your Dictionary well
  • AI: Claude Sonnet 4.6 or GPT-5.4
Available AI models are fetched from Vowen’s catalog and change over time. Open Configure your AI > Configuration to see the current list for your provider rather than relying on names printed here.