Skip to main content
The free plan includes 10 manual file transcriptions in total. Pro unlocks unlimited transcriptions, speaker identification, multi-file batches, watch folders, deletion and export.
Beyond live dictation, Vowen transcribes pre-recorded audio and video: interviews, podcasts, recorded meetings, voice memos, lecture recordings. Everything lives on the Transcribe page in the left sidebar.
Vowen's Transcribe page with the Transcribe button highlighted in the top right of the toolbar

The Transcribe button sits at the top right of the window. The gear beside it opens Transcribe settings.

Transcribing a file

1

Open the Transcribe dialog

Click the Transcribe button at the top right of the window. A dialog opens with the upload options.
2

Pick a model

Choose any configured local or cloud model from the Transcription Model dropdown. The choice is per-file, so you can push one recording through a higher-accuracy model without changing your dictation default.
3

Set language and speaker options

Pick a Language or leave it on Auto-detect. If the model supports it, toggle Identify Speakers to label who said what. Leave Add Timestamps on to get clickable per-segment times, or turn it off for a clean transcript. (Timestamps are hidden when Identify Speakers is on, since diarized output carries its own segment structure.)
4

Add your file

Drag and drop an audio or video file into the drop zone, or click to browse.
5

Click Transcribe

Vowen handles compression and chunking automatically. The result appears with [MM:SS] badges where the model provides them, and you can edit, copy, regenerate or export it.

Transcribing several files at once

Attaching more than one file per run is a Pro feature. On the free plan the button reads Add more files (Pro) and each run takes exactly one file.
With Pro, drop in as many files as you like. They queue up, run with the same model and language settings, and each becomes its own transcription.
A row on the Transcribe page showing a spinner beside a filename while the file is being transcribed

A queued file shows a spinner and its filename while it runs.

Transcribe settings

The gear icon in the Transcribe toolbar opens a settings modal with three tabs.
Pre-fills the Transcribe dialog so you are not re-picking the same options every time. Per-run changes in the dialog stay one-off and are never written back here.
  • Transcription Model, or leave it on “Same as dictation”, which follows whatever your dictation model is.
  • Language
  • Identify Speakers, plus Merge speaker segments and Number of speakers once it is on
  • Add Timestamps
  • Notify me when complete, which shows a banner when a transcription finishes
The Transcribe settings modal on the Defaults tab, showing Transcription Model, Language, Identify Speakers, Merge speaker segments and Notify me when complete

The Defaults tab pre-fills the Transcribe dialog. Per-run changes never write back here.

Automatically save each finished transcription to a folder on disk, in addition to keeping it in the app. Tick any combination of Plain Text (.txt), SubRip (.srt) and WebVTT (.vtt), and choose an export location (your Downloads folder by default).This applies to every job, including watch-folder ones.
The Transcribe settings modal on the Auto-export tab with toggles for Plain text, SubRip and WebVTT, and an Export Location row

Auto-export writes every finished transcription to a folder as well as keeping it in the app.

Drop a recording into a watched folder and Vowen transcribes it automatically. Each folder can carry its own model, language, speaker and tag rules, or inherit the Defaults tab.Files already sitting in the folder when you add it are skipped; only new arrivals are picked up. See Watch folders.
The Transcribe settings modal on the Watch folders tab showing one configured folder with its path and rules, and an Add Folder button

Each watched folder carries its own rules. Add Folder appends another.

Supported formats

Audio: mp3, wav, m4a, aac, ogg, flac, wma, opus Video: mp4, mov, avi, mkv, flv, wmv, webm, mpeg, mpg For video files, Vowen extracts the audio track before transcription.

Timestamps

Local Parakeet (V2, V3 and Japanese) and Whisper produce per-segment timestamps, shown as time badges you can click to jump the audio. Most cloud models do too. Two exceptions: Parakeet Mandarin returns plain text with no timings, and Nemotron streams the file through its live engine, which also produces no timings. Neither can offer subtitle export.

File size and duration limits

Vowen compresses and chunks large files for you. What it does depends on the provider. You never need to split a file by hand.
OpenAI enforces a hard 25 MB per-request limit. If a recording is still over 24 MB after compression and short enough that chunking by duration would not help, Vowen reports that the file is too large rather than failing silently.

Regenerating a transcription

Open a completed transcription and the sidebar shows a Regenerate Transcript panel listing every model you have configured, grouped by Local and Cloud.
  1. Open the transcription detail page.
  2. Pick any model from the panel.
  3. For diarization-capable models, toggle Identify Speakers.
  4. Click Regenerate Transcript.
The original is preserved as a version, so regeneration adds a new one rather than overwriting. Useful for comparing models on the same audio, or for a higher-accuracy pass after a quick first run.
A transcription detail page showing timestamped transcript text, a Regenerate Transcript panel with a model picker, and an Ask about the transcript input

A transcription's page: the text with timestamps on the left, the Regenerate Transcript panel on the right, and Ask AI along the bottom.

Export

Open a transcription and click Export. The dialog shows a live preview and lets you choose: Include title & date adds a header to document exports. Speaker labels and timestamps come along automatically when the transcript has them.
Subtitle formats only appear when the transcript has per-segment timing. Captions are timed straight from those segments, one cue per spoken segment, with speaker names where available. A transcript with no segment timing will not offer subtitles; use PDF, Text or Markdown.
Exporting is a Pro feature.
The more menu on a transcription detail page, listing Export, Copy and Delete

Export lives in the detail page's menu, alongside Copy and Delete.

Renaming

A transcription’s title defaults to the file’s name. Open it and edit the title field at the top of the detail page; changes save automatically and the list updates.

Organizing with tags

Manual transcriptions have their own tag system, separate from meeting-note tags.
  • Add tags to any card with add tag. Create a new coloured tag or reuse an existing one.
  • Use the tag filter bar above the list, with a Match: Any / All toggle when two or more tags are selected, plus sorting by Newest, Oldest or Recently tagged.
  • Use the search box to find a transcription by name.
The Transcribe page tag filter bar with tag chips on the left and a sort dropdown on the right

The tag filter bar above the list. Match Any / All appears once two tags are selected, and the sort control sits at the right.

Multi-select and bulk actions

Hover a finished card to reveal its checkbox. As soon as one is ticked, a bulk-action bar appears at the bottom of the window with: Only finished transcriptions can be selected; anything still processing has no checkbox.
The Transcribe list with two rows ticked and a bulk action bar reading 2 transcriptions selected, with Select all, Add tags, Export, Delete and Cancel

Tick two or more rows and the bulk bar appears with the actions that apply to all of them.

Batch export

The Export action exports everything you selected in one go. Pick Markdown, Plain Text or PDF, plus SubRip and WebVTT when every selected transcript carries segment timing. Batch export is Pro.

Editing a transcript

Open any completed transcription (or meeting note) to edit it segment by segment:
  • Edit text. Click into a segment and type. Undo and redo with Cmd/Ctrl+Z and Cmd/Ctrl+Shift+Z.
  • Split a segment. Open the segment’s ⋯ menu, place your cursor at the cut point, and choose Split at caret. The split snaps to the nearest word boundary and timestamps are interpolated.
  • Delete a segment. From the same menu.
  • Reassign a line to a speaker. On diarized transcripts, move a single line to a different or new speaker. To fold two speakers together entirely, use Merge Speakers.

Finding text in a transcript

Press Cmd+F (macOS) or Ctrl+F (Windows) inside a transcript to open the find bar. It searches case-insensitively across the whole transcript, highlights every match, and shows a current/total counter. Enter goes to the next match, Shift+Enter the previous (both wrap), Esc closes.

Playing back the audio

When a transcription has a saved audio file, a waveform player appears at the top of the detail page:
  • Play/Pause, plus Back 10s and Forward 10s.
  • Scrub by clicking or dragging anywhere on the waveform.
  • Click a segment’s timestamp to jump the audio there and start playing.
  • The current segment is highlighted and scrolled into view as audio plays, paused while you are editing.

Chat with the transcript

Once a transcription is complete you can ask AI questions about it from the chat panel: pull out action items, summarize a section, find a quote. See Chat with Transcriptions.