FAQ
Local transcription is free forever — Whisper, Parakeet, Ollama AI enhancement, unlimited history, custom presets, hotwords, all 90+ supported languages, hardware acceleration, and the global hotkey. No card required to sign up. Whisper Pro adds the Cloud features: OpenAI transcription (gpt-4o-mini-transcribe / gpt-4o-transcribe), Cloud AI enhancement (gpt-5-mini and friends), and web-aware AI with the OpenAI Responses API. Pro is $9.99/month, $79.99/year, or $69 lifetime. Same prices we've always had — we just made everything local free.
$69 once, Whisper Pro forever. That covers Cloud transcription with OpenAI's latest models, Cloud AI enhancement, web search, and every Pro feature we ship in the future — no recurring fees, no annual renewals. The free Local stack (Whisper, Parakeet, Ollama, presets, history, 90+ languages) is included by default for everyone, lifetime customer or not. Monthly ($9.99) and yearly ($79.99) plans are available if you'd rather start smaller. Pro plans include a 7-day Cloud trial; lifetime is one-time payment with no trial.
Absolutely. We offer a 7-day money-back guarantee on all purchases, no questions asked. If Whisper does not meet your expectations for any reason — whether it is a compatibility issue, performance concern, or simply not the right fit for your workflow — send us a quick email at whisper@remskill.com and we will process a full refund. We want you to feel completely confident in your purchase, and we would rather give your money back than have an unhappy customer.
Whisper gives you two fundamentally different ways to transcribe, and you can switch between them at any time. Local mode downloads an AI model to your computer and runs all transcription entirely on your machine — nothing ever leaves your device, no internet is needed, and it works completely offline. This is ideal if privacy is your top priority or you work in environments without reliable internet. Cloud mode connects to OpenAI's API using your own API key and unlocks a significantly more powerful experience: higher transcription accuracy powered by OpenAI's latest models, AI-powered text improvement that automatically cleans up grammar and removes filler words from your speech, a built-in web search that lets you research any topic by voice and get answers with sources, and custom AI instruction presets that can rewrite, translate, summarize, or transform your text in any way you define. Both modes use the same one-hotkey workflow — press, speak, and your text appears wherever your cursor is.
Whisper offers 8 local models split into two categories: English-only and Multilingual. English-only models are optimized exclusively for English and deliver faster, more accurate results for English speakers. They come in four sizes: Base (about 140 MB, fastest), Small (about 480 MB, recommended for most English users), Medium (about 1.5 GB, highest accuracy), and Turbo (about 1.5 GB, distilled model that runs 6x faster than the largest model while retaining 99% of its accuracy). Multilingual models support 90+ languages and come in four sizes as well: Small (480 MB), Medium (1.5 GB), Large V3 (3 GB, best accuracy), and Turbo (1.62 GB). If you only transcribe in English, always pick an English-only model — they are faster and more accurate than the equivalent multilingual model for English speech.
Multilingual models can transcribe speech in over 90 languages, including auto-detection where the AI identifies the spoken language automatically. However, there is an important performance tip: if you know which language you are speaking, always select that specific language in settings rather than relying on auto-detection. When you set a specific language, the model skips the language identification step and focuses entirely on accurate transcription, which makes it noticeably faster and more reliable. For the best multilingual experience, we recommend the Large V3 model (about 3 GB) with your language explicitly selected. If you only ever speak English, use an English-only model instead.
Think of the AI assistant as having ChatGPT built directly into every application on your computer, activated entirely by your voice. It goes far beyond basic transcription. You can ask factual questions and get detailed answers without opening a browser. You can dictate rough, unstructured thoughts and have AI rewrite them in a professional tone. You can ask it to summarize a paragraph, translate text to another language, extract action items from meeting notes, draft an email response, write code comments, convert text into bullet points, or follow any custom instruction you define. With web search enabled, it can look up real-time information and deliver a researched answer with source citations, pasted directly where your cursor is. All of this is triggered with a single hotkey press from any application on your computer.
When you enable web search in cloud mode settings, Whisper gains the ability to search the internet on your behalf, entirely by voice. Here is how it works: you press the hotkey, ask a question out loud, and the AI automatically detects that your question needs live information, searches the web, reads through multiple sources, synthesizes everything, and gives you a well-structured answer with source citations. The answer is pasted directly where your cursor is, whether that is a document, an email, a chat window, or a code editor. You never need to open a browser, switch tabs, or copy-paste anything. This is particularly useful for researchers, writers, developers looking up documentation, or anyone who needs quick factual answers without breaking their workflow.
AI text improvement is an optional feature in cloud mode that automatically processes your transcribed speech before it gets pasted into your application. When you enable it, the raw transcription is sent to an AI model that intelligently cleans it up: it removes filler words like "um", "uh", "like", and "you know", fixes grammar and punctuation errors, improves sentence structure, and produces polished, publication-ready text. You can also set custom instructions to control exactly how the AI processes your output. This feature is especially valuable for professionals who dictate emails, meeting notes, reports, documentation, or any content where presentation quality matters.
Yes, and this is one of the most powerful features available. You can create unlimited instruction presets — reusable AI commands that tell the AI exactly how to process your speech. For example, you might create presets like "Rewrite as a professional business email", "Translate to Spanish", "Summarize in three bullet points", "Extract action items and deadlines", or "Convert to a numbered list". Before you start recording, you select which preset to use from a dropdown, and the AI processes your transcription according to those specific instructions. You can also add custom vocabulary words (hotwords) to help the AI correctly recognize industry-specific terms, product names, technical jargon, or people's names.
Cloud mode is part of Whisper Pro. Once you're on a Pro plan, setup takes about two minutes — open Settings, go to the Model section, switch to the Cloud tab, and paste in your own OpenAI API key. You can grab one at platform.openai.com (free OpenAI account, no subscription, you only pay OpenAI for what you actually use). Cloud activates immediately. Your API key is encrypted and stored on-device using your machine's fingerprint — it never touches our servers. If you're on Free, the Cloud tab is locked behind an in-app upgrade card; nothing about your local setup changes.
Cloud mode uses your own OpenAI API key, so you pay OpenAI directly for actual usage on top of your Pro subscription. The OpenAI side is genuinely minimal: gpt-4o-mini-transcribe is about $0.003/minute of audio (an hour of dictation runs under $0.20), gpt-4o-transcribe is about $0.006/minute, and AI text improvement via chat completions costs a fraction of a cent per request. A user who dictates ~30 minutes a day, runs AI enhancement on every transcription, and does a few web searches typically spends $1–$3 per month with OpenAI. We take no cut, no markup, no platform fee — your API key, your bill.
No — and this is by design, not a missing feature. A team plan shares the subscription (every member's Cloud features unlock the moment you assign them a seat), but each member supplies their own OpenAI API key on their own machine. The key is encrypted on-device using a fingerprint derived from that specific computer's hardware, so the encrypted blob is not portable — even if we shipped your key to a teammate, it would not decrypt on their machine. We never store, sync, or proxy API keys through our servers. In practice: when a teammate signs in for the first time, Local mode (Whisper, Parakeet, Ollama AI enhancement) works immediately for free with no key. To use Cloud mode they open Settings → Model → Cloud and paste their own OpenAI key — same two-minute setup as a solo user. Everyone gets billed by OpenAI for their own usage, no admin gets surprised by a teammate's spending spree, and no single compromised device leaks a shared team key. If you need pooled billing where the admin pays for all team OpenAI usage, that requires a server-side proxy tier — reach out to us if that's a fit and we can talk about it.
Your privacy depends on which mode you choose, and you are always in full control. In local mode, your audio is processed entirely on your computer using a downloaded AI model. Nothing is sent to any server, ever — not to us, not to OpenAI, not to anyone. This makes local mode suitable for highly sensitive environments like healthcare, legal, finance, or classified work. In cloud mode, your audio is sent directly to OpenAI's transcription API through your own personal API key. The data goes from your computer straight to OpenAI — we are never in the middle, we never see your audio, we never see your transcription results. We do not collect, store, analyze, log, or monetize any of your voice data or transcriptions in either mode.
Yes — local mode works completely offline with no internet required at any point during transcription. The only time you need internet for local mode is the initial one-time download of your chosen AI model, which ranges from about 140 MB for the smallest model to 3 GB for the largest. Once the model is downloaded and saved to your computer, all transcription happens locally with zero network activity. You can use Whisper on a plane, in a basement with no signal, in a secure facility, or anywhere else without connectivity. Cloud features do require an active internet connection since they communicate with OpenAI's servers in real time.
Yes — Whisper uses a system-wide global hotkey that works in any application where you can type. On Windows the default hotkey is Ctrl+Space, and on macOS it is Cmd+Space. This works in word processors, email clients, messaging apps like Slack, Discord, and Teams, code editors like VS Code and JetBrains IDEs, note-taking tools like Notion and Obsidian, browser text fields, and any other application that accepts text input. When you press the hotkey and speak, your transcribed text — optionally enhanced by AI — is automatically pasted wherever your cursor is positioned. Both local and cloud modes work identically through this same global hotkey system.
Accuracy depends on the mode and model you choose. In local mode, transcription accuracy typically ranges from 95% to 99%. Larger models deliver higher accuracy. Several settings let you fine-tune quality: beam size controls how thoroughly the model searches for the best transcription, voice activity detection filters out silence and background noise, and custom vocabulary words help the model correctly recognize names, technical terms, and domain-specific jargon. In cloud mode, OpenAI's latest transcription models deliver even higher accuracy. On top of that, enabling AI text improvement adds another layer of correction. For the absolute best results, we recommend cloud mode with AI text improvement enabled.
Whisper supports over 90 languages in both local and cloud mode, including English, Spanish, French, German, Portuguese, Chinese, Japanese, Korean, Russian, Arabic, Hindi, Italian, Dutch, Turkish, Polish, and many more. If you only transcribe in English, use an English-only model — these are smaller, faster, and more accurate for English than the equivalent multilingual model. If you transcribe in other languages or switch between multiple languages, use a multilingual model and set your primary language explicitly in settings. In cloud mode, OpenAI's API supports all the same languages through a single model. Custom vocabulary words work in any language to help the AI recognize specialized terms, brand names, or jargon.
Whisper runs on Windows 10 or later and macOS 11 (Big Sur) or later, including full support for both Intel and Apple Silicon Macs. The app itself is lightweight at only about 25 MB. For local mode, you will need additional disk space for the AI model you choose: English-only models range from about 140 MB to 1.5 GB, and multilingual models range from 480 MB to 3 GB. We recommend at least 4 GB of RAM for basic use with smaller models and 8 GB or more for larger models. No dedicated GPU is required — all local transcription runs on your CPU. For cloud mode, the system requirements are minimal since all heavy processing happens on OpenAI's servers.
Whisper includes a built-in auto-update system that keeps your app current with zero effort. The app periodically checks for new versions in the background and shows a notification banner when an update is available. You can install it with a single click. If you purchased a lifetime license, you receive every update we ever release at no additional cost — this includes new features, new AI capabilities, new model support, performance improvements, additional language support, and bug fixes. Monthly and yearly subscribers also receive all updates for the duration of their subscription.