Getting started
Twenty minutes, most of it waiting for two signups. At the end you will have had Whispr answer a question out loud in a rehearsal, which is the only way to know it works before it matters.
1. Install
Whispr is an ordinary desktop app. It is not in the Mac App Store or the Microsoft Store, you do not need a developer account, and nothing here asks for administrator rights.
The trade for that is a warning on first launch, on both platforms: the build is not yet code-signed, and both operating systems say so about any app whose publisher they cannot name. It is worth knowing in advance, because it looks alarming and it is not a warning about anything Whispr does. Neither is a step you repeat — once you have allowed it, it opens normally forever after, including through updates.
On a Mac: open the DMG and drag Whispr to Applications. Open it from Applications — not from inside the DMG, or macOS will re-ask for permissions every launch.
The first time, macOS refuses and offers only Done. Go to System Settings → Privacy & Security, scroll to the bottom, and next to the line about Whispr being blocked choose Open Anyway. (On older versions of macOS, right-clicking the app and choosing Open does the same thing.)
On Windows: run Whispr Setup <version>.exe. It installs for you alone and
never asks for administrator rights. Windows shows a blue "Windows protected
your PC" panel on the first run — choose More info, then Run anyway.
The first launch takes a few seconds longer than the ones after it: the app starts a small background service that does the actual work, and that service keeps running if you close the window. That is deliberate — the session, the transcript and everything already recorded survive the window closing.
On Windows, keep the window open anyway. The service survives, but the recording does not: Windows capture happens inside the Whispr window, which is exactly why hearing the other side needs no permission there. Minimised is fine and behind your video call is fine — closed is not. On a Mac the window really is optional, and capture continues without it.
The browser version, and what it is not
There is a browser app at whispr.one/app. It hears the call and answers, on all three kinds of meeting, with no download — it is a real version of the answering half, not a demo.
What it does not have is everything that persists or watches around a meeting: pausing, a CV tailored to one role, the round format, the meeting record, goals, who said what, and live steering. Those need the background service that only the desktop app installs, and they are not coming to the browser. The app says so on its own first screen.
Use the browser one on a machine you cannot install software on. Use the desktop one everywhere else.
2. Give it the permissions it needs
Settings → Permissions, or whispr permissions in a terminal.
On a Mac, two:
- Microphone — so it hears you.
- Screen & System Audio Recording — so it hears them. There is no separate "system audio" permission on macOS; this is the one Apple puts it behind. Whispr reads no pixels.
Screen Recording needs the app restarted after you grant it, even once the toggle is on. macOS does not tell you that.
If you refused either by accident, macOS will not ask twice — the button opens the right System Settings pane instead.
On Windows, one: the microphone. Windows lets an application record the system mix without asking anybody, so there is nothing to grant for hearing them — that row simply passes, and it is passing rather than unchecked. The microphone prompt appears the first time Whispr opens the microphone, not at install. If you dismissed it, turn Whispr on under Settings › Privacy & security › Microphone.
If the call is on your phone
Whispr hears the other side from what this computer plays. A call held on a handset is not playing here, so their voice never reaches it — you would get a transcript of your own half, no questions detected and no answers, with every indicator green.
Three arrangements work:
- Join from this computer instead. If there is a dial-in bridge, join it from the desktop app — Zoom, Teams, Meet, Zoom Phone, Google Voice, WhatsApp or Signal desktop, FaceTime Audio on a Mac. The other side then plays through your output like any meeting, and nothing else needs changing. This is the reliable one.
- Route the phone call to the computer. macOS can take iPhone calls over
Continuity (FaceTime → Settings → Calls from iPhone); Windows can take
Android calls through Phone Link. Both should work, and neither is verified
here — prove it in a minute by starting a call and running
whispr verify audio --seconds 20, then checking the row for their side. On Windows the command cannot prove it, so watch the them meter move in the app instead. - Wire the phone in. Whispr can take the other side from an audio input
instead of your computer's own sound: run the phone's earpiece output into a
USB audio interface, then pick it under Settings → Your microphone → Where
the other side comes from. Your microphone stays you, so both sides are
still told apart by which wire carried them. Screen Recording is not used at
all on that setting, and echo cancellation is switched off on both inputs
while it is in use — wear headphones. From a terminal it is
whispr audio-devices --set-system <id>. - Not speakerphone. Putting the handset next to the laptop puts both voices on your microphone, and your microphone is structurally you — so their questions are recorded as things you said, no question is detected, and no answer ever appears. Whispr will now tell you when this is happening, but it cannot fix it.
Whichever you use, run the speech check on the pre-meeting screen. It is what lets Whispr tell your voice from the room, and without it anything your microphone picks up is recorded as something you said.
3. Add one transcription key and one answer key
Settings → Credentials. Everything runs on your own accounts: your keys, your billing, and we take no cut. Two are enough.
To hear the call — Deepgram. Sign up at console.deepgram.com, no card, and it starts with free credit. Whispr opens two streams (their voice and yours), so a 45-minute call is about 90 stream-minutes.
To answer — Cloudflare Workers AI. An account id and an API token with Workers AI · Read. The free daily allowance covers several calls a day.
Both screens tell you what the key buys, what it costs, and what happens if you
skip it. If you would rather nothing left your machine at all, a local Ollama
and a local whisper-server will do — slower, and free.
Press Test on each. A key that is stored but untested is the commonest reason a first meeting is silent.
4. Load your material
Prep. This is what makes an answer yours rather than generic.
For an interview: your CV, the job advert, and anything you have read about the company. For a sales call: what you sell (pricing, positioning, case studies) and who you are meeting. Drop in PDFs, Word documents, or paste text.
Your CV follows you between meetings. A job advert belongs to one meeting — which is why creating a new meeting stops the old advert grounding your answers, without deleting anything.
5. Run the pre-meeting check
Preflight. Do this once now, and again before anything that matters.
doctor proves the install; preflight proves this call — the microphone, the
level, this meeting's material, what your accounts have left, and what the
switched-on options will cost.
Do the level check. It asks you to read one sentence out loud, and it does two things: it proves your microphone is loud enough to be transcribed at all — a mic can be perfectly connected and far too quiet, and nothing else in the app can tell — and it teaches Whispr how loud you are, which is what stops a television or a colleague across the room being recorded as something you said.
Nothing on that screen refuses to start a meeting. A call happens once, and being told "no" at 09:59 is worse than a generic answer.
6. Rehearse before you rely on it
Practice. Do not let your first real question be your first question.
Pick a persona and press start. Whispr interviews you out loud, and afterwards tells you how long you ran, where the filler words were, and when your point actually arrived — measured from the transcript, not guessed. Ask for a model answer to compare yours against.
Ten minutes here is worth more than any amount of reading.
Try it for real, in the setup wizard, is the other half and proves the one thing practice cannot: that the other side's voice reaches this machine. Whispr listens to whatever your computer is playing, so a browser tab, a video or a voice note all count — you do not need a second device, though a phone in a throwaway meeting is the closest thing to the real arrangement. That sentence is also the limit: a call that only exists on your phone is not playing on this computer, so their side would not be recorded. See If the call is on your phone below. On one machine you can also press "I am speaking as the other side" while the test runs and play both parts yourself.
If that screen says Screen Recording is off, nothing the other side says can arrive, and no amount of talking will change it. Grant it there and restart Whispr — macOS does not apply that permission to a running app, even the first time.
7. The meeting itself
Create a meeting — the thing you prepare for. A second round at the same company is another session on the same meeting, so the prep and the history stay together.
Set Which round it is. For an interview that is behavioural, technical, system design or mixed; for a sales call, discovery, demo or pricing. Behavioural is worth setting deliberately — almost every behavioural round is scored against STAR, and that setting makes Whispr shape every answer that way.
Then press Start, and put Whispr on a second screen or beside the call window.
| Transcript | Both sides, live. Hover a turn to say who is speaking |
| Answer cards | One per detected question, answered by every engine in parallel |
| The record | What was decided, who owes what, the figures |
| ⚙ Quick settings | The microphone, the languages, whether it speaks |
Out of the box it says nothing aloud. You can turn reading-aloud on, and then it plays only into an output device you name — never your speakers by default, because anything your speakers play your microphone hears and your call transmits. Choose headphones, and wear them.
8. Afterwards
History has the transcript, which answers you used, and the record. Press Write a debrief for what went well, what to change, and what they raised — one model call on your own account, so it only happens when you ask.
If something is wrong
| It looks like | It usually is |
|---|---|
| Nothing is transcribed | The wrong microphone, or a transcription key that was never tested. Run Preflight |
| Whispr answers itself | Reading-aloud is on and playing through speakers. Quick settings → Read answers aloud → Off |
| Things you did not say | A television or someone nearby. Run the level check so it knows your voice |
| No answers at all | A general meeting starts in Listen. Switch to Answer |
| Answers are generic | No material loaded, or loaded on a different meeting |
whispr doctor checks everything an interview needs and says what to fix. The
Preflight tab does the same thing in the window, and goes further — it also
measures how loud you actually are, which no green tick can tell you.
If you want the terminal commands this manual uses, install them once from Settings → The whispr command; they are never required. Troubleshooting goes deeper.