Whispr

The Live view

Live is the screen to keep open during the call. On the left is the running transcript; on the right, one card per question the app heard, with a column for each answer engine you have switched on.

Press Start in the title bar to begin. It runs against the meeting named beside it, so check that name is the one you are about to walk into. Stop ends the session and writes it to History.

The transcript

Two speakers, and which is which is decided by where the sound came from rather than by how the voices sound: audio from your Mac's output is the other person, your microphone is you. Nothing is inferred, so nothing can be mistaken.

Rehearsing both sides on one Mac

That rule has one deliberate exception, and it is a button rather than a guess.

If you are testing the app by running both sides yourself — the interviewer in one browser, you in another — then when you speak as the interviewer it is still your microphone, so it would be recorded as you. Usually the meeting plays your voice back through the other browser and the app drops the duplicate; mute that tab and there is no playback to compare against.

I am speaking as the other side, above the transcript, records your microphone as the other person for as long as it is on. The header says Rehearsal — your mic is recorded as the other side. while it is, and I am the candidate again puts it back.

Two things worth knowing. It is off by default and never remembered, so a real call cannot inherit it. And those turns are marked in the stored session as a rehearsal capture, so History months later does not claim an interviewer said something you said yourself.

When a question is detected

A card appears with the question and a label for its kind — behavioral, technical, background or logistics. The enabled engines start answering in parallel, and each column fills as its answer streams in, with the time it took in the corner.

Detection deliberately leans towards emitting. A question that was not really a question costs one cheap model call; a question that was missed costs you the answer you needed.

An engine that has produced nothing yet still gets a column, showing either waiting… or the error it failed with. A slow or broken engine being visible is the point — an absent column would just look like a question nobody answered.

If the card found anything in your prep, a line underneath says how many snippets matched and opens to show them.

Typing a question yourself

The box along the bottom is the backstop. If something slipped past detection — or was asked over chat, or before you pressed Start — type it there and press Enter. It is answered exactly as a detected question is.

Choosing an answer

use this marks an answer as the one you actually used, and moves it to the front of the card so it is not somewhere you have to scan for mid-sentence. The same button clears the mark: the mistake worth being able to undo is choosing the wrong one.

Only one answer per question can be chosen. This is the most useful thing in History afterwards — "which answers did I actually use?" is the question that makes a review worth reading.

Asking for something else

Under a card that has settled, none of these fit? offers three:

Button What it asks for
shorter The same substance, sayable in about thirty seconds
more detail Fuller, with the specifics and numbers from your prep
different angle A different example or framing, not a reword

Each adds a new option to the card rather than replacing anything, so the answer you are reading from cannot vanish underneath you. again on a single answer asks that one engine for a different angle.

Every one of these buttons spends a call on your own account, which is why nothing is generated unasked. They are disabled while anything on the card is still streaming — an impatient double-click otherwise buys three near-identical options crowding the card you are trying to read from.

Two questions in one breath

When an interviewer asks two things at once, you get two cards grouped under a heading saying so. You can answer them separately without losing the fact that they arrived together.

The meters

While a session runs, two meters sit in the title bar — them and you.

A moving meter means audio is arriving, which is not the same as it being heard. A ⟳ beside a meter means that side's transcription is reconnecting and words may be missed; a ✕ means it is not being transcribed at all. Both keep showing until the channel recovers.

A flat meter with no mark is a capture problem rather than a transcription one — see Troubleshooting.

The cue, and voting on it

Above the answers, a card may carry a short cue — a bold head and three beats, like Scale · team of 4 → 30 | hired 12 | kept attrition at 8%. That is what the overlay shows you while you speak, and the engine that wrote it is named beside it.

Every enabled cue engine is asked at once and the first one back is shown, because a cue that arrives after you have started talking is worthless. Fastest is not the same as best, though, so the 👍 is how you say a cue was actually good. Press it and the vote is counted towards that engine.

The vote is on the cue, not the engine, so pressing twice is not two votes and the count survives you renaming or replacing the engine later. Settings → Anchor cue shows what the votes add up to: races won, voted good, refused, and median latency per engine, ordered by votes first and speed second. Nothing switches automatically — you choose, once the numbers say something.

If a cue engine keeps returning prose instead of the format, its cues are refused and nothing is drawn rather than a paragraph appearing on a window you are only glancing at. That is what the Refused column counts, and a few is ordinary — Late is the separate column for a cue that timed out, which is a different problem with a different fix.

Every engine's attempt is recorded, including the ones that arrived second. That is what the Beaten column is: without it, only the fastest engine could ever collect votes, so the numbers could never tell you that a slower one writes better cues.

Once the votes say something, Settings offers to ask one engine from then on — see Letting the votes decide. It changes which engine is asked, never what your engines cost.

Pausing, muting, and stopping

Above the transcript is one bar with everything you need mid-meeting, and it says in words whether Whispr is recording.

Pause stops Whispr listening to both sides without ending the meeting — a break, a side conversation, somebody sharing a screen you would rather not have transcribed. Resume continues into the same recording, so it stays one meeting rather than two halves. It is also in the title bar beside Stop, so it can be reached from any tab: audio is dropped before it is measured, before it is transcribed and before any engine sees it, so nothing is recorded, nothing is billed, and the meters go flat while it is on.

Mute my mic stops Whispr hearing you while the other side is still heard and answered. Worth knowing: muting yourself in Zoom or Meet does not do this — Whispr opens the microphone directly.

Answering / Listening only decides whether Whispr replies at all. Off, the transcript and the goal checklist keep running and nothing is generated or billed. A general meeting starts in Listening only, because most of a working meeting is not a question that needs an answer.

Stop ends the meeting and writes the recording. The transcript and cards stay on screen afterwards — that is the record, not a sign it is still running; the bar above them will say it has stopped.

Quick settings

⚙ Quick settings opens a panel over the meeting, holding the handful of things that actually go wrong in one — and nothing that needs reading, spends money, or cannot be undone. Everything else is still in Settings; this is what is fixable in the ten seconds before the next question.

Microphone, and its input level. The level is the one that catches people out: the same microphone at 60% and at 100% looks identical in any device list and is ten decibels apart in the audio, and the quiet one produces no transcript at all while every meter still moves. Changing the microphone takes effect on the next session — the current one keeps the device it opened.

Whispr plays through, and Read answers aloud. Nothing is ever spoken without a device chosen here, and there is no fallback to the system output, because on a laptop the system output is the built-in speakers — which your own microphone hears. If a meeting starts sounding like it is answering itself, this is the switch.

Languages spoken. Empty means detect, which is right for a genuinely multilingual room. On a quiet one, detection invents words in a language nobody present speaks, and naming yours stops it. Pick as many as apply, or type a code the list does not offer. One named language is sent as itself; two or more are sent to Deepgram as multilingual, and to whisper as the first of them, because whisper has no multilingual mode. Detect instead puts it back.

Headless, the same setting is whispr transcriber --languages en-GB,lv-LV, and whispr transcriber --languages detect clears it.

Transcribed by switches between the four transcribers. Like the microphone, it takes effect on the next session.

The meeting kind is shown and cannot be changed here, which is deliberate rather than an omission. A meeting keeps the kind it was created as: switching it after material is loaded would leave prospect notes grounding an interview, or a CV grounding a sales call, and neither is something you could undo. To work in a different kind, create a new meeting — the old one keeps its material and stays in your history.

The record — what a meeting keeps instead of answers

An interview earns its keep by answering questions: somebody asks, your engines race, you read one out. Most of a working meeting has no questions in it at all, and running that machinery over one produces a card per sentence saying "not asked for this one".

So a meeting keeps a record instead, above the cards: what was decided, what somebody now owes, and the figures — numbers, dates, names — you would otherwise have to ask for twice. It is useful thirty seconds later, when somebody says "what did we agree?", and it is what your debrief is built from afterwards.

Whispr also asks each part of a long turn to stand on its own. When somebody asks two things at once you get two cards — but a turn broken into fragments no longer produces a card for the halves that were never questions, which is where a burst of near-identical cards used to come from.

Most turns add nothing, and that is deliberate. A record that logs everything is a transcript, and there is one of those on the left. The lines are written by the same watcher that keeps your goal checklist, so the record costs no extra model call — it rides on one that was already being made every turn.

If the rail says the watcher stopped, the record stops too — it is written by the same thing. That watcher is deliberately the cheapest engine you have configured, which also makes it the slowest; if it keeps timing out, give the classifier role a faster model.

The watcher needs a model that can return JSON, because it answers in a small structured reply rather than prose. Small models often nearly manage it — leaving a value out where the answer is "nothing" — and Whispr reads those anyway. One it cannot read is ignored, on the assumption the model was being chatty; several in a row and the rail says so, because a record that quietly writes nothing is worse than one that admits it stopped.

whispr engines test now checks this before a meeting rather than leaving you to find out during one: a classifier that answers prose but not JSON is reported as failing, and names what it costs — the record, the objection rail and goal ticking.

Hover a line to remove it. You are the only judge of the record: a line you disagree with is wrong, whatever Whispr heard.

Answers are not switched off in a meeting — Answer works exactly as it does in an interview, and the typed question box always does. A general meeting just starts on Listen, because most of one is not a question.

Saying who is speaking

A meeting has more than two people in it, and Whispr cannot tell their voices apart — it deliberately does not try. What it knows structurally is the side: anything on your system audio is the other side, anything on your microphone is you. Which of four people on that side just spoke is something only you know.

So hover any of their turns and press + name. Two buttons, because there are two things you might mean:

Names you have already used appear as one-click buttons, and they carry across rounds of the same meeting, so you type each person once. Got it wrong? Open the same panel and press Clear — it undoes forward, exactly the way it was applied.

The name is shown next to the side, never instead of it, and it can never move a turn from one side to the other.

This is what makes the record worth having afterwards. "They said they would send the contract" is a sentence; "Ana said she would send the contract" is an action item with an owner. The names go into the debrief, into the commitments Whispr writes down while the meeting runs, and into the suggestions on the rail — which is why it is worth naming people early rather than at the end.

When somebody says your name

In an interview everything the other side says is aimed at you, so a question mark is the whole signal. In a meeting with five people it is not, and the one reliable marker that a turn is yours is your own name — the one in Settings.

Whispr watches for it, and being named counts even when nothing sounds like a question: "Intars — over to you" and a flat "Intars?" are both moments somebody is waiting for you to speak.

The question arrives on either side of the name, and often in the turn after it:

The last one is why a bare naming keeps looking forward for a few seconds: the turn that follows it is treated as addressed to you too. It carries to one turn only, so a conversation that moves on to somebody else does not keep counting.

Leave the name in Settings empty to switch all of this off.

The steering rail

Above the answer cards, Whispr keeps one small panel about the conversation rather than about any one question. It is silent unless there is something worth saying, which is most of a meeting.

It holds three things:

Whispr will not offer a second suggestion within thirty seconds of the last — a suggestion arriving while you are still reading the previous one is one you have no time to act on. An objection overrides that, because an objection is the moment a call turns, and hearing about it a minute late is hearing about a call you have already lost.

With assist off, the checklist and the objection keep updating and no suggestion is written — nothing is generated and nothing is billed. Reading the room while the record keeps score is exactly what listening-only is for.

If the watcher stops working, the rail says so plainly. Silence that looks healthy would be worse than the failure: your transcript and your answers are unaffected either way.

Read mode

Press Read on a card and the answer becomes an autocue: one breath-sized line, large, with the next line dimmed underneath. It advances when you reach the end of the line — Whispr is already transcribing your microphone, so it can be told rather than guessing at a scroll rate.

That is the whole reason this reads differently from a card. Reading a paragraph aloud sounds like reading; being paced to your own speaking does not.

It is a mode, not a replacement — the cards are one keypress away, and the other answers are still there to switch to.

Turning the answers off

The Assisting / Listening only button decides whether Whispr answers at all.

Most of a meeting is not a question that needs an answer. Turned off, everything that makes the recording valuable keeps running — capture, transcript, question detection, the history — and nothing talks back: no answers, no cue, nothing spoken. Small talk stays small talk, and you can read the room without three engines racing on your account.

The question box still works while it is off. Typing a question is asking for help, so that always answers.

Muting your microphone

Mute my microphone stops Whispr hearing you, without leaving the meeting.

Worth knowing why this exists: muting yourself in Zoom or Meet does not stop Whispr. It opens the microphone directly, so a side conversation with somebody in the room is otherwise recorded, transcribed, and attributed as you answering the interviewer — which is also the signal Whispr uses to decide a question has been answered.

While muted, nothing you say is transcribed, stored, or billed to your transcription provider. The other side is still heard and still answered throughout. It is offered only in a real meeting: a practice run hears nothing but your microphone, so muting there would end the exercise.

Hearing an answer instead of reading it

If you have turned on Read answers into my headphones (Settings), the fastest answer is spoken into the device you chose. It begins in well under a second, while the rest is still being generated, so it is speech you can start acting on rather than something that arrives after the moment has gone.

Only one answer per question is ever spoken, and only one engine's — the others stay on screen. If the speech stops part-way, a notice says so and no second engine starts speaking over it.

Nothing is spoken through your speakers, ever, under any failure. That is explained where you switch it on, and it is the reason there is no "system default" device to pick.

The overlay

The floating window is for the moment you are actually speaking. It is a glance, not a read, and it is described in The overlay.