Whispr

Practice

Practice is the app the other way round. Instead of listening for someone else and handing you an answer, it asks the question and listens to you.

Nothing is sent to an answer engine while you practise. Whispering the answer at somebody rehearsing defeats the exercise and spends their money doing it, so the app does not do it.

Practice runs against the meeting you have selected, using the material you loaded for it. Pick the meeting first.

Two modes

On screen — the question appears, you answer aloud, and you press a button when you are done. Silent, and nothing is spoken at either end.

Spoken — a voice agent asks out loud, listens, and follows up where an answer was vague. It needs your Deepgram key, and it needs headphones.

Both need transcription, to hear what you said. Both also use one of your answer engines — not to answer for you, but to write the questions in the first place and to judge how each attempt went. With no working engine the run still starts: you type your own questions, and you get the measured numbers below without the judged sentence.

Answering on screen

The header counts the questions and says what the app is doing: writing your questions, listening, or working out how that went. A dot lights while your microphone is hearing you.

Four buttons under each question:

Button What it does
I've finished answering Critiques this answer and moves on
Again, same one Critiques this answer, then asks the same question again
Skip this one Moves on without an answer
Stuck? Give me a cue A head and three beats to start from — never the answer

Stuck? Give me a cue

Practice deliberately shows you no answer. An answer on screen makes you a reader, and reading is the thing this is meant to train out of you — so nothing here asks an answer engine, which is also why practice costs so much less than a live call.

Being stuck with nothing at all is not useful either, so this asks for the cue instead: the same head-and-three-beats the overlay shows during a real call. It is a way in, not a way out, and the format is enforced rather than requested — prose is refused, so it cannot turn into a script however the engine misbehaves.

Each press is counted, and if you used any the summary at the end says how many. Needing fewer of them is the point of practising, so it is there as progress to watch rather than as a telling-off — and it stays out of the summary entirely on a run where you never pressed it.

The turn ends when you press, not when you go quiet. There is nobody else talking to mark the boundary, and cutting you off after a pause would cut off anyone who stops to think. A short grace window follows the press, because transcription runs slightly behind the speaking and the tail of an answer is usually where the point finally arrives.

Seeing what good sounded like

Under each critique there is Show me a strong answer. Press it and one example answer to that question is written — 120 to 180 words, the point in the first sentence, using your own prepared material, and ending with a line naming what your own attempt buried or left out.

It is deliberately after the critique and never before. Practice shows you no answer while you are still speaking, because an answer on screen makes you a reader and reading is the habit this is meant to train out of you. Afterwards it is the opposite: it is the lesson. A critique on its own can tell you the point arrived at 2m40s without ever showing you what arriving at 5s sounds like.

It says "a strong answer" rather than "the answer" on purpose. There is no single correct answer to a behavioural question, and one that claimed to be would be worse than useless.

A few things worth knowing:

Reading the critique

The numbers come first, and they always come. They are measured from the recording rather than judged, so they are exact and they arrive even when the judging engine is unavailable:

The sentence underneath is the judged half: whether it answered what was asked, and whether it drew on your prepared material. If that never arrives you are told so plainly — "no verdict on the content this time" — because the numbers above are still real and you should know which half is missing.

"The point arrived at 2m40s; everything before it was preamble" is the useful sentence, and it is why time-to-point is worth more attention than the total length. There is no encouragement in here, deliberately.

What you said opens the transcript of the attempt. When you have used "Again, same one", every attempt is kept and shown separately — comparing the two is the whole reason that button exists.

At the end, Finish and summarise writes up the run: how much you talked, your filler rate across everything, your longest answer, and the one that took longest to get to the point.

The spoken interviewer

Choose the voice, then who it is being and how hard it presses.

Which voice

Two ways to have an interviewer that talks, and they are billed on completely different scales:

Voice What it needs What it costs
Deepgram Your Deepgram key — the same one that transcribes your real calls Connection time, a few cents a minute for as long as it is on the call, so a 30-minute mock is a couple of dollars
Workers AI A voice Worker you deploy to your own Cloudflare account, plus the token you set on it Neurons: about 1,100 a minute, so a 30-minute mock is roughly $0.36 — but the free daily allowance only covers about nine minutes of it

Deepgram is the default and needs nothing deployed. The Workers AI option is much cheaper per minute and needs a Worker you maintain — docs/voice-worker.md in the repository is the whole procedure. The screen shows the costs of whichever one is selected at the moment you select it, with the free allowance and what a call actually consumes folded underneath.

The free-allowance line is the thing to read twice. Nine minutes a day is a short rehearsal, and on the free Cloudflare plan a run that crosses it does not get billed — inference is refused until 00:00 UTC, so the interviewer goes quiet mid-question. The socket stays up, so what you get is your Worker's own reason rather than a reconnect or a wrong-key warning, and the run keeps recording everything you said.

Everything after the choice is identical. The persona, the difficulty, the critique, the transcript and the history are the same either way, and so is the headphones requirement below.

Who it is being

For an interview: a friendly recruiter screen, a sceptical technical lead, or a senior stakeholder from outside your function. For a sales call: a sceptical economic buyer, a technical evaluator, or someone thinking about procurement. They ask the same subject differently, which is the point — practising only against the friendly one is how people get surprised.

Difficulty runs gentle, standard, intense. Intense follows up on every vague claim and will interrupt a rambling answer rather than waiting politely.

Once it starts, just answer. It decides for itself whether to press or move on.

Headphones are required for the spoken mode

The interviewer speaks aloud through this machine. On speakers, your microphone picks that voice up at close to half the level it comes out at, and the agent takes it for you starting to answer — it cuts itself off mid-question and talks over you.

There is no software guard here on purpose: the agent does its own turn-taking straight off the microphone stream. That is true of both voices — it is a consequence of the agent listening to your microphone itself, not of which service is behind it. The app asks you to confirm headphones once and then remembers.

Practice runs are kept

A practice session is recorded like any other, and its critiques are in History. It is kept because the critique is the entire point of having done it.

The only thing that is erased is the rehearsal inside the setup walkthrough — the one where you call yourself from a phone to check the audio works. That is a test of the equipment, not of your answers, and it never reaches History.