A prompter for macOS · private beta
The line, before you need it.
Souffleur listens to both sides of your call on your own machine and puts the next twelve words on screen about 2.6 seconds after the question lands. That is inside the pause you were going to take anyway.
Median of 23 turns of live speech on one Apple silicon Mac. p90 3.4 s.
From the last word you hear to a finished transcript on your own GPU.
Median time to twelve words you can start saying. Measured on live speech, p90 3.4 s.
Of audio uploaded. Not sampled, not buffered to a server, not kept.
Measured in the shipping app on my own calls, not in a benchmark. I built this because I needed it, and I have been using it since July. Here is how I measured it, including the four times my own instruments lied to me.
Why this exists
You cannot take notes and stay in the room.
Every good answer you have ever given on a call was assembled while you were also listening, watching a face, and deciding whether to push. The moment you look away to think, the other person hears it.
Notetakers arrive too late
A summary after the call is a record, not help. The decision you needed support for was made forty minutes ago.
Live tools are too slow to be live
Upload the audio, transcribe in a datacenter, come back. By the time words appear, the silence has already been noticed.
And they keep the call
Cloud copilots require an account, a bot in the meeting, or a recording on someone else’s disk. For a lot of conversations that is simply not allowed.
Where your call actually goes
Audio never leaves the machine.
Not a slogan. A line-by-line account, including the one part that does go out. If a tool will not tell you this plainly, assume the worst about it.
Captured, held in memory, never written to a server. No bot joins your meeting.
A local model runs on your Mac’s own GPU. This is also why it is fast: there is no upload.
Stored in a plain database in your home folder. Delete the file and it is gone.
To answer, the recent lines of the conversation are sent to the assistant you already pay for, under your own key. Nothing else is, and nobody sits in between: no account with me, no telemetry, no copy kept. Pause the assistant and not one word leaves.
On your screen
A panel only you can see.
It floats near your camera, so your eyes stay where the other person expects them. It never takes focus, so your typing still goes where you were typing. On the platforms tested it is excluded from screen sharing, and it draws nothing anywhere else.
listening
Six weeks, eleven teams. Retention held at 91 percent.
ready in 2.4 s
The panel is drawn to scale. Everything behind it is your call, untouched.
How it works
Three seconds, spent carefully.
The whole design is a fight over one budget: the natural pause after a question. Every number below was measured in the shipping app, on live speech, not in a benchmark.
It hears both sides, and can tell them apart
Your microphone and the call’s own audio are captured as two separate streams, so it always knows who asked and who answered.
It knows the question ended from the words, not the silence
Waiting for a long pause is how other tools cut people off mid-sentence. Souffleur reads the transcript instead, and has it ready 418 ms after you stop hearing them.
The answer streams onto a panel only you can see
Twelve usable words at 2636 ms, median, so you can start talking without reading ahead. The panel floats near your camera, never takes focus, and is excluded from screen sharing on the platforms tested.
Before you ask
The questions I would ask.
Short answers, including the ones that are not flattering.
What do I need to run it?
A Mac with Apple silicon. Speech recognition runs on the GPU, so an Intel Mac will not do. You also grant it screen and microphone permission once, the same way any recorder asks.
Does it work without internet?
Transcription does, entirely on your machine. Writing the answer needs a language model, and that one call goes out. Pause the assistant and nothing leaves at all.
Which meeting apps does it work with?
Any of them, because it listens to your system audio rather than plugging into an app. Screen-sharing invisibility is a separate matter: verified two-way on Zoom and Google Meet, on a real recording and a real preview. Teams is not verified yet. I would rather say that than imply it.
Which AI does it use, and whose account?
Yours. It plugs into the assistant you already pay for rather than reselling you one: Claude works today, and ChatGPT, DeepSeek and Gemini are the ones I intend to wire up next, in whatever order the waitlist asks for. You pay your provider directly and I take no cut of your usage. Speech recognition costs nothing either, because it never leaves your Mac.
Are those 2.6 seconds the same on every model?
No, and I am not going to average them into one flattering number. Every timing on this page was measured on Claude. A different provider means a different first token and a different cache, so each one gets measured separately and published as its own figure. The method, and what it cost me to get it right, is written up here.
Is this for cheating on job interviews?
It is built and sold for meetings and sales calls, where the person on the other end wants you to have the answer. I am not going to pretend the same software could not be pointed elsewhere, but that is not the product I am making, and it is not what the roadmap serves.
Do you keep my calls?
No. There is no server to keep them on. Audio and transcripts live in a file in your home folder, and deleting it is the whole story. I never see them, and there is no account for me to attach them to.
When can I actually install it?
No date, and I will not invent one. It works on my own calls today. Whether it becomes something you can install depends on whether enough people answer this page.
Private beta
It is not for sale yet. Tell me if it should be.
There is nothing to download today. I built this for my own calls, it works, and I am trying to find out whether anyone else wants it before turning it into a product. Three questions and an email address. That is the whole ask.