Case file 05 · Tooling · macOS
LocalFlow
Push-to-talk dictation for macOS that never leaves the laptop. Hold a key, speak, let go, and clean text appears at the cursor in whatever app you're in.
- Year
- 2026
- Role
- Solo: design and build
- Read
- 3 min
- Live
- Private repository · shipped
Context
Push-to-talk dictation for macOS that never leaves the laptop. Hold a key, speak, let go, and clean text appears at the cursor in whatever app you're in.
I built it because the tool I was paying for was cloud-only. Every word I dictated went to someone else's servers, at fifteen dollars a month, and it turned out the same job could be done entirely on the laptop.
The problem
Running everything locally means the model has to be small enough to be fast on a laptop and good enough to be worth using. Dictation is unforgiving: a wrong word takes longer to fix than it saved.
It also has to disappear. A dictation tool that steals focus, or needs its own window open, has already failed. You're dictating into something else, and the app's own UI is the one thing that must never get in the way.
What I built
A menu-bar app in Python. Speech recognition runs on the laptop through Parakeet on Apple Silicon. The text is cleaned up by rules and a personal dictionary, with an optional pass through a local language model that's off by default. Text goes in through the clipboard and a simulated paste, which works in every app, not just the ones that cooperate.
Around that: a transcript history in a local SQLite database, snippets that expand short phrases, per-app styles so dictation reads differently in a terminal than in a document, and an insights pane that works everything out from your own history on your own machine. The on-screen indicator is a panel that never takes focus, so it can show what's happening without pulling you out of what you're typing into.
Decisions & iterations
The language-model pass is off by default and skipped completely for anything under five words, because having a model rewrite a short phrase costs more than it helps. Its replies are checked, not trusted. A cleanup step that can make things up is worse than no cleanup at all.
Auto-learn was redesigned around things people actually do. It learns from a deliberate copy instead of silently reading text fields, because the version that watched everything was both creepier and less accurate. Hands-free became a lock on push-to-talk instead of a separate mode, which got rid of a whole class of bugs. Auto-stop on silence was considered and rejected, because it cuts people off mid-thought.
The most useful bug was in the numbers, not the audio. The insights pane was counting the model's words as mine, and a words-per-minute figure was quietly calculated from only half the history. I found both by measuring, not by reading the code, and in both cases the fix was to make the figure say what sample it's based on.
Outcome
It's what I dictate with. 477 tests cover it, including the insights chart, which is tested by running the real UI file and checking the bars it actually drew. It replaces a fifteen-dollar-a-month subscription, and the audio has never left the machine.