Hold a hotkey, talk, release - polished text lands wherever your cursor is.
No subscription, no account, and your audio never leaves your machine unless you opt in.
Everything the commercial dictation apps charge a monthly fee for, running entirely on your own hardware. Your voice never touches someone else's server.
Hold Ctrl+Shift → talk → release → polished text appears at your cursor.
A green microphone in the system tray means it's ready. First launch downloads the Whisper model once; after that it works fully offline.
| Echo Flow | Typical cloud dictation app | |
|---|---|---|
| Price | Free · MIT | $10 to $30 a month |
| Your audio | Stays on device (cloud is opt-in) | Uploaded every time |
| Works offline | Yes | No |
| Account required | No | Yes |
| Learns your corrections | Yes, locally | Limited, and cloud-side |
| Knowledge layer (notes · tags · graph · search) | Built in | No |
- Features
- How it works
- Screenshots
- Installation
- Daily workflow
- Configuration
- Voice commands (experimental)
- Teacher-model distillation (optional)
- Privacy & data flow
- Repository layout
- Troubleshooting
- iOS
- License & cost
| Feature | What it gives you |
|---|---|
| Local transcription | OpenAI Whisper on-device (tiny through large-v3-turbo, or auto by hardware). Nothing uploaded. |
| Local cleanup | A small LLM via Ollama (qwen2.5:3b-instruct) polishes raw output: punctuation, capitalization, filler removal. With no Ollama and no key you get deterministic rules-only cleanup, which still handles casing, punctuation and fillers. |
Re-paste (Ctrl+Shift+Win) |
Drops your last dictation into a new window: say it once in Slack, paste it again in email. |
| Snippets | Short codes expand after cleanup: btw → "by the way", lgtm → "looks good to me". Case- and word-boundary-aware. |
| App-aware profiles | Cleanup style can follow the focused app: casual in Slack, symbol-aware in VS Code, full sentences in Gmail. Ships with every profile on medium; differentiate them in Dashboard > Style. |
| Casing control | Learns a word's casing from one edit (tiktok → TikTok sticks forever, possessives included) and flattens Whisper's accidental "Every Word Capitalized" back to normal sentence case. |
| Hallucination guard | Length + RMS gate drops silent/short clips so Whisper can't invent "thank you for watching"; if the model goes off-track, your raw words are pasted (casing-normalized) instead. |
Every correction you make through the tray menu feeds back into cleanup. After a few hundred dictations it knows your jargon, names, and writing style.
flowchart LR
D["You dictate"] --> G["Self-grade<br/>0-100 quality"]
G --> H[("history.db")]
E["You fix it once<br/>(tray edit)"] --> L["Learn<br/>casings + patterns"]
L --> H
H -->|"few-shot + learned rules"| C["Next cleanup<br/>gets smarter"]
C --> D
L -.->|"enough signal"| F["LLM-free<br/>'learned' mode"]
| Capability | Detail |
|---|---|
| Self-grading | Every dictation gets a 0 to 100 quality score from four signals (Whisper confidence, hallucination guard, semantic coherence, pattern coverage). |
| Self-improving loops | Online weight calibration (SGD against your edits) + exponential pattern decay (14-day half-life) so stale jargon fades. |
| LLM-free mode | A learned cleanup provider built from your past corrections. It runs with no LLM at all once it has enough signal. |
| Auto-phasing | Progresses from Whisper + Ollama cleanup to fully LLM-free cleanup as your history grows. |
Two features on the My Voice page, different jobs:
| What it does | |
|---|---|
| My Voice (dictation) | A light-touch pass after cleanup that nudges your dictated text toward your own phrasing. Off by default; shadow previews it without changing anything. Meaning is preserved exactly, and the rewrite is dropped if it drifts at all. |
| Humanize (paste-in) | Paste AI-written text and get a human version back with the LLM tells stripped out (em-dash rhythm, delve / moreover / a testament to, "it's not just X, it's Y", tricolons, hedging stacks). Pick how it should sound. |
The humanizer gives you three targets, and needs no setup to start:
- A natural human. Plain prose with the AI tells removed. The default.
- Me. Also match your writing samples (add a couple on the same page). With none, it falls back to the natural-human rewrite and tells you why.
- A specific tone. Casual, professional, friendly, plain, confident, concise.
You also steer it: a strength slider (light / balanced / aggressive) and a custom tone box, plus a Try again button to re-roll.
It always gives you a result. A risky but readable rewrite (a number changed, meaning drifted) is shown with a warning rather than dropped; only genuinely broken output falls back to your original. It rewrites one paragraph at a time so structure survives, checks every number in both directions (a dropped figure is as wrong as an invented one), and shows a word-level diff plus an "AI tells: N → M" score so you can see exactly what it stripped. When the small default model struggles on a paragraph it escalates once to the next-larger model you have installed. All local by default.
| Feature | Detail |
|---|---|
| Notes | Pin any dictation to promote it to a long-lived knowledge object with title + description. |
| Tags | Three-signal auto-suggestion (cluster, similar, concept) with manual confirm. |
| Action items | Regex extraction of TODO-style phrases, with a blocklist for daily drivel. |
| Knowledge graph | D3.js force-directed view of dictations/notes/concepts, with tag filters, search, and a quality slider. |
| Semantic search | Find past dictations by meaning, not just keyword. |
| Review queue | Worst-quality-first list of un-edited dictations, one click from the tray. |
A native local window (Flask + PyWebView, server-rendered, no telemetry) at
http://127.0.0.1:8766 for managing everything: history, insights, custom
vocabulary, snippets, learned casings, style profiles, transforms, scratchpads,
voice-action shortcuts, settings, light/dark theme, and notification sounds.
- Loopback-only. Binds to
127.0.0.1only; the loopback boundary is the auth model, with aHost:header check on every request as DNS-rebinding defense. - Never blocks dictation. Flask runs in a daemon thread; the window runs in a separate process. A crash in either can't wedge the hotkey path.
- Works offline. No SPA framework, no Node toolchain.
- Keyboard-first. Press ⌘/Ctrl+K anywhere in the dashboard to jump to any page; the sidebar collapses to a drawer on narrow windows.
Open it from Tray → Open Dashboard, run_dashboard.bat, or a browser.
Everything in the dashed box runs on your machine. The only paths that leave it are opt-in and gated behind your own API key.
flowchart LR
A["Hold Ctrl+Shift<br/>push to talk"] --> B["Whisper STT<br/>local CPU / GPU"]
B --> C{"Cleanup"}
C -->|"default · local"| D["Ollama LLM"]
C -.->|"PE mode · opt-in"| E["Groq / Anthropic<br/>cloud, your key"]
D --> F["Paste at cursor"]
E -.-> F
F --> G[("history.db<br/>local, learns from you")]
G -.->|"few-shot examples"| C
P["iOS keyboard"] -.->|"Wi-Fi bridge"| B
DB["Dashboard<br/>127.0.0.1:8766"] --- G
subgraph LOCAL["Your machine · no network"]
B
C
D
F
G
DB
end
What happens when you talk, the live path end to end:
sequenceDiagram
autonumber
actor You
participant H as Hotkey listener
participant R as Recorder
participant W as Whisper · local
participant P as Cleanup · local
participant Cur as Cursor
You->>H: hold Ctrl+Shift
H->>R: start capture
You-->>R: speak…
You->>H: release
H->>R: stop
R->>W: audio buffer
W->>P: raw transcript (~0.5s)
P->>Cur: polished text pasted (~1s)
Note over R,Cur: end-to-end ≈ 1-2s · nothing leaves your machine
Note
Captured against a seeded demo database, so none of this is real dictation data.
Home. Your dictation inbox, quality-scored, with live time-saved, acceptance and latency stats.
Outcomes. How Echo shows up in your work: words per minute, the fixes it made, your app mix, a streak heatmap, and a quality trajectory.
![]() Knowledge graph. Dictations, notes & concepts, force-directed. |
![]() Dictionary. Learned casings ( github → GitHub) & custom vocabulary.
|
Privacy. A local-only audit ledger of exactly what touches the network. By default, nothing.
Fastest path: download EchoFlow-Daemon-Setup-<version>.exe from the
latest release and
run it. It installs the daemon with optional Windows autostart, and a .sha256
sits next to every asset so you can verify the download. The from-source path
below is what CI tests and is the way to go if you want to hack on it.
- Windows 10/11, or macOS 13+ from source (the installer is Windows-only; see macOS)
- Python 3.11+ on your PATH
- (Recommended) Ollama for local LLM cleanup
- (Optional) an NVIDIA GPU or an Apple silicon Mac, which Whisper uses automatically if present
git clone https://github.com/JOhnsonKC201/Echo_FLOW.git
cd Echo_FLOWscripts\setup.batOn macOS:
scripts/setup.shCreates a Python venv and installs dependencies from requirements.txt.
run.batOn macOS:
./run.shFirst launch writes a config.yaml in the repo root, copied from the factory
default in packaging/default/config.yaml. That copy is yours to edit and is
gitignored, so a clone can never inherit anyone else's settings: you start
local-first, exactly as described below. First launch also downloads the Whisper
model (a minute or two). When the green microphone appears in your system tray,
you're ready. Transcription runs
locally and nothing is uploaded.
Raw Whisper output gets a light polish from a local LLM via Ollama. Install Ollama, then pull the default model:
ollama pull qwen2.5:3b-instruct-q4_K_MThis is the default (cleanup.provider: ollama). Echo Flow starts Ollama for
you if it is installed but not running, so you do not have to add it to Windows
startup yourself.
No Ollama and no API key? Echo Flow still works. It falls back to rules-only cleanup, which is deterministic and needs no model at all:
| Rules-only cleanup does | It does not |
|---|---|
| Capitalization and end punctuation | Fix grammar or tense |
| Flattens Whisper's "Every Word Capitalized" runs | Reword or restructure |
| Collapses double-transcribed phrases | Add anything you did not say |
| Removes fillers: "um", "uh", "er", and comma-fenced "you know" / ", like," | |
| Applies any casings and phrases it has learned from your edits |
It is purely subtractive, so it can never put a word in your mouth. You lose
grammar cleanup, not dictation. Set cleanup.strip_fillers: false to turn the
filler pass off.
Regular dictation is 100% local. The one built-in cloud path is
Prompt-Engineering mode (Ctrl+Shift+Alt), which rewrites a short spoken
idea into a full engineered prompt using Groq, with your own key:
setx GROQ_API_KEY "gsk_..."Close and reopen your terminal so the variable loads (free key from https://console.groq.com, no credit card). Without a key, PE mode falls back to local Ollama. The same key powers the optional teacher-distillation loop.
| Command | Purpose |
|---|---|
run.bat |
Launch the daemon manually |
INSTALL.bat |
First-time setup with Windows autostart |
RESTART.bat |
Kill and relaunch. Run this after editing config.yaml or upgrading |
run_dashboard.bat |
Open the dashboard window |
UNINSTALL.bat |
Remove the autostart shortcut and optionally wipe data |
scripts\run_tests.bat |
Run the pytest suite |
On macOS the same jobs are ./run.sh, scripts/install_autostart.sh,
./restart.sh, ./run_dashboard.sh, scripts/uninstall_autostart.sh and
scripts/run_tests.sh.
Important
After pulling new code, run RESTART.bat (./restart.sh on macOS). The
daemon loads code once at startup, so fixes don't take effect until the
running tray process is relaunched.
Echo Flow runs from source on macOS 13 or later. What differs from Windows:
- Three permissions. The first hotkey press, the first recording and the
first paste each need one: Input Monitoring (to see the push-to-talk
keys), Microphone (to hear you) and Accessibility (to paste at your
cursor), all under System Settings, Privacy & Security. Grant them to
whatever runs Echo Flow: Terminal, iTerm, or python. Until all three are
granted the hotkey looks dead, the recorder hears silence and the text only
reaches the clipboard, with no error anywhere. The daemon lists whatever is
still missing when it starts, and
.venv/bin/python scripts/mac_permissions.pyanswers the same question without starting it (--openjumps to the right Settings pane). - Hotkeys. The default
ctrl+shiftworks as-is.cmd,commandandoptionare accepted inhotkey.comboif you would rather hold Command. - Voice commands use Mac chords. "Computer, undo that" presses Command+Z, "redo" presses Command+Shift+Z, "go to top" presses Command+Up. The command list is the same as on Windows; only the keys differ.
- App-aware profiles see the application name (Slack, Code, Safari) and, once Screen Recording is granted, the window title too. macOS gates window titles behind that permission; the app name alone is enough for the profiles.
- Whisper runs on the GPU on Apple silicon.
scripts/setup.shinstalls mlx-whisper on M-series Macs anddevice: autopicks it, with the samelarge-v3-turbomodel the CUDA path uses (an mlx-community conversion, downloaded on first start). If it cannot run, startup says so and falls back to faster-whisper on the CPU with thebasemodel. Intel Macs take that CPU path directly. MLX itself needs macOS 13.5 or later. Beam search and the built-in VAD are faster-whisper features; the mlx path decodes greedily and relies on the recorder's own silence guard. The startup lineWhisper ready: large-v3-turbo on mlxconfirms the GPU path is live. - Start at login.
scripts/install_autostart.shinstalls a per-user LaunchAgent (~/Library/LaunchAgents/com.echoflow.daemon.plist) that runs the daemon from this checkout at login and relaunches it if it crashes; a quit from the menu bar stays quit.scripts/uninstall_autostart.shremoves it. launchd runs the daemon as the Python interpreter rather than Terminal, so macOS asks for the three permissions again the first time and the grants go to python; the daemon's console, including the permission report, islogs/launchd.out.log. - Restart after changes.
./restart.shis theRESTART.battwin: through launchd when autostart is installed, otherwise it stops the daemon by its PID file and runs./run.shagain. - No installer yet. Without autostart,
./run.shin a terminal starts the daemon and the icon appears in the menu bar; quit from that menu. - Launchers:
scripts/setup.sh,./run.sh,./restart.sh,./run_dashboard.sh,scripts/run_tests.shand the two autostart scripts are the shell twins of the.batfiles above.
| Gesture | What happens |
|---|---|
| Ctrl+Shift (hold) | Record; release to transcribe + paste at the cursor |
| Ctrl+Shift+Win (hold, release) | Re-paste the last dictation into the current window |
| Ctrl+Shift+Alt | Prompt-Engineering mode: speak an idea, get a full engineered prompt |
| Tray icon | Pause, edit the last dictation, open the review queue, history, knowledge graph, dashboard |
It learns as you go. Every correction you make via the tray "edit last
dictation" dialog feeds back into cleanup. Fix tiktok → TikTok once and it
sticks forever; teach jargon, names, and your writing style over time.
Casing. Whisper sometimes hears a sentence as "Every Word Capitalized".
Echo lowercases mid-sentence words that aren't known proper nouns, so you get
normal sentence case. "Known" = casings you've taught, your Dictionary terms, a
bundled list of common brands/places/names, and I. If you would rather keep the stray
capitals than risk a surprise, set cleanup.casing.flatten_titlecase: false.
Everything lives in config.yaml. Most settings are also editable from the
dashboard Settings pages. Run RESTART.bat after editing the file directly.
| Key | Does |
|---|---|
hotkey.combo |
Push-to-talk combo (default ctrl+shift). |
whisper.model |
tiny · base · small · medium · large-v3-turbo · auto. Bigger = more accurate, slower. |
cleanup.provider |
ollama (local LLM, default) · learned (LLM-free, uses your corrections) · none (raw Whisper) · groq / anthropic (cloud, requires allow_cloud_cleanup). |
cleanup.allow_cloud_cleanup |
Opt in to cloud cleanup (Groq/Anthropic) for every dictation, which means your text leaves the machine. Off by default; falls back to local Ollama if the cloud call fails or the key is missing. Needs GROQ_API_KEY. |
cleanup.profiles |
App-aware cleanup styles (Slack vs VS Code vs Gmail). |
cleanup.casing |
flatten_titlecase, learn_from_edits, protect_common_nouns, all on by default. |
cleanup.snippets |
Your short-code to phrase expansions. |
dashboard.theme |
dark or light (also togglable in the UI). |
Off by default under the experimental: block in config.yaml. Both layers act
on a spoken prefix word (command_prefix, default "computer"). Command
Mode runs first and falls through to Action Mode on a miss.
Say "computer, select all", "computer, save", "computer, scroll down" and
Echo fires the keystroke from an allowlist instead of typing the words.
| Say… | It does |
|---|---|
| "computer, open spotify" | Launches an app from your action_apps allowlist (no shell-from-voice, ever) |
| "computer, open github.com" / "go to docs.python.org" | Opens a site (http/https/mailto only) |
| "computer, search the web for …" | Opens a web search |
| "computer, open email" | Opens your configured mail URL |
| "computer, open downloads folder" | Opens a folder from the action_folders allowlist (manage it on the dashboard Actions page) |
| "computer, summarize this pdf" | Summarizes the focused document with your local model, never a cloud call |
| "computer, create an event lunch with Sam tomorrow" | Writes a local .ics draft and never touches a calendar API |
| "computer, take a note that the build is green" | Saves a note |
| Media / volume | "play", "pause", "next", "previous", "mute", "volume up/down" via OS media keys |
Prefix-free (action_require_prefix: false): say the verb with no wake word
("open spotify"). It fires only when it resolves to a real shortcut/URL/search.
Anything else just types normally, so plain dictation is never swallowed. A
mis-heard wake word (jarvis → "Zalvis") is tolerated via fuzzy matching.
Phrasing fallback (action_intent_model, off by default): the tables above
are matched by strict patterns, so "launch spotify" or "play some music" miss
by a word. Turn this on and a missed command is retried by a small local
intent model that recovers common synonym/filler phrasings, but every guess is
re-validated through the same allowlist/URL guards, so it can never fire
anything the strict path wouldn't. Set it to shadow to log what it would
have done without acting, and use scripts/eval_intent.py to tune
action_intent_min_conf from data before trusting it. Toggle it (Off / On /
Shadow) on the dashboard Settings → Experimental page.
Two backends power that fallback (action_intent_backend): keyword
(the default, dependency-free verb-synonym rules) or model, a tiny local
embedding + logistic-regression head that generalizes to phrasings no rule
anticipated ("hush" → mute, "make a memo that…" → note). It reuses the same
on-device sentence-transformers embedder as the RAG layer (nothing leaves the
machine) and trains out-of-the-box from a shipped seed corpus; build or sharpen
it with python scripts/train_intent.py --train (see --eval / --probe).
Safety model: the allowlist and URL-scheme checks are the
sole authority on what executes. Nothing in Action Mode deletes, sends, or
pays. Every attempt is logged to the voice_actions table.
After each dictation, Echo Flow can re-clean the raw text via a stronger cloud
LLM in the background and store it as a source='teacher' row. The pattern miner
learns from both your edits and the teacher's, so the system improves toward a
reference model, not just toward you. No added latency on the live path (the
teacher runs in a daemon thread); a quality gate only persists the pair when the
teacher grades at least as well as your version.
setx GROQ_API_KEY "gsk_..." :: one-timeThen enable it under Dashboard → Settings → Vibe → Teacher model. Bootstrap from existing history without waiting for new dictations:
python scripts\backfill_teacher.py --apply --limit 500Review the pairs at http://127.0.0.1:8766/teacher before trusting the loop.
- Local by default. No telemetry, no analytics, no auto-update phone-home.
Transcripts, embeddings, and learning data live in
data/history.dbon your machine. Audio is never written to disk at all: it stays in memory for the length of the utterance and is dropped. - Five paths can call a cloud API, all gated behind a key you set yourself:
Prompt-Engineering mode (
Ctrl+Shift+Alt), the teacher loop,cleanup.allow_cloud_cleanup,cleanup.verify.escalate_cloud(the second cleanup pass), andexperimental.humanize_use_cloud. The third one is the one to watch, because unlike a deliberate keystroke it applies to every dictation. Separately,update.check_on_startup(off by default) makes one anonymous GET toapi.github.comat launch; it carries no dictation data. Check which are live on the dashboard's Privacy page, or readcleanup.providerandcleanup.allow_cloud_cleanupinconfig.yaml. - Your speech is written to a log in plain text. Every dictation appends its
raw transcript and its cleaned output to
data/wispr.log(rotating, 5 MB x 5) at INFO level. It never leaves the machine and it is covered by.gitignore, but note that the dashboard's Wipe button clears thedictationstable only, and the export zip containsconfig.yamlandhistory.dbonly: neither one touches this log. Deletedata/wispr.logby hand if you want a transcript gone from the disk entirely. - No keys are ever logged. Startup audits which cloud features are enabled and warns on a missing key, without printing the key.
- Bridge & dashboard stay loopback-only unless you deliberately change the
bind address. Read
docs/MOBILE_BRIDGE.mdbefore exposing the bridge to your LAN.
curl http://127.0.0.1:8766/api/healthzReturns daemon liveness, current phase, and which optional features are wired, without exposing keys.
app.py entry point
config.yaml the only thing you normally edit (created on first run; gitignored)
src/ the app: daemon, dashboard, voice pipeline
├── main.py daemon: hotkey, recording, transcription, dispatch
├── cleanup.py LLM/learned cleanup + casing/punctuation polish
├── transcribe.py Whisper wrapper
├── learn.py pattern + casing learning from your edits
├── hotkey.py global push-to-talk listener
├── inject.py paste/type at the cursor
└── dashboard/ Flask app, routes, templates, static assets
tests/ pytest suite (run: scripts\run_tests.bat; status: CI badge above)
scripts/ setup, backfills, helpers
docs/ architecture, dashboard, mobile, audits, action-layer specs
assets/ app icons
installer/ Windows installer + code-signing
ios/ iOS keyboard-extension port (see ios/README.md)
*.bat / *.vbs Windows launchers (run / install / restart / uninstall)
*.sh macOS and Linux launchers (setup / run / restart / dashboard / autostart / tests)
*.spec PyInstaller build specs
Where to read more: PRODUCT_OVERVIEW.md for the big
picture · CHANGELOG.md for feature history ·
docs/ROADMAP.md for what is next · docs/
for deeper specs (start at docs/README.md).
| Symptom | Fix |
|---|---|
| My fix/setting didn't take effect | Run RESTART.bat (./restart.sh on macOS). The daemon loads code & config at startup; a running process won't reflect changes until relaunched. |
| Whisper invents "thank you for watching" on silence | Already guarded (length + RMS); very short/quiet clips are dropped. |
| Recording starts when I only wanted to re-paste | The Ctrl+Shift+Win combo has a veto: add Win within a frame and recording aborts, paste fires instead. |
| Ollama "connection refused" | Start the Ollama app or run ollama serve. |
| "Couldn't reach the local model" after every reboot | Fixed in 0.3.2. Echo Flow autostarts at login but Ollama does not, so the daemon used to come up with its model backend down. It now starts Ollama itself when the binary is installed (cleanup.autostart_ollama, on by default). If it still cannot, you get rules-only cleanup and a toast saying so rather than silence. |
| Hotkey dead after a Windows update | pynput's global listener sometimes needs a restart. Run RESTART.bat. |
| macOS: hotkey does nothing, recordings are silent, or text only reaches the clipboard | A permission is missing. Run .venv/bin/python scripts/mac_permissions.py from the terminal you start Echo Flow in; it names the missing grant (Input Monitoring, Microphone or Accessibility) and --open takes you to the pane. Grant it to the app running Echo Flow, then start it again. |
| Pasting lags in some Electron apps | Clipboard restore runs in a background thread; usually fine, occasionally a ~100ms hiccup. |
| Every word comes out Capitalized | Fixed in current code; if you still see it, RESTART.bat so the running daemon picks up the casing pass. |
A custom keyboard you install via Settings. Hold to dictate, release to insert.
It talks to your desktop's local bridge over Wi-Fi. Without the bridge it
defaults to Groq cloud transcription with on-device Whisper as the fallback;
pin "on-device only" in the host app to keep audio on the phone. Building it
needs a Mac with Xcode; see ios/README.md.
MIT. See LICENSE.
Nothing if you run fully local. Groq is free at single-human speaking volumes. Anthropic/OpenAI cost real money per API call, so only use them if you want their cleanup quality and don't mind the bill.




