Skip to content

ppXD/Wisp

Repository files navigation

Wisp

Local-first, real-time meeting transcription — on-device, private, GPU-accelerated.

Your microphone and the meeting's audio become a live, per-speaker transcript with an AI copilot beside it — entirely on your machine. No cloud. No upload. No account.


CI Release Platforms Built with Rust Tauri v2 License: MIT

English · 简体中文


Download for macOS   Download for Windows


Wisp — live transcription with per-speaker labels and the realtime AI Assist copilot

Wisp turns any conversation into a clean, timestamped, speaker-labelled transcript in real time — and runs a live AI assistant alongside it. Every byte of audio and every model stays on your device by default. It installs in seconds, needs zero extra setup (no virtual audio drivers, no kernel extensions), and is tuned to the metal on each platform — Metal + the Neural Engine on Apple Silicon, native loopback on Windows.

🔒 Private by design · ⚡ Optimized to the architecture · 🎛️ Yours to configure · 🪶 Install-and-go

✨ Highlights

🎙️ Live + AI copilot Sub-second streaming transcript with per-speaker labels, plus an AI assistant that summarizes, extracts action items, and coaches in real time
🔒 100% on-device Audio and models never leave your machine. Cloud engines are strictly opt-in, with your own keys
🍎 Apple-Silicon-native Metal GPU inference, Apple Neural Engine via Core ML, Unified Memory zero-copy, and on-device Apple SpeechAnalyzer
🔌 Zero dependencies One-click system-audio capture — no BlackHole, no kernel extensions, no virtual devices
🎛️ Truly customizable Choose, configure, and delete models freely. Language, accuracy/speed, decoding params — all in your hands
📄 SOTA file transcription Accurate batch transcription with diarization, word-level timestamps, custom vocabulary, and structured export
🧠 Private searchable memory Every meeting auto-saved to a local library you can search by meaning, not just keywords — on-device embeddings (Qwen3, BGE-M3, multilingual E5) with hybrid keyword + semantic ranking
🪶 Tiny footprint An 8–22 MB installer; multi-GB models stream in only when you ask for them

🎙️ The live meeting copilot

This is the heart of Wisp — and it's fast.

  • Streaming transcript, sub-second. Words appear as they're spoken (live partials → finalized lines), each timestamped. No "press stop to see the result."
  • Knows who's talking. On-device diarization labels every line live — You vs Them, or Speaker 1 / 2 / 3 — with running speaker centroids that stay stable across a long call.
  • Captures both sides. Your microphone and the meeting's system audio are fused onto a single timeline, with WebRTC AEC3 echo cancellation so the remote voices don't bleed back through your mic. One click — no loopback driver to install.
  • 🤖 AI Assist — your second brain in the call. A live copilot panel that streams as it thinks:
    • Rolling summaries, action items, decisions, and open questions that update as the meeting unfolds
    • Follow-up email drafts, ready to send
    • Real-time hints and service-industry templates — sales coaching, support guidance, live sentiment/tone monitoring
    • Diarization-aware context so it knows who said what, with controlled cadence so it's helpful, not noisy

Bring your own LLM endpoint (local or cloud) — the assistant is model-agnostic and fully parameterized (temperature, penalties, max tokens, and more).

📄 State-of-the-art file transcription

Drop in any audio or video file and get a transcript you can trust:

  • Accuracy-first by default — Whisper large-v3-turbo, with heavier or quantized variants a click away.
  • Speaker diarizationwho said what, with word-level timestamps and per-word speaker attribution.
  • Custom vocabulary / term biasing — feed in names, products, and jargon so they transcribe correctly.
  • Cleaner audio in — neural denoising and VAD gating drop non-speech before the model sees it.
  • Optional local LLM cleanup — tidy punctuation and disfluencies without leaving the device.
  • Structured Markdown export — summary, speakers, and a timestamped timeline, ready to share.
  • Live progress even on opaque decode phases, so you're never staring at a frozen bar.

Wisp — file transcription with the AI Notes sidebar open beside the transcript

🧠 A private, searchable memory

Every meeting Wisp transcribes is saved to a local notes library — and you can search across all of them by meaning, fully on-device.

  • Semantic + keyword, fused. Search "the budget discussion" and Wisp finds the right note even if those exact words were never spoken — SQLite FTS5 keyword matching fused with vector similarity (hybrid). Every hit shows why it matched: keyword highlights and a semantic confidence score.

  • SOTA local embeddings — your pick. Download, switch, and delete on-device embedding models right in Settings:

    • Qwen3-Embedding 0.6B — a state-of-the-art instruction-tuned decoder embedder
    • BGE-M3 — 1024-dim, 100+ languages, strong CJK
    • plus GTE multilingual, Multilingual E5 (small / base / large), and BGE-Chinese

    Switching models re-indexes your library atomically — an interrupted or failed switch never corrupts the existing index.

  • Nothing leaves the device. Notes are embedded and searched locally through ONNX Runtime. Prefer a hosted embedder for a job? OpenAI-compatible cloud models are opt-in, with your own key.

🔒 Local-first & genuinely private

Transcription, diarization, denoising, and VAD all run on-device through sherpa-onnx and Metal/ANE — no audio ever touches a network. Models are pulled from Hugging Face only when you choose to install them, then cached locally.

Need a hosted model for a specific job? Cloud engines (OpenAI, Gemini, Groq, Qwen, Speechmatics) are available as opt-in — you add your own key, stored locally, and Wisp stays local everywhere else.

🍎 Tuned for Apple Silicon

Wisp doesn't just "run on a Mac" — it's optimized to the architecture:

  • Metal GPU inference — the Whisper engine (whisper.cpp) executes on the GPU via Metal for large speedups over CPU.
  • Apple Neural Engine — the Whisper encoder runs through Core ML so the heaviest stage lands on the ANE, freeing the GPU and CPU.
  • Unified Memory Architecture — Apple Silicon's shared memory means zero-copy handoff between CPU, GPU, and ANE: no PCIe round-trips, lower latency, less power.
  • Apple SpeechAnalyzer — on macOS 26, Wisp can use Apple's built-in on-device speech framework (ANE-accelerated, zero model download) as a first-class engine.
  • ScreenCaptureKit — native, permissioned system-audio capture with no kernel extension.

The engine, model, and decoding cadence are auto-selected to your machine (cores, RAM, GPU/Neural-Engine tier), so it's fast out of the box and tunable when you want control.

🎛️ Yours to configure — not the other way around

No mandatory engine downloads. No forced multi-gigabyte "summary engine" gating you before you can start. Wisp runs immediately, and you decide what to add:

  • Pick any model from the catalog — speed-first or accuracy-first — and switch per mode (Live vs File).
  • Delete models you don't need, right from the picker, to reclaim disk in one click.
  • Configure everything — transcription language, accuracy/speed profile, VAD gating, denoising, decoding thresholds, diarization, custom vocabulary — instead of being locked to one preset.
  • Honest picker — models your machine can't run are clearly marked, with size and hardware hints before you download.

🔌 Zero setup, no dependencies

  • macOS: system audio via ScreenCaptureKitno BlackHole, no Loopback, no virtual audio device.
  • Windows: system audio via native WASAPI loopback.
  • Microphone + meeting audio captured together, echo-cancelled, with one click. Install the app and start — that's the whole setup.

🪶 And more

  • Dictation — global push-to-talk speech-to-text injected into any app via native text insertion.
  • Tiny installers — 8 MB (Windows) / 22 MB (macOS); the ML runtime is the floor, models stream on demand.
  • Multilingual UI — English, 简体中文, 繁體中文.
  • Cross-platform — macOS (Apple Silicon) and Windows (x64), one codebase.

📦 Install

Grab the latest build from Releases:

macOS (Apple Silicon)

  1. Download Wisp_<version>_aarch64.dmg, open it, and drag Wisp to Applications.
  2. The build is currently unsigned, so clear the quarantine flag once:
    xattr -cr /Applications/Wisp.app
  3. Launch Wisp. Grant Microphone and Screen Recording (for system audio) when prompted.

Apple Silicon only (M1 or newer). Intel Macs are not supported — the GPU/Neural-Engine and system-audio paths require the modern Apple-Silicon SDKs.

Windows (x64)

  1. Download Wisp_<version>_x64-setup.exe and run it.
  2. Launch Wisp and grant microphone access when prompted.

🛠️ Build from source

Requires Rust (stable), Node 20+, and platform build tools (Xcode + meson/ninja on macOS; MSVC on Windows).

git clone --recurse-submodules https://github.com/ppXD/Wisp.git
cd Wisp/app
npm install
npm run tauri dev      # run the app
npm run tauri build    # produce an installer

💻 Platform support

Platform Status Transcription GPU System audio Echo cancel
macOS (Apple Silicon) ✅ Released sherpa-onnx · whisper.cpp · Apple SpeechAnalyzer Metal + ANE ScreenCaptureKit WebRTC AEC3
Windows (x64) ✅ Released sherpa-onnx DirectML (auto) · CPU fallback WASAPI loopback cross-stream dedup

📄 License

MIT © Wisp contributors.

Built with Rust · Tauri · sherpa-onnx · whisper.cpp

About

Local, real-time, cross-platform meeting transcription — on-device, private, model-swappable.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages