Voice interface for Claude Code. Talk to Claude instead of typing — Flow handles speech-to-text, intent parsing, and text-to-speech so you can code hands-free.
Flow runs as a terminal UI (built with Textual) that sits alongside Claude Code. It operates in two modes:
- SDK mode — launches Claude Code directly via the Claude Agent SDK, managing the full conversation lifecycle
- Bridge mode — attaches to an already-running Claude Code session through a WebSocket bridge (
flow-bridge), letting you add voice control to any existing terminal session
┌─────────────┐ ┌──────────────────────────────┐
│ Microphone │────▶│ STT Provider (Sarvam/etc.) │
└─────────────┘ └──────────────┬───────────────┘
│ transcript
┌──────────────▼───────────────┐
│ Voice Intent Parser │
│ (yes/no, option selection) │
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ State Machine │
│ IDLE → LISTENING → SENDING │
│ → PROCESSING → SPEAKING │
└──────────────┬───────────────┘
│
┌────────────────────┴────────────────────┐
│ │
┌──────────▼──────────┐ ┌─────────────▼──────────┐
│ Claude Agent SDK │ │ WebSocket Bridge │
│ (SDK mode) │ │ (Bridge mode) │
└──────────┬──────────┘ └─────────────┬──────────┘
│ │
└────────────────────┬────────────────────┘
│ response
┌──────────────▼───────────────┐
│ TTS Provider (Sarvam/etc.) │────▶ Speaker
└──────────────────────────────┘
Requires Python 3.10+ and uv.
# Clone and install
git clone https://github.com/satish-rathod/flow.git
cd flow
uv tool install .This installs two commands:
flow— the voice interface TUIflow-bridge— WebSocket bridge for attaching to running Claude Code sessions
Create a .env file in the project root, or at ~/.config/flow/.env for global installs:
# Required — Sarvam AI API key for STT/TTS
SARVAM_API_KEY=your_key_here
# Optional
STT_PROVIDER=sarvam # sarvam (default)
TTS_PROVIDER=sarvam # sarvam (default)
CLAUDE_MODEL=claude-sonnet-4-20250514 # override Claude model
CLAUDE_PERMISSION_MODE=default # Claude permission modeGet a Sarvam API key at sarvam.ai.
# Start with session picker
flow
# Start a new session
flow --new
# Resume a specific session
flow --resume <session-id>
# Continue the most recent session
flow --continue# In one terminal, start the bridge wrapping Claude
flow-bridge claude
# In another terminal, attach Flow to the bridge
flow --attachYou can also attach to a specific WebSocket URL:
flow --attach ws://localhost:9712Once running, just speak naturally. Flow understands:
- Direct instructions — "create a hello world Python script"
- Yes/no answers — "yes", "go ahead", "no", "deny" (for Claude's permission prompts)
- Option selection — "first", "second", "option three" (for numbered choices)
flow [options]
Options:
--language CODE Language for STT/TTS (e.g. en-IN, hi-IN)
--speaker NAME TTS voice (e.g. shubh, meera)
--new Start a new session directly
--resume ID Resume a specific session
--continue Continue the most recent session
--attach [URL] Attach to a running bridge (auto-discover or specify URL)
# Install dev dependencies
uv sync --extra dev
# Run tests
uv run pytest
# Run Flow from source
uv run flowflow/
├── main.py # CLI entry point
├── config.py # Configuration and .env loading
├── state_machine.py # Conversation state machine
├── voice_intent.py # Yes/no and option selection parser
├── claude_client.py # Claude Agent SDK integration
├── bridge.py # WebSocket bridge (pty + ws server)
├── ws_client.py # WebSocket client for bridge mode
├── session_summary.py # Post-completion session summaries
├── audio.py # Audio device management
├── stt.py # STT event types
├── tts.py # TTS abstraction
├── discovery.py # Bridge auto-discovery
├── summarizer.py # Regex-based output summarizer
├── providers/
│ ├── base.py # Provider interface
│ └── sarvam.py # Sarvam AI STT/TTS provider
└── ui/
├── app.py # Main Textual application
├── styles.tcss # TUI stylesheet
├── widgets/ # UI components (voice indicator, input bar, etc.)
└── screens/ # Session picker, permission/question dialogs
MIT