Skip to content
 
 

Latest commit

 

History

69 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Flow

Voice interface for Claude Code. Talk to Claude instead of typing — Flow handles speech-to-text, intent parsing, and text-to-speech so you can code hands-free.

How it works

Flow runs as a terminal UI (built with Textual) that sits alongside Claude Code. It operates in two modes:

  • SDK mode — launches Claude Code directly via the Claude Agent SDK, managing the full conversation lifecycle
  • Bridge mode — attaches to an already-running Claude Code session through a WebSocket bridge (flow-bridge), letting you add voice control to any existing terminal session

Architecture

┌─────────────┐     ┌──────────────────────────────┐
│  Microphone  │────▶│  STT Provider (Sarvam/etc.)  │
└─────────────┘     └──────────────┬───────────────┘
                                   │ transcript
                    ┌──────────────▼───────────────┐
                    │    Voice Intent Parser        │
                    │  (yes/no, option selection)   │
                    └──────────────┬───────────────┘
                                   │
                    ┌──────────────▼───────────────┐
                    │    State Machine              │
                    │  IDLE → LISTENING → SENDING   │
                    │  → PROCESSING → SPEAKING      │
                    └──────────────┬───────────────┘
                                   │
              ┌────────────────────┴────────────────────┐
              │                                         │
   ┌──────────▼──────────┐               ┌─────────────▼──────────┐
   │  Claude Agent SDK   │               │  WebSocket Bridge      │
   │  (SDK mode)         │               │  (Bridge mode)         │
   └──────────┬──────────┘               └─────────────┬──────────┘
              │                                         │
              └────────────────────┬────────────────────┘
                                   │ response
                    ┌──────────────▼───────────────┐
                    │   TTS Provider (Sarvam/etc.) │────▶ Speaker
                    └──────────────────────────────┘

Installation

Requires Python 3.10+ and uv.

# Clone and install
git clone https://github.com/satish-rathod/flow.git
cd flow
uv tool install .

This installs two commands:

  • flow — the voice interface TUI
  • flow-bridge — WebSocket bridge for attaching to running Claude Code sessions

Configuration

Create a .env file in the project root, or at ~/.config/flow/.env for global installs:

# Required — Sarvam AI API key for STT/TTS
SARVAM_API_KEY=your_key_here

# Optional
STT_PROVIDER=sarvam          # sarvam (default)
TTS_PROVIDER=sarvam          # sarvam (default)
CLAUDE_MODEL=claude-sonnet-4-20250514  # override Claude model
CLAUDE_PERMISSION_MODE=default         # Claude permission mode

Get a Sarvam API key at sarvam.ai.

Usage

SDK mode (direct)

# Start with session picker
flow

# Start a new session
flow --new

# Resume a specific session
flow --resume <session-id>

# Continue the most recent session
flow --continue

Bridge mode (attach to existing Claude Code)

# In one terminal, start the bridge wrapping Claude
flow-bridge claude

# In another terminal, attach Flow to the bridge
flow --attach

You can also attach to a specific WebSocket URL:

flow --attach ws://localhost:9712

Voice commands

Once running, just speak naturally. Flow understands:

  • Direct instructions — "create a hello world Python script"
  • Yes/no answers — "yes", "go ahead", "no", "deny" (for Claude's permission prompts)
  • Option selection — "first", "second", "option three" (for numbered choices)

CLI options

flow [options]

Options:
  --language CODE    Language for STT/TTS (e.g. en-IN, hi-IN)
  --speaker NAME     TTS voice (e.g. shubh, meera)
  --new              Start a new session directly
  --resume ID        Resume a specific session
  --continue         Continue the most recent session
  --attach [URL]     Attach to a running bridge (auto-discover or specify URL)

Development

# Install dev dependencies
uv sync --extra dev

# Run tests
uv run pytest

# Run Flow from source
uv run flow

Project structure

flow/
├── main.py              # CLI entry point
├── config.py            # Configuration and .env loading
├── state_machine.py     # Conversation state machine
├── voice_intent.py      # Yes/no and option selection parser
├── claude_client.py     # Claude Agent SDK integration
├── bridge.py            # WebSocket bridge (pty + ws server)
├── ws_client.py         # WebSocket client for bridge mode
├── session_summary.py   # Post-completion session summaries
├── audio.py             # Audio device management
├── stt.py               # STT event types
├── tts.py               # TTS abstraction
├── discovery.py         # Bridge auto-discovery
├── summarizer.py        # Regex-based output summarizer
├── providers/
│   ├── base.py          # Provider interface
│   └── sarvam.py        # Sarvam AI STT/TTS provider
└── ui/
    ├── app.py           # Main Textual application
    ├── styles.tcss      # TUI stylesheet
    ├── widgets/         # UI components (voice indicator, input bar, etc.)
    └── screens/         # Session picker, permission/question dialogs

License

MIT

About

Voice interface for Claude Code

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages