CLI toolkit for the NVIDIA NIM free tier. Query models, run chat completions, generate embeddings, and manage your API key from the terminal. Zero dependencies beyond Python 3.7+.
- List models: See which NIM models are available to your API key
- Chat completions: Send messages and get responses from LLMs
- Text completions: Use the legacy completions endpoint
- Embeddings: Generate vector embeddings for text
- Health check: Verify your API key and connectivity
- Config management: Store, show, and remove your API key locally
- Model aliases: 10 built-in shortcuts for popular free-tier models
- Python 3.7 or newer
- An NVIDIA NIM API key (free at https://build.nvidia.com/)
No pip packages needed. The tool uses only the Python standard library.
git clone https://github.com/cappy-dev/nvidia-nim-tools.git
cd nvidia-nim-tools
chmod +x nvt.pyOptional: symlink it into your PATH:
ln -s $(pwd)/nvt.py ~/.local/bin/nvtGet a free API key from https://build.nvidia.com/ and store it:
python3 nvt.py config --set-key nvapi-xxxxxOr export it as an environment variable:
export NIM_API_KEY=nvapi-xxxxxOr pass it on every command:
python3 nvt.py --api-key nvapi-xxxxx modelspython3 nvt.py modelspython3 nvt.py knownThis prints 10 shortcut names mapped to their full model IDs so you do not need to type long paths.
python3 nvt.py chat llama3-70b-instruct --prompt "Explain quantum entanglement in one paragraph"With a custom system prompt and temperature:
python3 nvt.py chat llama3-70b-instruct \
--prompt "Write a haiku about Raspberry Pi" \
--system "You are a poet." \
--temperature 0.9Read the prompt from a file:
python3 nvt.py chat llama3-70b-instruct --file prompt.txtpython3 nvt.py complete mixtral-8x22b-instruct \
--prompt "The capital of France is"python3 nvt.py embed --text "Hello world"Save the full response to a file:
python3 nvt.py embed --text "Hello world" --output embedding.jsonUse a different embedding model:
python3 nvt.py embed arctic-embed-l --text "search query"python3 nvt.py healthPrints the number of visible models and latency. Good for confirming your key works.
python3 nvt.py config --set-key nvapi-xxxxx # store key
python3 nvt.py config --show # show masked key
python3 nvt.py config --unset # remove keyllama-3.1-nemotron-70b-instruct-> meta/llama-3.1-nemotron-70b-instructllama3-70b-instruct-> meta/llama3-70b-instructllama3-8b-instruct-> meta/llama3-8b-instructmixtral-8x22b-instruct-> mistralai/mixtral-8x22b-instruct-v0.1mistral-large-> mistralai/mistral-large-2-instructgemma-2-27b-> google/gemma-2-27b-itphi-3.5-mini-> microsoft/phi-3.5-mini-instructstarcoder2-15b-> bigcode/starcoder2-15b-instructarctic-embed-l-> nvidia/arctic-embed-lnvidia-llama3-chatqa-1.0-70b-> nvidia/nvidia-llama3-chatqa-1.0-70b
You can always pass the full model ID directly. The aliases are shortcuts for convenience.
NVIDIA NIM free tier has per-model rate limits (typically 40 requests per minute and 10,000 tokens per minute). The CLI will print an error if you hit a 429 response. Wait a minute and retry.
Your API key is stored locally at ~/.config/nvidia-nim-tools/config.json. The tool never sends your key anywhere except integrate.api.nvidia.com. No telemetry, no analytics, no phone-home.
MIT