A web application that answers natural-language questions (Hebrew and English) by discovering and extracting datasets from Israel's open data portal — data.gov.il.
- You type a question in the search box (e.g. "כמות הגשמים שירדה באשקלון אתמול" or "gas stations in Jerusalem").
- An OpenAI-powered agent interprets the question, searches data.gov.il's CKAN API for relevant datasets, selects the best resource, and extracts the data.
- The app presents a direct answer, an evidence table of extracted records, full source provenance, and any warnings or assumptions.
The LLM agent has access to 6 server-side functions via OpenAI function calling:
| Tool | Purpose |
|---|---|
search_datasets |
Search data.gov.il for datasets by keyword |
get_dataset |
Get full metadata for a dataset |
list_resources |
List all resources in a dataset |
preview_resource |
Preview fields and sample rows from a DataStore resource |
fetch_resource |
Fetch DataStore data with filters, search, sort, pagination |
download_and_sample |
Download and parse CSV/JSON files (allowlisted domains only) |
app/
main.py # FastAPI routes, rate limiting, Jinja2 templates
models.py # Pydantic models
utils.py # Retry/backoff, caching, domain allowlist, sanitization
smoke_test.py # CLI smoke test
templates/
index.html # Single-page RTL-aware UI
services/
openai_agent.py # LLM agent loop with tool calling
ckan_client.py # CKAN Action API wrappers
resource_resolver.py # Resource ranking and selection
filters.py # Date normalization, local filtering
answer_formatter.py # Structured response builder
fetchers/
datastore_fetcher.py # DataStore-backed resource fetcher
csv_fetcher.py # CSV download and parse
json_fetcher.py # JSON download and parse
tests/ # 31 unit tests
Dockerfile
render.yaml # Render deployment config
requirements.txt
- Python 3.11+
- An OpenAI API key
# Clone the repo
git clone https://github.com/YoniLabell/DataGovAI.git
cd DataGovAI
# Install dependencies
pip install -r requirements.txt
# Set your OpenAI API key
export OPENAI_API_KEY="sk-..."
# Run the app
uvicorn app.main:app --reload --port 8000Open http://localhost:8000 in your browser.
| Variable | Required | Default | Description |
|---|---|---|---|
OPENAI_API_KEY |
Yes | — | OpenAI API key |
OPENAI_MODEL |
No | gpt-4.1-mini |
OpenAI model to use |
PORT |
No | 10000 |
Server port |
LOG_LEVEL |
No | INFO |
Logging level |
RATE_LIMIT_PER_MINUTE |
No | 15 |
Max requests per IP per minute |
pytest tests/ -vAll 31 tests use mocked HTTP responses and require no API keys.
Run a question end-to-end from the command line (requires OPENAI_API_KEY):
python -m app.smoke_test "כמות הגשמים שירדה באשקלון אתמול"
python -m app.smoke_test "gas stations in Jerusalem"- Push this repo to GitHub.
- Go to Render Dashboard and click New > Blueprint.
- Connect the repo — Render will detect
render.yamlautomatically. - Set the
OPENAI_API_KEYenvironment variable in the Render dashboard. - Deploy.
- Create a new Web Service on Render.
- Connect your GitHub repo.
- Set Runtime to Docker.
- Add environment variable
OPENAI_API_KEY. - Set Health Check Path to
/health. - Deploy.
The app runs on Render's free tier with no heavy dependencies.
- Domain allowlist: Only fetches data from
data.gov.iland its subdomains. Arbitrary URL fetching is blocked. - Input sanitization: User input and CKAN data are sanitized before processing.
- Rate limiting: Per-IP rate limiting (15 requests/minute by default).
- Tool enforcement: The LLM proposes actions; the server validates and executes only allowed tool calls.
- Backend: Python, FastAPI, httpx, Pydantic
- LLM: OpenAI API with function calling
- Data source: data.gov.il CKAN Action API
- Frontend: Server-rendered HTML with vanilla JavaScript
- Deployment: Docker, Render