diff --git a/README.md b/README.md index 4403e5b..fcd1864 100644 --- a/README.md +++ b/README.md @@ -146,6 +146,7 @@ The full command surface. The top level has eight groups — `auth`, `workspaces | `databases show` | Show details for an instant database | | `databases create` | Create a new instant database | | `databases fork` | Fork a database into a new, independent database | +| `databases lineage` | Show a database's fork ancestry and direct forks | | `databases attach` | Attach a catalog so its tables are queryable | | `databases detach` | Detach a previously attached catalog | | `databases use` | Set the current (default) database | diff --git a/skills/hotdata/SKILL.md b/skills/hotdata/SKILL.md index fa93683..1b3df33 100644 --- a/skills/hotdata/SKILL.md +++ b/skills/hotdata/SKILL.md @@ -98,6 +98,7 @@ hotdata databases list [--workspace-id ] [--output table|json|yaml hotdata databases count [--workspace-id ] [--output table|json|yaml] hotdata databases create [--name ] [--catalog ] [--table ...] [--schema public] [--expires-at ] [--workspace-id ] [--output table|json|yaml] hotdata databases fork [] [--name ] [--expires-at ] [--workspace-id ] [--output table|json|yaml] +hotdata databases lineage [] [--forks-limit ] [--workspace-id ] [--output table|json|yaml] hotdata databases use hotdata databases unset hotdata databases [--workspace-id ] [--output table|json|yaml] @@ -119,10 +120,11 @@ hotdata databases tables remove
[--database ] [--schema public] [--w - `list` — all instant databases in the workspace. Active database is marked with `*` under the DEFAULT column; CREATED shows when each database was made. - `count` — the total number of instant databases in the workspace, across **all** pages (`list` shows one page). Prints a bare integer by default so it drops straight into scripts (`$(hotdata databases count)`); `--output json|yaml` render `{"count": N}` / `count: N`. - `create` — creates a new instant database. `--name` is an optional human-readable display name. `--catalog` sets the SQL alias used in queries (`SELECT … FROM .schema.table`); must be `[a-z_][a-z0-9_]*`. `--expires-at` accepts relative durations (`24h`, `7d`, `90m`) or an RFC 3339 timestamp; omitting means no expiry. Repeat `--table` to declare tables up front. -- `fork` — creates a new instant database that is an independent deep copy of an existing one (same schemas, tables, and data); the source is left unchanged and the two diverge freely afterwards. The source defaults to the active database; pass the database `` to fork another. `--name` defaults to `-fork` (so the two stay distinguishable in `list`); `--expires-at` accepts a relative duration or RFC 3339 timestamp, and when omitted a still-future source expiry is carried over. The fork becomes the active database on success. The fork answers to the **same catalog alias** as its source inside its own scope; catalogs attached to the source are **re-attached** to the fork, but indexes are **not** carried over. Only databases created with the current (DuckLake) storage engine can be forked — older parquet-backed databases return an error. +- `fork` — creates a new instant database that is an independent deep copy of an existing one (same schemas, tables, and data); the source is left unchanged and the two diverge freely afterwards. The source defaults to the active database; pass the database `` to fork another. `--name` defaults to `-fork` (so the two stay distinguishable in `list`); `--expires-at` accepts a relative duration or RFC 3339 timestamp, and when omitted a still-future source expiry is carried over. The fork becomes the active database on success. The fork answers to the **same catalog alias** as its source inside its own scope; catalogs attached to the source are **re-attached** to the fork, but indexes are **not** carried over. Only databases created with the current (DuckLake) storage engine can be forked — older parquet-backed databases return an error. The fork's output (and its `databases ` view) records **`forked_from`** provenance: source id, the source's name at fork time, the copied snapshot, and when. +- `lineage` — renders a database's fork family tree (defaults to the active database): the ancestor chain down from the root, the queried database marked `← this database`, and its **direct** forks one level below (a fork of a fork appears in its own parent's lineage). Lineage is a historical record, not a live link — the databases stay independent, and a **deleted** generation stays in the chain (marked `deleted`). Forks made before the server recorded lineage carry none. `--forks-limit ` pages the direct-fork list (server clamps to 1–100); a truncated list closes with `⋯ N more`, and `-o json`/`yaml` always report the true `fork_count`. - `use` — saves the database **id** as the active database. Subsequent `databases tables` and `databases context` commands use it automatically. Note that a successful `fork` also updates this: the fork becomes the active database. - `unset` — clears the active database from config. -- `` — inspect one database (returns id, catalog, name, expires_at). +- `` — inspect one database (returns id, catalog, name, expires_at; a fork also shows its `forked_from` record). - `remove` — removes the instant database; clears the active-database config if it matched. - `load` (top-level shorthand) — loads parquet into `--catalog.--schema.--table`. Accepts `--file`, `--url`, `--upload-id`, or `--result-id` (load a saved query result by id — from `hotdata databases results` or a query's `[result-id: …]` footer — instead of a file; the result must belong to the target database). Replaces the table by default; pass `--append` to add rows to the existing table instead. If the table was not declared at create time, the CLI automatically deletes and recreates the database with the table declared, then retries the load. - `tables list` — lists tables with `TABLE` (`..
`), `SYNCED`, `LAST_SYNC`. Uses active database when `--database` is omitted. @@ -207,14 +209,14 @@ hotdata databases context push [--database ] [--dry-run] ### Execute SQL Query ``` -hotdata query "" [--workspace-id ] [--database ] [--output table|json|csv] +hotdata query "" [--workspace-id ] [--database ] [--dialect hotsql|duckdb|postgres|snowflake] [--output table|json|csv] hotdata query status ``` - Default output is `table` (row count and execution time). - **A query runs inside one instant database** (active database or `--database`); with none set it fails *"a database is required."* The scope sees the database's own catalog **plus any attached catalogs only**. To query an attached catalog's tables or join across catalogs, attach the catalog first — see [Querying across catalogs (attach)](#querying-across-catalogs-attach). - Use `hotdata databases tables list` and `hotdata databases tables show` for discovery — not `information_schema` via `query`. (Discovery lists every workspace table; queryability still requires the table's catalog to be in the active database's scope.) -- **PostgreSQL dialect.** Quote non-lowercase columns with double quotes. +- **PostgreSQL dialect.** Quote non-lowercase columns with double quotes. To write DuckDB/Postgres/Snowflake SQL instead, pass `--dialect` (server-side transpile, read-only queries) — details in **`hotdata-analytics`**. - Async runs return `query_run_id` → poll with `query status ` (do not re-run the same heavy SQL). `query status` exit codes: `0` succeeded, `1` failed, `2` still running (poll again), `3` succeeded but the result is a truncated/incomplete preview. - **Large results are complete, not a preview.** The server returns inline rows only up to a bounded cap and persists the full set out-of-band; `hotdata query` transparently fetches the full result, so the printed rows and row count are the complete set. (If the full result can't be retrieved, the CLI prints the preview and a `warning:` to stderr.) - **Backpressure is handled.** Under heavy concurrent load the server may shed a query with HTTP 429 (`OVERLOADED`); the CLI auto-retries (honoring `Retry-After`) before surfacing an error — no manual retry needed. @@ -227,7 +229,7 @@ hotdata jobs list [--workspace-id ] [--job-type ] [--status hotdata jobs [--workspace-id ] [--output table|json|yaml] ``` - `list` shows only active jobs (`pending`, `running`) by default. Use `--all` to see all jobs. -- `--job-type`: `data_refresh_table`, `data_refresh_connection`, `create_index`, `managed_load`. +- `--job-type`: `create_index`, `managed_load`. (`data_refresh_table` / `data_refresh_connection` were retired server-side — the engine is push-only; nothing refreshes.) - `--status`: `pending`, `running`, `succeeded`, `partially_succeeded`, `failed`. - Use `hotdata jobs ` to inspect a specific job's status, error, and result. diff --git a/skills/hotdata/references/WORKFLOWS.md b/skills/hotdata/references/WORKFLOWS.md index dece853..86003a7 100644 --- a/skills/hotdata/references/WORKFLOWS.md +++ b/skills/hotdata/references/WORKFLOWS.md @@ -140,7 +140,7 @@ hotdata databases fork --expires-at 24h # deep copy; becomes the active databa hotdata databases load --catalog sales --table orders --file ./risky.parquet # hits the fork ``` -**Capture both ids.** After the fork, both databases answer to the same catalog alias (here `sales`), so ids are the only unambiguous way to refer to either one — the source id comes from `databases list` up front, the fork id from the `fork` output. The shared alias means experimental SQL runs unchanged against the fork. Attached catalogs are re-attached to the fork; indexes are not carried over. When done, keep the fork (`databases use ` to switch back to the source) or delete it (`databases remove `). Only DuckLake-backed databases can be forked — see `fork` in the main skill for details. +**Capture both ids.** After the fork, both databases answer to the same catalog alias (here `sales`), so ids are the only unambiguous way to refer to either one — the source id comes from `databases list` up front, the fork id from the `fork` output. The shared alias means experimental SQL runs unchanged against the fork. Attached catalogs are re-attached to the fork; indexes are not carried over. When done, keep the fork (`databases use ` to switch back to the source) or delete it (`databases remove `). If the ids get lost, `hotdata databases lineage` recovers the pair: it renders the fork tree (ancestors and direct forks), and `databases ` shows a fork's `forked_from` record. Only DuckLake-backed databases can be forked — see `fork` in the main skill for details. ---