Skip to content

Export bounded grouped SQL inserts for local DB dumps - #2893

Merged
graphite-app[bot] merged 1 commit into
mainfrom
arda/rai-2649-bounded-grouped-sql-dumps
Sep 30, 2026
Merged

graphite-app[bot] merged 1 commit into
mainfrom
arda/rai-2649-bounded-grouped-sql-dumps

Conversation

@findolor

@findolor findolor commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Chained PRs

Motivation

Filtered local DB dumps still contain one INSERT per row. That creates a large number of Rust SQL statement objects during browser bootstrap and adds avoidable parser work.

Part of RAI-2649.

Solution

  • Group consecutive rows from the same table into multi-row INSERT statements, with at most 256 rows or 256 KiB per statement. A single row above the byte limit remains a standalone statement so the exporter preserves it.
  • Keep each statement on one line inside the existing single BEGIN/COMMIT dump transaction, so current line-based dump importers can read it.
  • Add tests for row and byte limits, SQL escaping, and round-trip imports into SQLite.

This PR does not change the manifest schema, browser importer, or sqlite-web. The bounded atomic browser import API is tracked in sqlite-web#35; wiring it into raindex remains a separate follow-up.

Checks

  • nix develop .#rust-shell --offline -c cargo fmt --all -- --check — passed.
  • nix develop .#rust-shell --offline -c cargo clippy --workspace --all-targets -- -D warnings — passed.
  • nix develop .#rust-shell --offline -c cargo test --workspace -- --test-threads=2 — passed. The default parallel run timed out in unrelated local Anvil fixture tests; limiting concurrency resolved it.
  • CI-equivalent nix develop .#wasm-shell --offline Wasm test command — passed; this host reported no runnable Wasm tests.
  • Three read-only simplification passes, two read-only local Codex reviews, and a CodeRabbit review — no actionable findings.

Review focus: confirm multi-row VALUES remains compatible with the current one-statement-per-line dump importer. Unrelated unstaged benchmark experiments in the workspace are excluded from this PR.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⚙️ Run configuration

Configuration used: Repository: rainlanguage/raindex/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 6c740bcb-9cea-4e2f-8e32-02ecc53e2796

📥 Commits

Reviewing files that changed from the base of the PR and between a018a4b and 984785c.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: rainlanguage/raindex/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 8d4e860f-028e-4526-8109-be222485704b

📥 Commits

Reviewing files that changed from the base of the PR and between 66adc15 and a018a4b.

📒 Files selected for processing (1)
  • crates/common/src/local_db/export.rs

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

SQL export now groups row tuples into multi-row INSERT statements. Each statement is limited to 256 rows and 256 KiB, except that a single oversized row is emitted on its own. Tests check limit handling and row preservation.

Changes

SQL Export

Layer / File(s) Summary
Bounded insert builder
crates/common/src/local_db/export.rs
The builder groups row tuples into multi-row INSERT statements and flushes before exceeding the row or byte limit. An oversized row is emitted alone. Tests verify splitting, SQL execution, value order, and the default row limit.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Feature

Merge Risk: ⚪ Minimal · up to a018a

No identified issue prevents merging after the normal checks and stack requirements are satisfied.

Security Architecture Review

Security architecture risk: 🔵 Low · up to a018a

The grouped inserts remain one statement per line within the existing dump transaction. No new exposed operation or access-control change is evident, though interruption and concurrency behavior is not fully tested.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The directly evidenced exposure is the generated local database dump and its existing importer, not a new service entrypoint or privilege path.

Trust Boundaries and Controls

  • observed — Rows still pass through the existing SQL-value formatter before grouping; the new builder joins the resulting tuples rather than changing that conversion path.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 1 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: bounded grouping of SQL inserts for local database dumps.
✨ Finishing Touches 💡 2
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@findolor findolor changed the title Export bounded grouped SQL inserts Export bounded grouped SQL inserts for local DB dumps Sep 28, 2026
@linear

linear Bot commented Sep 28, 2026

Copy link
Copy Markdown

RAI-2649

findolor commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator Author

How to use the Graphite Merge Queue

Add the label Raindex-queue to this PR to add it to the merge queue.

You must have a Graphite account in order to use the merge queue. Sign up using this link.

An organization admin has enabled the Graphite Merge Queue in this repository.

Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue.

This stack of pull requests is managed by Graphite. Learn more about stacking.

@findolor

Copy link
Copy Markdown
Collaborator Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

Comment on lines +1288 to +1297
let by_bytes = build_insert_statements_bounded("items", &columns, &rows, 256, 75)
.expect("bound statement bytes");
assert!(by_bytes.lines().count() > 1);
assert!(by_bytes.lines().all(|line| line.len() < 75));

for sql in [by_rows, by_bytes] {
let conn = Connection::open_in_memory().expect("open database");
conn.execute_batch("CREATE TABLE items (value TEXT NOT NULL);")
.expect("create table");
conn.execute_batch(&sql).expect("import grouped inserts");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

minor: No test covers the two behaviors the description highlights. With max_bytes = 75, no single row is above the limit, so the oversized-row path is never taken. The flush happens at 84 > 75, far from the edge. If you change > to >=, remove the + 2, or remove the statement_rows > 0 guard (which would write a bare ; line), the tests still pass. Also, execute_batch(&sql) parses statements across lines, so it does not test the real importer (value.lines() into one statement per line in pipeline/engine.rs). Add one row that is larger than max_bytes and one tuple that lands exactly on max_bytes. Then import with sql.lines(), one statement per line, as the engine does.

Comment on lines 184 to +185
let values_sql = format_row_values(row, &column_names).map_err(LocalDbError::from)?;
output.push_str(&format!(
"INSERT INTO \"{table}\" ({quoted_columns}) VALUES ({values_sql});\n"
));
let tuple = format!("({values_sql})");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

minor: This problem existed before this PR, but it affects the line-importer question you asked about. format_sql_value only doubles ', so a TEXT value that contains \n puts the statement on two lines. erc20_tokens.name/symbol come from on-chain name()/symbol() without sanitising, so anyone can deploy a token with a newline in its name. value.lines() in pipeline/engine.rs then gives two broken fragments, and the whole bootstrap transaction rolls back. Grouping does not make this worse, because one bad row already failed the import. It is fine as a follow-up: emit such literals with char(10)/replace(...) or as a hex cast, or split statements in the importer in a way that knows about quotes.

Comment on lines 159 to 172
table: &str,
columns: &[TableInfoRow],
rows: &[Value],
) -> Result<String, LocalDbError> {
build_insert_statements_bounded(table, columns, rows, MAX_INSERT_ROWS, MAX_INSERT_BYTES)
}

fn build_insert_statements_bounded(
table: &str,
columns: &[TableInfoRow],
rows: &[Value],
max_rows: usize,
max_bytes: usize,
) -> Result<String, LocalDbError> {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: build_insert_statements only forwards the two constants to build_insert_statements_bounded. One build_insert_statements(table, columns, rows, max_rows, max_bytes), with the constants passed at the call site on line 80 and in the default-limit test, removes the wrapper and the second name.

@graphite-app

graphite-app Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Merge activity

## Chained PRs

- Depends on #2887.

## Motivation

Filtered local DB dumps still contain one `INSERT` per row. That creates a large number of Rust SQL statement objects during browser bootstrap and adds avoidable parser work.

Part of [RAI-2649](https://linear.app/makeitrain/issue/RAI-2649/produce-bounded-grouped-sql-dumps-and-versioned-browser-manifests).

## Solution

- Group consecutive rows from the same table into multi-row `INSERT` statements, with at most 256 rows or 256 KiB per statement. A single row above the byte limit remains a standalone statement so the exporter preserves it.
- Keep each statement on one line inside the existing single `BEGIN`/`COMMIT` dump transaction, so current line-based dump importers can read it.
- Add tests for row and byte limits, SQL escaping, and round-trip imports into SQLite.

This PR does not change the manifest schema, browser importer, or sqlite-web. The bounded atomic browser import API is tracked in [sqlite-web#35](rainlanguage/sqlite-web#35); wiring it into raindex remains a separate follow-up.

## Checks

- `nix develop .#rust-shell --offline -c cargo fmt --all -- --check` — passed.
- `nix develop .#rust-shell --offline -c cargo clippy --workspace --all-targets -- -D warnings` — passed.
- `nix develop .#rust-shell --offline -c cargo test --workspace -- --test-threads=2` — passed. The default parallel run timed out in unrelated local Anvil fixture tests; limiting concurrency resolved it.
- CI-equivalent `nix develop .#wasm-shell --offline` Wasm test command — passed; this host reported no runnable Wasm tests.
- Three read-only simplification passes, two read-only local Codex reviews, and a CodeRabbit review — no actionable findings.

Review focus: confirm multi-row `VALUES` remains compatible with the current one-statement-per-line dump importer. Unrelated unstaged benchmark experiments in the workspace are excluded from this PR.
@graphite-app
graphite-app Bot force-pushed the arda/rai-2648-skip-unused-event-rows branch from 66adc15 to 1e8019b Compare September 30, 2026 11:11
@graphite-app
graphite-app Bot force-pushed the arda/rai-2649-bounded-grouped-sql-dumps branch from a018a4b to 984785c Compare September 30, 2026 11:11
@graphite-app
graphite-app Bot changed the base branch from arda/rai-2648-skip-unused-event-rows to main September 30, 2026 11:25
@graphite-app
graphite-app Bot merged commit 984785c into main Sep 30, 2026
19 of 20 checks passed
@github-actions

Copy link
Copy Markdown
Contributor

@coderabbitai assess this PR size classification for the totality of the PR with the following criterias and report it in your comment:

S/M/L PR Classification Guidelines:

This guide helps classify merged pull requests by effort and complexity rather than just line count. The goal is to assess the difficulty and scope of changes after they have been completed.

Small (S)

Characteristics:

  • Simple bug fixes, typos, or minor refactoring
  • Single-purpose changes affecting 1-2 files
  • Documentation updates
  • Configuration tweaks
  • Changes that require minimal context to review

Review Effort: Would have taken 5-10 minutes

Examples:

  • Fix typo in variable name
  • Update README with new instructions
  • Adjust configuration values
  • Simple one-line bug fixes
  • Import statement cleanup

Medium (M)

Characteristics:

  • Feature additions or enhancements
  • Refactoring that touches multiple files but maintains existing behavior
  • Breaking changes with backward compatibility
  • Changes requiring some domain knowledge to review

Review Effort: Would have taken 15-30 minutes

Examples:

  • Add new feature or component
  • Refactor common utility functions
  • Update dependencies with minor breaking changes
  • Add new component with tests
  • Performance optimizations
  • More complex bug fixes

Large (L)

Characteristics:

  • Major feature implementations
  • Breaking changes or API redesigns
  • Complex refactoring across multiple modules
  • New architectural patterns or significant design changes
  • Changes requiring deep context and multiple review rounds

Review Effort: Would have taken 45+ minutes

Examples:

  • Complete new feature with frontend/backend changes
  • Protocol upgrades or breaking changes
  • Major architectural refactoring
  • Framework or technology upgrades

Additional Factors to Consider

When deciding between sizes, also consider:

  • Test coverage impact: More comprehensive test changes lean toward larger classification
  • Risk level: Changes to critical systems bump up a size category
  • Team familiarity: Novel patterns or technologies increase complexity

Notes:

  • the assessment must be for the totality of the PR, that means comparing the base branch to the last commit of the PR
  • the assessment output must be exactly one of: S, M or L (single-line comment) in format of: SIZE={S/M/L}
  • do not include any additional text, only the size classification
  • your assessment comment must not include tips or additional sections
  • do NOT tag me or anyone else on your comment

findolor added a commit to rainlanguage/sqlite-web that referenced this pull request Sep 30, 2026
## Related PRs and issue

- [RAI-2650](https://linear.app/makeitrain/issue/RAI-2650/add-bounded-atomic-sql-dump-import-to-sqlite-web-and-release-it)
- Producer grouped SQL change: [raindex#2893](rainlanguage/raindex#2893). This PR adds an upstream import API for the subsequent browser integration; neither PR needs the other to merge.
- Builds on the existing `transaction(statements)` API from [sqlite-web#28](#28).

## Motivation

The browser bootstrap currently materializes Rust and JavaScript statement objects for the SQL dump and sends the array through one transaction callback. Grouping rows on the producer side reduces this work, but the browser still needs a bounded way to stream SQL text into one atomic import without exposing partial data across tabs.

## Solution

- Add `beginSqlDumpImport`, `appendSqlDumpChunk`, `finishSqlDumpImport`, and `cancelSqlDumpImport` to the public wasm API. The database worker owns one `BEGIN IMMEDIATE` transaction across chunks and commits only after a successful finish; SQL errors, invalid chunks, cancellation, and abandoned sessions roll back.
- Parse SQL incrementally across chunk boundaries, including strings, comments, and trigger bodies. Bound each UTF-8 chunk to 512 KiB and an unfinished statement to 16 MiB. Reject row-returning statements during import to avoid collecting unbounded results.
- Reject connection PRAGMA settings, `EXPLAIN`, and `ATTACH`/`DETACH` before preparation so failures cannot leave connection changes that transaction rollback would not undo. Accept and skip the standard `PRAGMA foreign_keys=OFF;`/`=0` dump header while preserving existing foreign-key enforcement.
- Reject unpaired UTF-16 surrogates before conversion or dispatch. A cached JavaScript regex validates each chunk with one call across the Wasm boundary.
- Run all core import tests against SQLite's memory VFS without silently skipping failed opens, alongside real OPFS browser integration. Cover partial lexical states, triggers, statement limits, rollback, cancellation, and caller-visible leader-loss outcomes.
- Route import actions through the existing leader/follower coordinator. Reject competing queries and transactions while an import is active, limit outstanding import work per client instance, and expire abandoned imports after 120 seconds of inactivity on the next request or 30-second watchdog tick. An import response is reported only after the database worker resolves it, so a delayed commit cannot be reported as a follower timeout.
- Document the API and limits; add browser integration tests and an isolated synthetic benchmark.

This PR does not change the producer, the browser bootstrap call site, the dump format, or the published package version. The repository's main-branch release workflow publishes the next patch version after merge.

## Checks

- [x] `nix develop -c build-submodules`
- [x] `nix develop -c local-bundle`
- [x] `nix develop -c rainix-rs-static`
- [x] Core WASM tests in Chrome: 124 passed (15 core import tests execute through the memory VFS)
- [x] Public WASM tests in Chrome: 42 passed
- [x] `nix develop -c npm test` in `svelte-test`: 170 passed, 4 existing skips
- [x] `nix develop -c npm run lint-format-check` in `svelte-test`
- [x] `git diff --check`

The latest WASM suites passed under the standard `nix develop -c test-wasm` wrapper with a matched Chrome for Testing / ChromeDriver 147 pair configured locally. The initial run with the system ChromeDriver 144 failed before tests started; the matched pair passed all tests. The packaged browser integration suite passed under Playwright Chromium.

`nix develop -c npm run test:benchmark` also passed. Its three sequential cases import 10,000 equal rows: single-row transaction including construction 59.6 ms, grouped transaction 12.4 ms, and identical grouped SQL through import chunks 15.9 ms. This supersedes the earlier two-case comparison: grouping and the import API must be measured separately. This is a fixed-order smoke measurement with warm storage and small chunks, not an end-to-end or production speedup estimate; download, large-dump object allocation, and indexes require raindex benchmarks.

Review focus: cross-tab serialization, transaction-marker handling, containment of connection state, and rollback behavior when a chunk or finish fails. A leader change before an import response reports an unknown outcome to the caller. Worker failure during finish can leave the outcome unknown even when an error response arrives; check/reset before retrying. Any lost finish response requires checking/resetting the database before retrying; the new test deliberately stalls a DB-worker response to make that path deterministic.

Local Codex review completed three initial rounds with four reviewers, followed by two follow-up rounds with three reviewers, with no OpenCode supplement. Three simplification passes checked the follow-up fixes. All seven latest review comments are addressed: skipped headers are excluded from counts, empty dumps are consistently rejected, unmatched markers have a specific error, the unused queue guard is removed, benchmark APIs use identical grouped SQL, and the supported dumps and liveness/unknown-outcome limitations are documented. The earlier fixes also close the `EXPLAIN PRAGMA` bypass and prevent attachment changes from surviving cancellation.



<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **New Features**
  * Added support for importing SQL dumps in chunks, with controls to begin, finish, or cancel an import across connected clients.
  * Imports validate chunk encoding and SQL statements, reject unsupported operations, and roll back on errors or after 120 seconds of inactivity. Regular database operations are blocked while an import is active.
* **Documentation**
  * Added guidance on streaming SQL dumps, import constraints, failure handling, and checking the database before retrying when the commit outcome is unknown.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
findolor added a commit to rainlanguage/rain.local-db.remote that referenced this pull request Sep 30, 2026
## Motivation

The remote producer still pins raindex before the unused-row and grouped-insert changes. Browser downloads therefore continue to contain raw RPC logs and signed take-order context rows, with one INSERT per retained row.

Part of [RAI-2652](https://linear.app/makeitrain/issue/RAI-2652/update-remote-db-producer-for-optimized-browser-artifacts).

## Solution

- Update `lib/raindex` from `6a78f29134f3125800ebae46db738ecb19077c69` to merged raindex main commit `984785c8a620072f515dc070fec0e732c9bfd156`.
- Include [raindex#2887](rainlanguage/raindex#2887), which stops new raw/context writes and omits `raw_events`, `take_order_contexts`, and `context_values` from exports, including rows inherited from older dumps.
- Include [raindex#2893](rainlanguage/raindex#2893), which groups retained INSERT rows with limits of 256 rows or 256 KiB per statement; oversized individual rows remain standalone.

The SQL dump filename, gzip encoding, manifest schema, database schema, and publication flow remain compatible. The current browser transaction importer already supports these multi-row INSERT statements. This producer update does not depend on sqlite-web#35 or enable deferred browser indexes.

## Rollout

After merging, dispatch the existing deploy workflow for the current deployment ID and settings URL. It rebuilds and installs the CLI from this pinned submodule, then starts the producer.

The first successful run downloads the existing published dump, imports it into a temporary database, indexes newer blocks, and exports a filtered, grouped replacement. The runner uploads the dumps before publishing the manifest. No separate filtering script, empty manifest, or full historical reindex is needed.

Verify each target's published dump omits the three unused tables, contains grouped INSERTs, and retains its watermark. Verify a fresh browser import and order/vault queries. Failed targets can retain their previous manifest entries and old dumps. Keep the previous manifest/dump objects available before rollout for rollback.

## Checks

- [x] Confirm the pinned revision includes both upstream changes.
- [x] `git diff --check`.
- [x] `bash -n prep.sh nixos/local-db-remote-run.sh`.
- [x] `nix run .#build-raindex-cli` — release CLI build passed on aarch64 Darwin with Rust/Cargo 1.94.
- [x] CLI `--help` and `local-db sync --help` smoke checks.
- [ ] Live deployment, publication, and fresh-browser verification; this PR does not execute them.


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->

## Summary by CodeRabbit

* **User-facing changes**
  * No user-visible changes are described in the available summary. Any impact of this update is unclear, so no feature or behavior changes are listed.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants