feat(search): check the index schema on startup and refuse breaking changes - #3197
Conversation
Up to standards ✅🟢 Issues
|
| Metric | Results |
|---|---|
| Duplication | -51 |
🟢 Coverage 67.84% diff coverage · +0.12% coverage variation
Metric Results Coverage variation ✅ +0.12% coverage variation (-1.00%) Diff coverage ✅ 67.84% diff coverage Coverage variation details
Coverable lines Covered lines Coverage Common ancestor commit (d38fbc8) 85544 19909 23.27% Head commit (0e2e287) 85770 (+226) 20064 (+155) 23.39% (+0.12%) Coverage variation is the difference between the coverage for the head and common ancestor commits of the pull request branch:
<coverage of head commit> - <coverage of common ancestor commit>Diff coverage details
Coverable lines Covered lines Diff coverage Pull request (#3197) 342 232 67.84% Diff coverage is the percentage of lines that are covered by tests out of the coverable lines that the pull request added or modified:
<covered lines added or modified>/<coverable lines added or modified> * 100%
NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.
a1b0535 to
c7340ac
Compare
8f27800 to
d05047b
Compare
d05047b to
29b0b27
Compare
68cf5b1 to
e9fc6ae
Compare
e9fc6ae to
c70d884
Compare
c70d884 to
43d41fa
Compare
…hanges Both engines now diff the stored/live index schema against the schema generated from code when the service starts. A shared recursive classifier in the mapping package is the single oracle: - equal: start normally. - additive (new fields without any indexed data): applied in place. OpenSearch gets a PUT _mapping with the full code properties, bleve persists the code mapping into the index (SetInternal + reopen) so the new fields are properly typed immediately and later startups classify equal. A startup warning lists the new fields because documents indexed before the upgrade lack them until re-indexed. - breaking (changed definitions or analyzers, removed or renamed fields, or new fields that already contain data of unknown form): refuse to start with an error describing the rebuild procedure (delete the index, start, run "opencloud search index --all-spaces") and the OC_EXCLUDE_RUN_SERVICES=search escape hatch. PUT _mapping is deliberately only the apply mechanism, never the judge: its merge semantics cannot see removals or renames and it accepts in-place updatable param changes with an ack. bleve additionally checks idx.Fields() so previously dynamically indexed data (which leaves no schema trace in bleve) is caught, matching by exact name and by path prefix. While at it: the OpenSearch startup check runs with a real, minute-bounded context instead of context.TODO(), bleve indexes are opened with a 5s bolt_timeout so a second process fails fast instead of hanging on the file lock, and the reversed errors.Is arguments in bleve.NewIndex were fixed. #3092
… delete step Addresses the two Copilot review comments on the PR: the additive opensearch log now matches the bleve warning (level and re-index hint), and the refuse message spells out how to delete the index per engine (DELETE /<name> vs removing the bleve directory).
… tests - shorten the multi-line doc comments flagged as too verbose - add reconcile_test.go: direct unit tests for Reconcile incl. the persisted-but-errored and classify-error branches (previously only reached indirectly through the engine integration tests) - convert the 11 near-identical Classify It blocks to a DescribeTable
- export mapping.SortedUnionKeys and reuse it in bleve.compareKeysExcept instead of a copied union-of-keys block - add Classification.AddBreaking to fold engine-specific breaking reasons and force the verdict, replacing the identical block in the bleve and opensearch Classify paths
The opensearch-go bump renamed the mapping-get accessor, the schema check takes a context and a logger now, and the golden bleve mapping carries the word-broken Name and Title.
The refuse specs use registered analyzers (fulltext is gone), the golden regenerates via UPDATE_GOLDEN, MappingGetResp grew an accessor, and the parity suite passes the new NewBackend signature.
Pins the shipped schema as a reviewable diff; regenerate with UPDATE_GOLDEN=1.
9341733 to
e11ce95
Compare
The failure runs the classifier on golden vs generated: additive means regenerate only, breaking means bump too.
|
🤔 so how do we migrate in k8s? we can update all pods. the old ones will answer from the old index, the new ones will automagically create a new _v3 index, which is empty. a change in a space will trigger a reindex of that space. Or the admin executes a reindex cmd manually ... so forward update. search requests from new pods will go to the new index, from old pods to the old ... that is at least consistent. |
Yes, you are right about all of that - BUT that's an already existing issue. Completely independent from my first refactor and also this one. This PR is just a mechanism to detect divergence and add fields to the index A real waterproof approach would probably take care of migration, possibly with event replay for different versions, I don't know. I agree it would be nice to have, but totally out of scope right now |
Stacked on #3345 (base
refactor/search-mappingis #3345's branch, so this PR shows only its own commits; retarget tomainonce #3345 merges, both ship together). Refs #3092.Two mechanisms keep the index schema and the code in sync.
Versioned indices (from #3345). The index name carries
search.SchemaVersion(opencloud-resource-v<N>,bleve-v<N>). An intended breaking change bumps the constant: the service targets a fresh, empty index and leaves the old one in place. No migration or index-to-index reindex; content lives only in the index and the files are the source of truth, soopencloud search index --all-spaces --force-rescanrebuilds everything.Startup schema check (this PR). The bump is manual, so this is the safety net for a schema change made without one. On startup both engines diff the existing index schema against the schema generated from code, via one shared recursive classifier:
PUT _mapping; bleve persists the code mapping (SetInternal+ reopen). A startup warning lists the new fields (documents from before the upgrade lack them until re-indexed).SchemaVersionor revert) and lists the diff.Why this shape:
PUT _mappingis just the apply step (its merge semantics hide removals/renames), so it must not judge.idx.Fields(): a now-explicit field that already holds dynamic data of unknown shape is breaking (no search beats wrong search).Known and accepted:
bolt_timeout) instead of hanging.Deliberate consequences (not bugs):