Skip to content

[mirror] feat: deterministic N-Triples serialization via an opt-in canon parameter - #1

Open
jdsika wants to merge 1 commit into
mainfrom
feat/deterministic-ntriples
Open

jdsika wants to merge 1 commit into
mainfrom
feat/deterministic-ntriples

Conversation

@jdsika

@jdsika jdsika commented Sep 14, 2026

Copy link
Copy Markdown

Mirror of RDFLib#3546 — tracking only. Do not merge.

Upstream PR: RDFLib#3546

Why this PR exists

Every fix this organisation depends on is written as a self-contained, single-commit
upstream contribution rather than a local patch, so that a fork is only needed for as
long as its contributions are unmerged.

That convention has one weakness: an upstream PR is easy to lose track of, because it
lives in someone else's repository and closes on someone else's schedule. This mirrored PR
is the counterweight — an open PR in our own fork for every contribution still pending
upstream, so the set of things we are carrying is visible from our own repository list.

It is closed when the upstream PR is merged or rejected, not before.

Contents

A single commit, identical in content to the upstream PR head:

feat/deterministic-ntriples -> main

Where it is used

Not yet consumed: the pipeline pins released rdflib==7.6.0. Relevant to diffable-rdf, which re-serializes canonicalized graphs through rdflib.

…eter

The N-Triples serializer writes statements in the order the store iterates
them, which follows set iteration and so varies between processes. The same
graph therefore serializes to different bytes in different runs, which shows
up as diff noise in version-controlled RDF.

PR RDFLib#3008 solved this for longturtle with an opt-in `canon` parameter that
canonicalizes the graph and sorts, fixing RDFLib#1890. This gives NTSerializer the
same parameter, with the same default of False, so ordinary serialization is
untouched and still streams statement by statement.

`canon=True` canonicalizes with to_canonical_graph and writes the rendered rows
in sorted order. Sorting the rendered rows rather than the triples keeps the
order total: two distinct terms can share a str(), which would leave their
relative order down to the store iteration order that this is meant to remove,
whereas identical rows mean identical statements.

Measured over six PYTHONHASHSEED values on a graph with a hub, a shared list
tail and a deep blank-node chain: 5 distinct documents before, 1 after.
NT11Serializer inherits it by subclassing.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant