Re-scope 000031 away from DIET: the contact/nanowire mechanism was never sourced - #262
Merged
Conversation
…ver sourced (#256) Codex review of #256 plus caching the discovery study's OA full text (PMID:28287150, PMC5347079) showed this is not a two-papers-disagree controversy. The contact-dependent DIET mechanism was never supported by the record's own cited source. - "pili"/"nanowire" appear nowhere in that paper's abstract, and in the full text only as generic background about Geobacter species citing prior work. - Its discussion states the opposite of what this record claimed: pre-cultures were "constituted of both nanowire-rich aggregates and nanowire-poor planktonic cells", and growth on conductive material is proposed as a future option "to ensure electrical connections" -- connection was NOT established. - Several evidence items cited that generic background sentence, or the unrelated "sole electron acceptor" sentence, as if they demonstrated contact-mediated transfer between these two organisms. One was also truncated mid-word. Changes: community_category DIET -> SYNTROPHY (the schema defines DIET specifically as "Direct interspecies electron transfer", a mechanism claim no source supports here, while SYNTROPHY is the neutral bucket); pili/nanowire assertions removed from description, environment notes, taxon notes and the interaction; both interactions renamed to mechanism-neutral forms with the downstream edge and discussion anchors updated; "Cell Contact and Nanowire Formation" environmental factor becomes "Electrical Connection Between Cells", now carrying a REFUTE item quoting the authors' own future-work sentence; DIET-as-shorthand softened throughout the prose. The competing mechanism is NOT asserted in its place -- PMID:34939136 calls its own cobamide model "putative", and is not open access, so its cell-free spent-medium result stays unverified. The discussion is rewritten from "contested, curator decision pending" to a record of what was re-scoped and why. Left to a curator: the record `name` and filename still say "DIET"; changing them affects external references. Verification: validate clean; snippet audit MATCH 4107 -> 4108 with MISMATCH unchanged at 166; network-integrity unchanged at 40, and the new dangling-anchor check confirms the renames left no broken references. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Jul 28, 2026
Closed
realmarcin
added a commit
that referenced
this pull request
Jul 29, 2026
…thread) (#263) Reconciled against merged PRs since the stale 2026-07-21 date. Marked DONE: - 000031 re-scoping (#256) — was "curator decision still open"; resolved by PR #262. Added a dedicated section recording what changed and why it was a defect rather than a controversy. - Li et al. 2024 ingestion (#259) — was "still not ingested"; PR #261 cached it via the new --from-file path and curated 000068. - §2 "apply modeled_environment matching to the ingredient suggester" — this was never actually pending: suggest_related_ingredients.py has read modeled_environment since PR #220 that created it. §2/#30 now has no actionable remainder here. Corrected a wrong claim that was sitting in the file: the Li 2024 summary said "5-10 mM promotes, >=30 mM inhibits". The >=30 mM part came from the Edison report and describes the paper's anaerobic-sludge system; in the coculture the response is non-monotonic (30 mM recovers in the later stage, only 50 mM inhibits). Also noted that one quoted snippet from that report appears nowhere in the paper. Newly logged: PR #255 (Suillus-Bacillus thiamine SynCom) has been open since 2026-07-26, non-draft, MERGEABLE/CLEAN with all gates green, and was absent from this file entirely. Still open and unchanged: #257 (validate-references reporting), #258 (14 dangling edges, triage), #259 (automated retrieval from blocking publishers), and the upstream-blocked term-minting items in sections 0/1. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
realmarcin
added a commit
that referenced
this pull request
Jul 29, 2026
Follow-up to #262, which re-scoped the record after finding the contact-dependent DIET mechanism was never supported by its cited source. The human-readable label and path still asserted it. - kb/communities/Geobacter_Clostridium_DIET.yaml -> kb/communities/Geobacter_Clostridium_Interspecies_Electron_Transfer_Coculture.yaml - name: "Geobacter-Clostridium DIET Community" -> "Geobacter-Clostridium Interspecies Electron Transfer Coculture" The id CommunityMech:000031 is UNCHANGED. It is the stable cross-repo key; only the label and path moved, so external references by id keep resolving. Derived artifacts were regenerated rather than hand-edited: just gen-browser, gen-html (305 communities + browser + landing), gen-umap, gen-community-pages, and scripts/generate_validation_report.py. docs/community_graph.html is the one exception -- no generator recipe exists for it in this repo, so its three references were replaced textually; it should be regenerated from source if a recipe is added. The remaining "DIET" strings in NEXT_TASKS.md are the intentional before -> after notation in the rename note. Verification: record validates; no stale references outside research/ artifacts. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Aug 2, 2026
realmarcin
added a commit
that referenced
this pull request
Aug 2, 2026
Two records carried a duplicate mapping key. PyYAML keeps the last of a pair and reports nothing, so in each case one curated value was being discarded at parse time while every gate stayed green. **Geobacter/Clostridium — a retracted claim was winning.** The evidence item for PMID:28287150 had two `explanation` values. Git shows why: PR #262 ("Re-scope 000031 away from DIET: the contact/nanowire mechanism was never sourced") *inserted* its correction but left the old line in place as context, so the parsed value was the very claim #262 set out to retract — its re-scoping was inert in the data while looking applied in the file. Which one survives is settled by the record itself, not by preference: the surviving text must be the one beginning "PARTIAL -", because the item's own `supports: PARTIAL` agrees with it, and because that text explicitly describes the other as the wording it replaced. The stale line is deleted. **Trichodesmium/Alteromonas — an orphaned note, not a redundant one.** The iron(2+) metabolite had two `notes`; the ROS one was winning, so the note explaining the iron CHEBI grounding was lost. But the ROS note is not surplus: the same interaction cites "detoxification of reactive oxygen species" as evidence, and the record already curates reactive oxygen species as a compound (CHEBI:26523) with its own relevance statement. The interaction's metabolite list was simply missing that third entry. So rather than delete a curated statement, the note is given its proper home — a `reactive oxygen species` metabolite grounded to the CHEBI term the record already uses — and the iron note is restored to iron(2+). KNOWN_DUPLICATES is now empty: zero duplicate keys repo-wide. The comment above it explains the bar for adding one back. Filed separately, not fixed here: #295 — the same DIET-background snippet is still cited as `supports: SUPPORT` with no explanation in this record's second interaction, so #262's judgement was applied to one occurrence and not the other. 589 tests pass; both records pass schema and id<->label validation. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
realmarcin
added a commit
that referenced
this pull request
Aug 4, 2026
…mis-tiered items The review's most severe finding was right: the file asserted "#319 is decided" while the issue body still says "unresolved" and had zero comments, so the claim existed in no citable place. In a decision-support document that is the worst failure — it tells a reader to skip a decision the tracker says is open. The decision is now recorded as a comment on #319 and cited by permalink, and the file says explicitly to cite the comment rather than the body. Its numbers were also unreconciled. Re-measured on `main`: 13 host/antagonist participant slots across **9** records (not 12) and 14 placeholder slots across 9. Both differ from the issue body's 23-across-17, which predates #345 and used a broader criterion; the file now says so instead of quietly disagreeing. **Two items were in the wrong tier, both by my own stated criterion.** #295 is not decision-free: the issue asks for PARTIAL *or* dropping the citation and names the curator who made #262's call as the decider. It is also not a clean pair — the SUPPORT occurrence is 150 chars and truncated mid-word against the other two at 188, so it needs a truncation repair too. Moved to Tier 2. #350's done-when hid a judgement. All 4 isolates fail term validation, but the failures are mostly wrong *id*, not wrong label — CHEBI:30319 recorded as "dicyanoaurate(1-)", ENVO:00000072 as "mine tailing", GO:0055114/GO:0055065 obsolete. Picking the right id per term is what id-label-correspondence reserves for a curator. Moved to Tier 2 with the note that the brief must choose which branch to take. #358 was listed as "ready now" while the same file declared it blocked on #357. It moves to its own queued bucket, and now states the byte question plainly: 4015 bytes on main is already over 4000, so if the ceiling counts bytes the file has been over all along — which is the question #358 exists to settle. Corrected numbers: #352a is 7 records without a page, not 1 (the loop would have had to decide commit-all vs hand-pick unbriefed); #306 is 62 stems with both .md and .txt exactly, 63 folding case; #325 is 310 of 312, not 311. Also noted that a #352a PR cannot close #352, since that issue carries the duplicate-SPRUCE question too — so the loop's "issue closed" finish condition will not fire. Pointers: NEXT_TASKS.md's link moved off the "Last reconciled:" line, since a naive `s/^Last reconciled:.*/` bump would have deleted it (verified it now survives); CLAUDE.md listed the derived file but not the primary backlog, and now lists both.
realmarcin
added a commit
that referenced
this pull request
Aug 4, 2026
#360) * Add NEXT_TASKS_LOOP.md: which open issues suit an autonomous /goal run `NEXT_TASKS.md` says what is deferred. It does not say what can be handed to a loop that will not stop to ask, and that is a different question — an item needing a curation or schema decision stops on the loop's first substantive step and wastes the run. All 26 open issues are classified into three tiers plus a never-loop set, with the criterion stated up front: a machine-checkable definition of done, no curation decision, bounded blast radius, and a premise that survives measurement. Every claim was re-measured against `main` today rather than copied from the issue text, which matters because half the issues in this repo have turned out wrong on inspection (#273, #276, #310, #346). Verified here: `uv sync --group dev` still fails; the DIET snippet is still cited at both PARTIAL and SUPPORT; `NCBITaxon:1125` is still the one ungrounded taxon of four in that record; 4 of 4 isolates fail term validation; 63 references have both a .md and a .txt in the cache; 312 records against 305 generated pages; 4 dangling wiki-links. Tier 1 is eight items with green/red finish conditions, recommending #290 first — one line, exits 0 or doesn't, and it retires a gotcha the goal prompt has to carry. Tier 2 is three that are automatable only with a brief that constrains judgement; #347 in particular needs an explicit "use exact substrings, delete what you cannot source" or the failure mode is fabricating evidence. Tier 3 lists eleven where the decision needed is named, so it can be answered in one pass. Also records the ordering constraints: #358 waits on PR #357, the three SPRUCE issues all edit one file, and #314 should precede #294 so the enum backfill has correct data under it. Linked from CLAUDE.md and NEXT_TASKS.md — a new doc nothing references is invisible, which was a review finding on the last one (#344). * NEXT_TASKS_LOOP: quote main's goal-prompt size, not the unmerged branch's The #358 row cited 3995 chars / 4021 bytes / 5 spare — the numbers from PR #357, which is still open. This file merges into main, where the prompt is 3987 / 4015 with 13 spare, so a reader measuring it would have concluded the row was wrong. Now states main's figures and flags what #357 changes them to. * Address the review of #360: record the #319 decision, and demote two mis-tiered items The review's most severe finding was right: the file asserted "#319 is decided" while the issue body still says "unresolved" and had zero comments, so the claim existed in no citable place. In a decision-support document that is the worst failure — it tells a reader to skip a decision the tracker says is open. The decision is now recorded as a comment on #319 and cited by permalink, and the file says explicitly to cite the comment rather than the body. Its numbers were also unreconciled. Re-measured on `main`: 13 host/antagonist participant slots across **9** records (not 12) and 14 placeholder slots across 9. Both differ from the issue body's 23-across-17, which predates #345 and used a broader criterion; the file now says so instead of quietly disagreeing. **Two items were in the wrong tier, both by my own stated criterion.** #295 is not decision-free: the issue asks for PARTIAL *or* dropping the citation and names the curator who made #262's call as the decider. It is also not a clean pair — the SUPPORT occurrence is 150 chars and truncated mid-word against the other two at 188, so it needs a truncation repair too. Moved to Tier 2. #350's done-when hid a judgement. All 4 isolates fail term validation, but the failures are mostly wrong *id*, not wrong label — CHEBI:30319 recorded as "dicyanoaurate(1-)", ENVO:00000072 as "mine tailing", GO:0055114/GO:0055065 obsolete. Picking the right id per term is what id-label-correspondence reserves for a curator. Moved to Tier 2 with the note that the brief must choose which branch to take. #358 was listed as "ready now" while the same file declared it blocked on #357. It moves to its own queued bucket, and now states the byte question plainly: 4015 bytes on main is already over 4000, so if the ceiling counts bytes the file has been over all along — which is the question #358 exists to settle. Corrected numbers: #352a is 7 records without a page, not 1 (the loop would have had to decide commit-all vs hand-pick unbriefed); #306 is 62 stems with both .md and .txt exactly, 63 folding case; #325 is 310 of 312, not 311. Also noted that a #352a PR cannot close #352, since that issue carries the duplicate-SPRUCE question too — so the loop's "issue closed" finish condition will not fire. Pointers: NEXT_TASKS.md's link moved off the "Last reconciled:" line, since a naive `s/^Last reconciled:.*/` bump would have deleted it (verified it now survives); CLAUDE.md listed the derived file but not the primary backlog, and now lists both. * Address the review of #360: the #319 counts didn't follow the criterion I stated The review's P1 is right, and it is the worst kind of error for this file: the counts I published contradicted the criterion published beside them. My comment on #319 said the figures were "restricted to participants that resolve to no taxonomy entry", then gave 13 and 14 — which came from the *auditor's* rule (no name match AND the id is not unique), not that one. Re-measured over all 1127 participant slots in 312 records: criterion non-placeholder NCBITaxon:2 total id appears nowhere in that taxonomy 10 / 8 rec 13 / 8 rec 23/16 unresolved by the auditor 13 / 9 rec 14 / 9 rec 27/18 The first is the criterion for this decision, and it reproduces the 23 in the issue body exactly — only the record count moved, 17 to 16, after #345. So my aside that the issue "used a broader criterion" was backwards: the issue's was the tighter one, mine was looser. The four extra slots are name variants of members already in taxonomy — "Olsenella (Actinobacteriota)" against an id listed twice, "Variovorax" against one listed six times, "Bacillus SynCom" against one listed four times, plus Saanich Inlet's aggregate. They need a rename, not a new entry, and calling them "host/antagonist" was wrong: Variovorax and Olsenella are ordinary members and "Bacillus SynCom" is an aggregate belonging with the placeholders. So "the 13 each need a grounded term and a snippet" was false for at least three of them. The GitHub comment is corrected too, since this file tells readers to cite it in preference to the issue body — fixing only the file would have left the citable record wrong. **#277 moves to Tier 2.** It fails the same test that demoted #295 and #350: the issue offers three mutually exclusive remedies, the substance lives in a memory directory the issue records as absent, and "no dangling links" is satisfiable by deletion — which discards what the issue calls load-bearing. **#358 moves out of "Tier 1 — ready now"** into its own Queued section. Listing a blocked item under a heading that says ready is exactly the trap a loop reading top-down falls into. Also: the `id-label-correspondence` claim was overstated — the skill never reserves that call for a curator, it prescribes `validate_ncbitaxon_ids.py` and `term_fix_apply.py`; the demotion stands on its other ground. And #359 is dropped from "Never loop these", having been closed as filed-on-a-false-premise. * NEXT_TASKS_LOOP: #358 is unblocked now that #357 has merged #357 merged as 5a1d60b, so the Queued section it justified is gone and #358 returns to Tier 1. Its figures are re-measured against the new main: 3998 chars and 4028 bytes, leaving 2 characters of headroom — and the bytes now exceed 4000 by 28, which sharpens rather than settles the char-or-byte question #358 exists to answer.
realmarcin
added a commit
that referenced
this pull request
Aug 7, 2026
* Apply #262's DIET judgement to the occurrence it missed (#295) The record cites the same PMID:28287150 snippet three times, not twice as the issue says. #262 re-scoped two to PARTIAL on the grounds that it is the paper's generic background about Geobacter species in prior literature, not a finding about this coculture. The third stayed SUPPORT, so the record simultaneously held that the snippet establishes nothing here and that it fully supports an interaction. Worse than the issue recorded: that third occurrence also had no explanation and a snippet truncated mid-word at "with o". #262's note on the first occurrence says "Snippet also completed here, having previously been truncated mid-word" - the same truncation, fixed in one place and left in another. Re-scoped to PARTIAL with the same reasoning, snippet completed, explanation added. The interaction is unaffected either way: its other two items, the metabolic-shift quantification and the electron-uptake measurement, are what carry it - which is what the issue predicted. Writing that explanation reintroduced #400's defect, in the file I was fixing: an unquoted "#262" in a plain multi-line scalar starts a YAML comment, so everything after it was swallowed and the file stopped parsing. `just validate` caught it. Quoted, and swept the KB for the same pattern - three continuation lines contain " #", all inside quoted scalars, and validate-scalars reports 0 truncated across 318 files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address the #465 review: eight truncated snippets, not one I claimed the KB was clean on the basis that validate-scalars reports 0 truncated across 318 files. That check structurally cannot see this defect - it flags a plain scalar swallowed by a "#" comment, and these are well-formed scalars that merely stop early. It reports 0 correctly and says nothing. The reviewer's sweep - resolve each snippet against its cached source, flag any that matches verbatim but is followed in the source by a letter - found seven more, in five records: "...ammonia-oxidizing bac" -> bacteria (x2) "...and Stenotrop" -> Stenotrophomonas (x5) "...elemental sulf" -> sulfur "...were potential" -> potentially "...propionate usi" -> using protons as the electron acceptor "...were upregu" -> upregulated in D. vulgaris "...base of the Thermo" -> Thermoplasmatales within the Euryarchaeota All completed verbatim from references_cache/, and the sweep is now a test so the class cannot recur silently. Digits following a snippet are excluded - those are citation markers the cached markdown ran together with the preceding word - and one letter case is allowlisted with its reason: "parvus" + "cocultivated" is a missing space in the cache, not a cut quote. Also stopped the new explanation vouching for a sibling item that does not do what its own explanation claims: that snippet is the paper's study-scope sentence and quantifies nothing, despite an explanation saying it quantifies the shift toward 1,3-propanediol. Same defect class as this one; noted rather than fixed, since it needs the curator's judgement. Filed #466: just validate-references performs zero checks and passes vacuously, which is how all eight of these survived. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Resolves #256.
I filed #256 as a curator judgement call between two disagreeing papers. It isn't one. A Codex second opinion plus caching the discovery study's OA full text (PMID:28287150, PMC5347079 — it was previously abstract-only) showed the contact-dependent DIET mechanism was never supported by the record's own cited source.
What the source actually says
pili/nanowireappear nowhere in the abstract, and in the full text only as generic background about Geobacter species citing prior work (refs 1,2,4,6,7).The discussion states the opposite of what this record claimed:
Electrical connection was a proposed future improvement, not an achieved condition.
Evidence items cited that generic background sentence — and the unrelated "sole electron acceptor" sentence — as though they demonstrated contact-mediated transfer between these two organisms. One was also truncated mid-word (
…couple the electron balance with o).So the record asserted a mechanism its own source contradicts. That's a defect, not a controversy.
Changes
community_categoryDIET→SYNTROPHYecological_interactions[0]ecological_interactions[1]The schema defines
DIETspecifically as "Direct interspecies electron transfer" (communitymech.yaml:148-149) — a mechanism claim, not a coarse bucket — whileSYNTROPHY= "Syntrophic metabolic cooperation" sits beside it as the neutral option.Also: pili/nanowire assertions stripped from
description,environment_term.notes, taxon notes and the interaction; the downstream edge and both discussion anchors updated for the renames; DIET-as-shorthand softened throughout; the misused snippets downgraded toPARTIALwith honest explanations, and the truncated one completed.What is deliberately NOT asserted
The cobamide mechanism does not replace the old claim. PMID:34939136 calls its own model "putative", and is not open access, so its 0.22-µm cell-free spent-medium result stays unverified against full text. The record now says the route is unresolved and names both hypotheses as putative.
Left to a curator
The record's
nameand filename still say "DIET". Changing those affects external references, so I left them and noted it in the discussion.Verification
just validateclean;just test266 passedDANGLING_ANCHORcheck added in Dangling-edge detection, DOI full-text caching, and a correction to the validate-references note #260 confirms the renames left no broken references🤖 Generated with Claude Code