Skip to content

[P1] sync.py downloads data and then drops it (updated_at, sources, tags, collection labels, hub metrics) #7

Description

@michellzappa

Problem

scripts/sync.py downloads more than it writes out:

  • updated_at is selected from Supabase but never copied into the record built by build_enriched(), so it is null in all rows of indexes/technologies.json.
  • Sources: technologies_sources is downloaded and never written. The CMS has also moved sources into technology_evidence, so the legacy table may be out of date.
  • Tags: per-technology tag1/2/3 are only exported as per-hub counts in indexes/tags.json, not on each entry.
  • collection_label isn't exported, so cities collections show up as IDs like M7CFmLD9Qx2KxloytEYe6w.
  • Hub fields about, indicator_explainers and metrics_config are downloaded but not written to indexes/hubs.json, which causes [P0] trl / impact / investment are mislabelled for 9 hubs (986 entries) #5.
  • Hubs with no published technologies (e.g. lexicon) are listed with a count of 0.

Fix

  • Pass updated_at through in build_enriched(). While there, fix the misindented "sources" line.
  • Read sources from technology_evidence, falling back to technologies_sources.
  • Add collection_label and tags to the frontmatter, and sources to the body as a ## Sources list.
  • Add about and metrics_config to hubs.json.
  • Leave out hubs whose research row is unpublished or that have 0 published technologies.
  • Add a unit test in scripts/test_sync.py for each of these.

Runbook: docs/improvements/03-schema-v2.md

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions