Skip to content

Grouped aggregate over variable-length paths without materializing walks - #1080

Open
adsharma wants to merge 2 commits into
mainfrom
opt/grouped-varlen-agg
Open

adsharma wants to merge 2 commits into
mainfrom
opt/grouped-varlen-agg

Conversation

@adsharma

@adsharma adsharma commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Closes the query shape in #475:

MATCH (a:NodeType01)-[:EdgeType01*0..3]->(b:NodeType01)
RETURN b.active AS active, count(b) AS total, avg(b.score) AS avg_score
ORDER BY total DESC, active

The plan materialized every (a, b) walk in RECURSIVE_EXTEND before the hash join + aggregate (2.18B walks / ~85 GiB for a two-row answer on the issue graph -> OOM-killed). The existing REACHABLE_COUNT rewrite only covers single-source COUNT(DISTINCT b).

Commits

  1. fix(gds): emit single length-0 row when varlen source reaches itself — PathsOutputWriter::write(tableID) re-ran the full DFS for the source when lowerBound==0, duplicating every non-empty source-to-source walk and dropping the length-0 row (self-loop + *0..3 counted 6 instead of 4). Gate the re-emit on the source having no parent entry, with a writeCoveredSourceWalk hook (no-op default) that the variable-length writer overrides to emit just the empty walk. Regression test in test/test_files/recursive_join/n_n.test.

  2. opt(count): grouped aggregate over variable-length paths without materializing walks — new CountRelTableOptimizer::tryRewriteGroupedReachableCount + Logical/PhysicalGroupedReachableCount. The physical operator counts walks per distinct end node with a level DP over CSR (O(up * E) time, O(V) memory) and folds the counts into the GROUP BY accumulators, reading each end node's properties once. COUNT/AVG weight every walk, matching row-wise aggregation. Scope (v1): full-table forward WALK on one node table, GROUP BY end-node properties/ID, optional COUNT + optional AVG. Test: OptimizerTest.GroupedReachableCount (incl. self-loop + negative shapes).

Validation

  • Unoptimized vs rewritten plans return identical rows on small graphs (incl. self-loops, NULL scores, string keys) and on a 20K-node / 120K-edge random graph (exact match; this caught the engine bug above).
  • Spot check 200K nodes / 1.6M edges: rewritten 1.0s / 62MB peak vs unoptimized 60s / 8GB (second run OOMs in the buffer pool).
  • optimizer_test.cpp syntax-checked locally; full suites (optimizer, e2e incl. recursive_join/tck) left to CI.

Resulting plan (LBUG_DUMP_LOGICAL=1)

=== LOGICAL PLAN (before optimization) ===
EXPLAIN
  PROJECTION [active,total, ... (3 total)]
    ORDER_BY [total,active]
      PROJECTION [active,total, ... (3 total)]
        AGGREGATE [keys=active aggs=total,avg_score]
          PROJECTION [b._ID,b.score, ... (3 total)]
            HASH_JOIN [INNER keys=b._ID]
              HASH_JOIN [INNER keys=a._ID]
                PATH_PROPERTY_PROBE
                  RECURSIVE_EXTEND [VAR_LEN_JOINS]
                SCAN_NODE_TABLE [N a._ID]
              SCAN_NODE_TABLE [N b._ID]
=== END LOGICAL PLAN ===
=== LOGICAL PLAN (after optimization) ===
EXPLAIN
  PROJECTION [active,total, ... (3 total)]
    ORDER_BY [total,active]
      PROJECTION [active,total, ... (3 total)]
        GROUPED_REACHABLE_COUNT [E dir=fwd bound=_0_a nbr=_2_b range=0..3 keys=1 COUNT AVG]
=== END LOGICAL PLAN ===

PathsOutputWriter::write(tableID) re-emitted the source through the full
DFS whenever lowerBound==0, duplicating every non-empty source-to-source
walk and dropping the length-0 row (e.g. a self-loop with *0..3 counted 6
instead of 4). Gate the re-emit on the source having no parent entry and
add a writeCoveredSourceWalk hook (default: historical re-emit, a no-op
for writers that skip the source) which the variable-length writer
overrides to emit just the missing empty walk.
…rializing walks

MATCH (a:T)-[:R*lo..up]->(b:T) RETURN b.k, count(b), avg(b.x) materialized
every (a, b) walk in RECURSIVE_EXTEND before the hash join + aggregate
(2.18B walks / ~85GiB for a two-row answer on the issue-475 graph, OOM
killed); the existing REACHABLE_COUNT rewrite only covers single-source
COUNT(DISTINCT b).

Add CountRelTableOptimizer::tryRewriteGroupedReachableCount plus a
Logical/PhysicalGroupedReachableCount operator pair. The physical
operator counts walks per distinct end node with a level DP over CSR
(O(up * E) time, O(V) memory) and folds the counts into the GROUP BY
accumulators, reading each end node's properties once instead of once
per walk. COUNT weights every walk; AVG weights by walk multiplicity,
matching row-wise aggregation over the unoptimized plan.

Scope (v1): full-table forward WALK from/to the same single node table,
GROUP BY on end-node properties (or end-node ID), one optional COUNT and
one optional AVG on an end-node property. Filters, DISTINCT, grouping by
the start node, masks, limits and non-native storage keep the original
plan.
@adsharma adsharma changed the title opt/grouped varlen agg Grouped aggregate over variable-length paths without materializing walks Sep 29, 2026
@adsharma
adsharma marked this pull request as ready for review September 29, 2026 23:42

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant