distinct_count
Generic Description
distinct_count computes the distinct_count aggregate over the rows currently grouped by Cypher. Documenting it separately matters because aggregate placement changes row grain, grouping behavior, and whether state/merge forms are valid.
Simple example:
MATCH (d:Document)
RETURN distinct_count(properties(d).prize) AS value
Consumer-Level Explanation
Use distinct_count when the graph pattern has already selected the row grain and the next step is a trustworthy metric, not another traversal. Keep the aggregate in Cypher when planner visibility, grouped execution, materialized aggregate views, or assistant-authored query validation matter.
More Detailed Explanation
distinct_count operates after MATCH, WHERE, WITH, procedure output, or document/map projection has shaped rows. For planner tooling the important distinction is whether this page documents a final aggregate, a state producer, or a state merger. Final aggregates return business-facing values; state functions return execution-facing summaries; merge functions combine those summaries and should not be treated as ordinary scalar math.
Advanced Example
This example puts distinct_count after graph matching and time bucketing so the aggregate summarizes an explicit business grain rather than an accidental stream of rows.
MATCH (u:User)-[e:VIEWED]->(d:Document)
TIME e.ts BETWEEN datetime('2025-01-01T00:00:00Z') AND datetime('2025-02-01T00:00:00Z')
WITH date_trunc('day', e.ts) AS day_bucket, d
RETURN day_bucket,
distinct_count(properties(d).prize) AS metric_value
ORDER BY day_bucket
Real Use Cases
- daily or cohort-level graph metrics built from matched relationships
- document or payload profiling after extracting map/list fields into rows
- feature and report queries that need the aggregate to remain visible to the planner
Real Limitations And Tradeoffs
- aggregate results only summarize rows that reached the aggregate; broad patterns can still create expensive or misleading groups
- state and merge values are implementation-oriented summaries, not display-friendly business fields
- statistical aggregates require meaningful numeric inputs and enough observations to support interpretation