Skip to content

Estimate distinct counts after a filter with the survival formula in FilterExec #26052

Description

@asolimando

Is your feature request related to a problem or challenge?

FilterExec::statistics_from_inputs caps each column's distinct count at the filtered row count. When values repeat, this overestimates. For example, 1000 rows with 100 distinct values (10 rows per value) filtered to 200 rows gives min(100, 200) = 100, while the expected number of surviving values is about 89.

The deprecated FilterStatisticsProvider (#25969) implements a better estimate, but only as an opt-in provider.

Describe the solution you'd like

Apply the survival formula in FilterExec itself, for columns the predicate does not constrain:

NDV_after = NDV * (1 - (1 - selectivity)^(rows / NDV))

capped at the current estimate. This is ndv_after_selectivity in operator_statistics/mod.rs.

Additional context

Evaluating it needs distinct counts in the stats benchmark data, see #25576.

Activity

  1. asolimando commented on Oct 5, 2026

    @asolimando
    MemberAuthor

    take

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions