Is your feature request related to a problem or challenge?
FilterExec::statistics_from_inputs caps each column's distinct count at the filtered row count. When values repeat, this overestimates. For example, 1000 rows with 100 distinct values (10 rows per value) filtered to 200 rows gives min(100, 200) = 100, while the expected number of surviving values is about 89.
The deprecated FilterStatisticsProvider (#25969) implements a better estimate, but only as an opt-in provider.
Describe the solution you'd like
Apply the survival formula in FilterExec itself, for columns the predicate does not constrain:
NDV_after = NDV * (1 - (1 - selectivity)^(rows / NDV))
capped at the current estimate. This is ndv_after_selectivity in operator_statistics/mod.rs.
Additional context
Evaluating it needs distinct counts in the stats benchmark data, see #25576.
Is your feature request related to a problem or challenge?
FilterExec::statistics_from_inputscaps each column's distinct count at the filtered row count. When values repeat, this overestimates. For example, 1000 rows with 100 distinct values (10 rows per value) filtered to 200 rows givesmin(100, 200) = 100, while the expected number of surviving values is about 89.The deprecated
FilterStatisticsProvider(#25969) implements a better estimate, but only as an opt-in provider.Describe the solution you'd like
Apply the survival formula in
FilterExecitself, for columns the predicate does not constrain:NDV_after = NDV * (1 - (1 - selectivity)^(rows / NDV))capped at the current estimate. This is
ndv_after_selectivityinoperator_statistics/mod.rs.Additional context
Evaluating it needs distinct counts in the stats benchmark data, see #25576.