Skip to content

feat: Only filter records that directly match a kept record - #89

Merged
Pringled merged 12 commits into
mainfrom
fix/complete-duplicate-reporting
Sep 28, 2026
Merged

Pringled merged 12 commits into
mainfrom
fix/complete-duplicate-reporting

Conversation

@Pringled

@Pringled Pringled commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

This PR makes exact and near duplicates follow the same contract: each filtered record references one retained canonical, and selected_with_duplicates provides the complete groups. To make this work, it also reworks the way we consider duplicates by changing to a "direct" deduplication approach. To illustrate the change here's an example:

Suppose we have a dataset with records A, B, C, D, E and threshold=0.9 where the similarities can be represented like this:

  A B C D E
A – 0.92 0.85 0.85 0.64
B 0.92 – 0.92 0.92 0.69
C 0.85 0.92 – 0.85 0.64
D 0.85 0.92 0.85 – 0.92
E 0.64 0.69 0.64 0.92 –

In other words, the following pairs are duplicates:

    C
    |
A — B — D — E

In this case, there are two "clean" ways to deduplicate:

  • direct: this keeps A, C, D and filters B (dupe of A) and E (dupe of D)
  • transitive: this only keeps A and filters B, C, D, E, because they are all connected through a chain of duplicates, even though E is not similar to A at all.

What we on main is a bit different: we keep A and E, and filter B, C, and D. The logic is this (when going through the records in the example) :

  • A: no neighbour seen → kept, seen = {A, B} (since B is a direct duplicate of A)
  • B: neighbour A seen → dropped
  • C: neighbour B seen → dropped, reported as duplicate of B only (removed, and 0.85 to A)
  • D: neighbour B seen → dropped
  • E: its only neighbour D is not in seen (since dropped records never add their neighbours) → kept

So it's kind of an inbetween of direct and transitive deduplication where we follow the chain up to two steps, so a record is dropped if it is a neighbour of a kept record, or a neighbour of one of that kept record's neighbours.

I've also changed the API a bit. Instead of giving duplicates in the format of [(duplicate, score)...] for every filtered, we now just give a duplicate_of and score, see example below. Also, you can still just get the duplicates of any selected item via selected_with_duplicates, the format change is in filtered since each filtered record now lists only the selected record it was removed for, instead of all its neighbours, e.g.:

Before:

result.selected   # ['A']
result.filtered   # [DuplicateRecord(record='B', duplicates=[('A', 0.94), ('C', 0.82)]),
                  #  DuplicateRecord(record='C', duplicates=[('A', 0.97), ('B', 0.82)])]
result.selected_with_duplicates  # [A → [('B', 0.94), ('C', 0.97)]]

After:

result.selected   # ['A']
result.filtered   # [DuplicateRecord(record='B', exact=False, duplicate_of='A', score=0.94),
                  #  DuplicateRecord(record='C', exact=False, duplicate_of='A', score=0.97)]
result.selected_with_duplicates  # [A → [('B', 0.94), ('C', 0.97)]]

result.filtered[0].duplicates    # DeprecationWarning, returns [('A', 0.94)]

Anyway, I think doing direct is both simpler and more honest. This is the least "aggressive" way of filtering, but it also has the nicest guarantees for the final dataset. What I don't like about the current logic is that some records get filtered from the final dataset, but if you then check them against the finalized dataset they are actually not a duplicate of anything in the final dataset. On the benchmarks, this change leads to 1-3% less dropped records on average. I'll rerun them in a followup since I plan on making a few more changes in followup PRs.

@Pringled
Pringled marked this pull request as draft September 25, 2026 14:28
…eighbor limit

Neighbors propose their canonical (in both directions) and each candidate is
verified by direct cosine similarity, so matches stay non-transitive. When a
record's neighbor list is truncated at MAX_NEIGHBORS without a match, it is
compared against all selected records directly.
@codecov

codecov Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
semhash/datamodels.py 100.00% <100.00%> (ø)
semhash/index.py 100.00% <100.00%> (ø)
semhash/records.py 100.00% <ø> (ø)
semhash/semhash.py 100.00% <100.00%> (ø)
semhash/utils.py 100.00% <100.00%> (ø)
tests/test_datamodels.py 100.00% <100.00%> (ø)
tests/test_semhash.py 100.00% <100.00%> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Zero embeddings (e.g. empty text) produced NaN normalized vectors, which could
make the all-selected fallback pick NaN and keep an extra canonical.
Cross-dataset deduplication now verifies ANN neighbors with exact cosine
similarity, like self-deduplication, so scores are exact and zero vectors
(e.g. empty text) are never reported as near duplicates.
@Pringled
Pringled marked this pull request as ready for review September 26, 2026 11:26
@Pringled
Pringled requested a review from stephantul September 26, 2026 11:26
@Pringled Pringled changed the title fix: Use consistent canonical references for exact and near duplicates feat: Only filter records that directly match a kept record Sep 26, 2026
@Pringled
Pringled merged commit 9bdbcd0 into main Sep 28, 2026
5 checks passed
@Pringled
Pringled deleted the fix/complete-duplicate-reporting branch September 28, 2026 06:56
Pringled added a commit that referenced this pull request Sep 28, 2026
…h benchmarks and bump to 0.5.0

Representative selection encoded its top candidates again after ranking. The
ranking now returns the embeddings in ranked order (index vectors in self mode,
the query embeddings in cross mode) and diversification reuses them.

Benchmarks are rerun on a MacBook Pro (Apple M5, 48 GB RAM) to reflect #89 and
#90, and the version is bumped to 0.5.0 for their breaking changes.
Pringled added a commit that referenced this pull request Sep 30, 2026
…h benchmarks and bump to 0.5.0

Representative selection encoded its top candidates again after ranking. The
ranking now returns the embeddings in ranked order (index vectors in self mode,
the query embeddings in cross mode) and diversification reuses them.

Benchmarks are rerun on a MacBook Pro (Apple M5, 48 GB RAM) to reflect #89 and
#90, and the version is bumped to 0.5.0 for their breaking changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants