Skip to content

feat: Reuse ranking embeddings when selecting representatives, refresh benchmarks - #91

Merged
Pringled merged 1 commit into
mainfrom
perf/reuse-representative-embeddings
Sep 30, 2026
Merged

Pringled merged 1 commit into
mainfrom
perf/reuse-representative-embeddings

Conversation

@Pringled

Copy link
Copy Markdown
Member

find_representative and self_find_representative encoded their top candidates a second time, even though those embeddings were already computed while ranking. The ranking now returns them and diversification reuses them. This barely matters with our default model but does have a big effect when you use a slower encoder.

I also reran the benchmarks.

…h benchmarks and bump to 0.5.0

Representative selection encoded its top candidates again after ranking. The
ranking now returns the embeddings in ranked order (index vectors in self mode,
the query embeddings in cross mode) and diversification reuses them.

Benchmarks are rerun on a MacBook Pro (Apple M5, 48 GB RAM) to reflect #89 and
#90, and the version is bumped to 0.5.0 for their breaking changes.
@Pringled
Pringled requested a review from stephantul September 30, 2026 06:43
@codecov

codecov Bot commented Sep 30, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
semhash/semhash.py 100.00% <100.00%> (ø)
tests/test_semhash.py 100.00% <100.00%> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@Pringled
Pringled merged commit 19f9c59 into main Sep 30, 2026
5 checks passed
@Pringled
Pringled deleted the perf/reuse-representative-embeddings branch September 30, 2026 07:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants