Skip to content

sync(hindsight): merge upstream vectorize-io 0.8.3 (resolves #12 conflicts) - #13

Merged
kkroo merged 224 commits into
mainfrom
sync/upstream-0.8.3
Jun 18, 2026
Merged

kkroo merged 224 commits into
mainfrom
sync/upstream-0.8.3

Conversation

@kkroo

@kkroo kkroo commented Jun 18, 2026

Copy link
Copy Markdown

Conflict-resolved resync of vectorize-io/hindsight main (0.7.2 → 0.8.3, 100 commits / 1377 files). Supersedes #12 — that PR's head is upstream's own branch (unpushable), so this is the same content on a Blockcast-side branch with conflicts resolved.

Risk: LOW (source-mirror only)

The deploy is version-pinned (install.sh CHART_REF=v0.6.2 + upstream :0.6.2 image), so merging this does not deploy 0.8.3 and runs none of the 11 new alembic migrations. The worker slot model (the #809/#820 values.yaml fix) is byte-identical at 0.8.3 and unaffected.

Conflicts resolved (3)

  • engine/providers/openai_compatible_llm.py → took upstream. Its rate-limit handling (_rate_limit_retry_at + _raise_provider_quota_defer: parses Retry-After and body reset-at timestamps and windows, and defers the task on cooldowns > max_backoff instead of blocking the worker slot) is a superset of the Blockcast-only _retry_after_seconds (header-only, cap-and-sleep). The auto-merge had left both retry paths coexisting — --theirs makes it coherent. (Net behavior change: long rate-limit cooldowns now defer rather than sleep — strictly better under our rpm:2 throttle.)
  • tests/test_liveness_backpressure.py (Blockcast-only) → dropped the now-removed _retry_after_seconds import + test (+ its unused helpers); kept the 2 unrelated tests (health-degrade, retain-truncation). Both pass locally.
  • paperclip/README.md → union: upstream's dynamicBankId/bankId rows + Blockcast's seedBank* fork rows (tagged (Blockcast fork)).
  • paperclip/package-lock.json → took upstream (matches upstream package.json).

⚠️ The upgrade-DEPLOY is a separate, higher-risk change

Actually running 0.6.2→0.8.3 in prod requires: shared-brain-DB backup + planned run of the 11 migrations (widen bank_id→text, split history tables, drop archive embedding col, invalidated_memory_units, maintenance routines), bump install.sh ref + image tag, and review the config.py +272 changes. Do that as its own PR with a maintenance window — not implied by merging this.

🤖 Generated with Claude Code

r266-tech and others added 30 commits June 8, 2026 10:35
…eration status endpoint (vectorize-io#2037)

PR vectorize-io#2013 added a durable progress snapshot (OperationProgress: stage/at/
processed/total/detail) plus an updated_at heartbeat and an include_payload
query param yielding task_payload to GET .../operations/{operation_id}, but
the 'Get operation status' docs had no response-field prose for any of them
(the example even passes include_payload without explaining it). Added a
response-fields subsection sourced from http.py. Regenerated the skills mirror.
…io#2028) (vectorize-io#2038)

OpenCode >=1.16 iterates every plugin-entry export and throws on any
non-function value; the re-exported DEFAULT_HINDSIGHT_API_URL string
bricked plugin load. Drop it from the entry (still exported from
./config) and add a regression test that the entry is function-only.
… table (vectorize-io#2036)

FireworksLLM overrides supports_batch_api()->True (fireworks_llm.py:106),
and provider=="fireworks" dispatches to FireworksLLM (llm_wrapper.py:424),
but the base OpenAICompatibleLLM grants batch only to openai/groq
(openai_compatible_llm.py:1236) so the override is load-bearing. The
capabilities matrix in llmProviders.json was missing the fireworks
batchApi flag, rendering it as '-' (not supported) and understating the
provider. Regenerated the CI-enforced skills mirror (models.md).
…aliases (vectorize-io#2031)

* docs(configuration): document shared HINDSIGHT_API_COHERE_API_KEY / LITELLM_API_BASE / LITELLM_API_KEY fallback aliases

* docs(configuration): document shared HINDSIGHT_API_COHERE_API_KEY / LITELLM_API_BASE / LITELLM_API_KEY fallback aliases
…onfig.py (vectorize-io#2030)

* docs(models): sync gemini + vertexai default models to 3.x matching config.py

* docs(models): regenerate skills mirror default-model table (gemini+vertexai 3.x)

* docs(models): sync Vertex AI walkthrough + gemini examples to 3.x (complete vectorize-io#2030 scope)

The defaults table fix (vectorize-io#2030) left the env-var examples and Vertex AI
setup walkthrough still handing users the retired gemini-2.0-flash-001
(404 on Vertex) and stale gemini-2.0-flash. Sync the prose surface:
- Vertex AI examples + google/ prefix note -> gemini-3.1-flash-lite (vertexai default)
- Gemini AI Studio example -> gemini-3.5-flash (gemini default)
Regenerated the CI-enforced skills mirror.
…torize-io#2029)

The prompt template used **what**, **when** etc. as field labels.
This Markdown bold syntax leaked into LLM outputs causing non-JSON
responses across all tested models (GPT-4, Ollama models: gemma4,
kimi-k2, llama3.2, qwen3.5, glm-5.1).

Replaced **field** with "field" — same visual emphasis for the model
but no Markdown syntax to confuse JSON output parsing.

Fixes vectorize-io#1138

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…ectorize-io#2027)

---
updated-dependencies:
- dependency-name: pyarrow
  dependency-version: 23.0.1
  dependency-type: indirect
  dependency-group: uv
- dependency-name: aiohttp
  dependency-version: 3.14.0
  dependency-type: indirect
  dependency-group: uv
- dependency-name: idna
  dependency-version: '3.15'
  dependency-type: indirect
  dependency-group: uv
- dependency-name: starlette
  dependency-version: 1.0.1
  dependency-type: indirect
  dependency-group: uv
- dependency-name: starlette
  dependency-version: 1.0.1
  dependency-type: indirect
  dependency-group: uv
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: indirect
  dependency-group: uv
- dependency-name: python-multipart
  dependency-version: 0.0.27
  dependency-type: indirect
  dependency-group: uv
- dependency-name: aiohttp
  dependency-version: 3.14.0
  dependency-type: indirect
  dependency-group: uv
- dependency-name: idna
  dependency-version: '3.15'
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
* fix(retain): expose retain outcome metadata

* fix(retain): avoid double-counting batch extraction errors; drop dup json parse

- _write_batch_extraction_errors overwrites extraction_errors_* instead of
  folding in stored counters, which double-counted on batch crash recovery
  (resumed batch reprocesses all results and recomputes errors from scratch).
- Remove now-unused _parse_result_metadata helper and merge_errors method.
- Log retain-outcome-metadata write failures at warning (not debug): a missing
  write silently regresses clients to the ambiguous pre-fix behaviour.

---------

Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
…sql|oracle) (vectorize-io#2024)

* docs(configuration): document HINDSIGHT_API_DATABASE_BACKEND (postgresql|oracle)

* docs(configuration): document HINDSIGHT_API_DATABASE_BACKEND (postgresql|oracle)
…ectorize-io#2039)

* feat(recall): make semantic threshold configurable

* refactor(recall): rename semantic_threshold to semantic_min_similarity

Align the new semantic gate with its sibling BM25_MIN_SCORE: per-strategy
prefix, and 'min_similarity' since the value is a cosine similarity. Renames
the env var (HINDSIGHT_API_SEMANTIC_MIN_SIMILARITY), config field, and the
build_semantic_arm parameter (min_similarity).

---------

Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
…maintenance loop (vectorize-io#1969) (vectorize-io#2019)

* feat(consolidation): periodic reconcile + cross-tenant retention via maintenance loop (vectorize-io#1969)

Add a single background MaintenanceLoop (engine/maintenance.py) started in
MemoryEngine.initialize(), replacing the two per-recorder retention sweep tasks.
One ~60s tick runs each job on its own interval:

- Consolidation reconcile (HINDSIGHT_API_CONSOLIDATION_RECONCILE_INTERVAL_SECONDS,
  default 300, 0=off): re-schedules consolidation for banks with eligible-but-
  unscheduled facts and no in-flight consolidation, recovering facts stranded
  when a consolidation operation failed terminally (vectorize-io#1969).
- Retention sweeps (hourly) for audit_log and llm_requests, now across ALL tenant
  schemas (the old sweeps only swept the base schema).

Cross-tenant discovery uses server-side PL/pgSQL routines (migration
e5f6a7b8c9d0): public.banks_needing_consolidation() and
public.schemas_with_expired_rows(table, ts_col, days) — one round-trip each
instead of a per-schema query storm at scale. Config gating resolves the full
hierarchy per returned bank (global/tenant/bank); Tenant gains an optional
tenant_id so tenant-layer overrides are honored.

* fix(consolidation): gate maintenance loop to PostgreSQL

The retention sweeps target PG-only tables and the reconcile relies on PG-only
PL/pgSQL routines, so on Oracle every tick would call non-existent functions and
spam warnings. Skip starting the loop when the backend is Oracle (mirrors the
PG-only migration).

* test(consolidation): 100-tenant maintenance loop targeting test

Provisions 100 tenant schemas (cloning the five tables the loop touches) and
verifies each job affects only the tenants it should: audit-log and llm-request
retention purge expired rows only in schemas that have them (recent rows kept
everywhere), and the consolidation reconcile enqueues only the eligible banks
into their own schema — skipping auto-consolidation-disabled, in-flight, and
already-consolidated banks.

* fix(migration): chain maintenance routines after the split-history head

After rebasing onto main, the maintenance-routines migration and vectorize-io#2007's
split-history migration (a7b8c9d0e1f2) both pointed at d3e4f5a6b7c8, creating two
alembic heads (test_single_head failed). Re-point down_revision to a7b8c9d0e1f2
so the tree is a single linear head again.

* fix(maintenance): create public routines once + stop loop racing tests

Two CI failures from the maintenance work:

1. Migration ran CREATE OR REPLACE FUNCTION public.* on every per-schema
   migration; concurrent tenant provisioning collided on the pg_proc catalog
   ('tuple concurrently updated'). Create the shared public routines only on the
   base-schema run (target_schema unset); tenant runs skip them.

2. The maintenance loop auto-starts in every test engine (llm-trace retention is
   on by default), and its background sweep deleted llm_requests rows that
   test_maintenance_multitenant had just inserted. Disable llm-trace retention in
   the test env too, so with reconcile already off and audit retention off by
   default no job is enabled and the loop never starts; tests drive it directly.
…int log, surfaced errors (vectorize-io#2047)

* fix(opencode): observable logging — config-only debug, resolved-endpoint log, surfaced errors

OpenCode users (notably on Windows) could see tool calls register but no
memories land, with zero signal as to why: every retain/recall failure was
swallowed via debugLog, the resolved API URL/bank was only logged when debug
was on, and HINDSIGHT_DEBUG is unreliable to set for OpenCode's plugin runtime.

- Add a Logger that routes through OpenCode's server log stream
  (client.app.log, service=hindsight) — TUI-safe, visible via --print-logs and
  the OpenCode log files. Falls back to console.error when no client.
- error/warn/info are always emitted; debug is gated on config.debug.
- Always log the resolved endpoint + bank at init (a common 'memories aren't
  saving' cause is silently defaulting to Hindsight Cloud).
- Surface retain/recall/hook failures as errors instead of swallowing them;
  hooks still never throw, so OpenCode is not affected.
- Drop the HINDSIGHT_DEBUG env override; 'debug' is now a config-only option
  (opencode.json plugin options or ~/.hindsight/opencode.json).
- Tests for the logger; update config tests; document the change.

Refs vectorize-io#1758

* style(opencode): prettier-format plugin.test.ts (pre-existing drift)

* docs(opencode): document config-only debug + default error/endpoint logging
…ectorize-io#2046)

The existing recall suites only exercise the temporal retrieval arm
incidentally. This adds a dedicated 'recall-temporal' suite that stamps
all memories with one event_date and augments every query with a 1-day
window on it, so the temporal entry-point scan matches (near-)all rows —
the dense-temporal-zone regime from vectorize-io#1958 that vectorize-io#1983 bounded.

- _populate_bank gains an optional event_date for the clustered regime
- registered in SUITES; runs by default in the daily all-suites job
- added to the workflow_dispatch suite choices for manual single runs

Results flow to the perf dashboard automatically (publish script keeps
the full suites[] array); a matching 'Recall + temporal' page has been
added there.
…ectorize-io#2045)

* test(ci): de-flake TEI parallelism timing + disposition judge reruns

Two pre-existing flaky tests that failed unrelated to their subject:

- test_tei_cross_encoder::test_parallel_requests asserted absolute elapsed
  < 0.08s to prove parallelism; CI scheduling jitter pushed it to 0.10s.
  Widen the simulated latency and assert comfortably below the serial time
  (max_concurrent_observed > 1 remains the deterministic parallelism proof).

- test_quality_integration::test_high_skepticism_response_is_more_hedged_than_low
  is a judge-evaluated disposition comparison that exhausted its 2 reruns in CI;
  bump to 3 (matching the heaviest LLM tests).

* fix(ci): prettier-format opencode plugin.test.ts (verify-generated-files)

CI runs prettier --write across all integrations and found opencode/src/
plugin.test.ts drifted from the shared .prettierrc.json (it was last hand-edited
in vectorize-io#2038), failing verify-generated-files on every PR. Apply the formatting the
generator expects (collapses a wrapped .toBe(...) to one line).
…works (vectorize-io#2049)

* fix(opencode): call OpenCode app.log as a method so logging actually works

0.2.3 routed logs through client.app.log but extracted it to a detached
reference (const log = client.app.log; log(...)). OpenCode's app.log is a class
method that uses `this` internally, so the detached call threw
'this._client is undefined' — swallowed by the try/catch, and the console
fallback was skipped because the reference was truthy. Net effect: 0.2.3 logged
nothing in real OpenCode (no resolved-endpoint line, no surfaced errors).

- Call app.log as a method on app so `this` is preserved.
- On synchronous failure, fall through to the console.error fallback instead of
  swallowing.
- Regression test with a this-dependent app.log (mirrors OpenCode's client).

Verified live against OpenCode 1.16.2: 'service=hindsight ... Hindsight plugin
initialized' and 'Injected recall context' now appear in the log stream.

* chore(opencode): sync package-lock version to 0.2.3

* fix(opencode): make autoRecall independent of session.created ordering (vectorize-io#1758)

autoRecall keyed off session.created marking recalledSessions and
system.transform consuming it — which silently disabled recall if
system.transform fired first (the relative order is an undocumented OpenCode
detail that has differed across versions; vectorize-io#1758 item 2).

Recall now runs on the first system.transform per session, using
recalledSessions purely as a dedup marker for sessions already recalled into.
session.created no longer participates. Behaviour is identical on 1.16.2 (where
created fires first) but no longer breaks if the order flips.

Verified order-independence with unit tests (recall before/after/without
session.created) and a built-plugin harness.
…orize-io#2050)

Nearly all hs_llm_core flakiness comes from the judge: a single temperature-0
call to the judge model occasionally flips its verdict on borderline phrasing,
failing a test whose system output was actually fine.

Harden the shared judge (used by ~49 assertions across 24 files) so every
judge-based test benefits at once:

- When the primary (temp-0) verdict is 'not met', collect N independent
  higher-temperature second opinions and uphold the failure only if the majority
  still agrees. Verdicts that pass on the first call return immediately, so
  passing tests are unchanged in cost and behaviour, and genuine failures (all
  judges agree) still fail. Tunable via HINDSIGHT_TEST_JUDGE_CONFIRMATIONS /
  _CONFIRM_TEMPERATURE.
- Retry transient judge-call errors (rate limits, 5xx) so judge-infra hiccups
  don't fail the test under evaluation (HINDSIGHT_TEST_JUDGE_CALL_ATTEMPTS).

Also add the standard @pytest.mark.flaky backstop to the mental-model
tag-security test, which lacked one.
…llery + sidebars (vectorize-io#2048)

* docs(integrations): single source of truth for sidebar + guardrails

Make src/data/integrations.json the single source for the Integrations
sidebar across every docs version, and add build-time guardrails so it
can't drift.

- Inject the Integrations sidebar category at render time from
  integrations.json via a DocRoot/Layout/Sidebar swizzle. Every docs
  version (current + frozen 0.3-0.7) now shows the same list, and adding
  one JSON entry is all it takes - no per-version sidebar edits. The
  sidebar files keep only a positional placeholder category (a link to
  the gallery), which the swizzle replaces.
- check-integrations.mjs, wired into `npm run build`:
  - forward: fail if a JSON entry has no docs-integrations/<slug> page
    (the injected sidebar isn't covered by Docusaurus link-checking).
  - reverse: fail if a released integration tag is missing from the JSON
    (skips gracefully without tags; excludes private cloudflare-oauth-proxy).
- Add the released-but-undocumented integrations to the JSON so the
  gallery + sidebar show them: claude-agent-sdk and superagent (with new
  doc pages) and paperclip.
- CI: fetch tags (fetch-depth: 0) in the docs build jobs so the reverse
  check can see them.

One name + one icon per integration come straight from the JSON; display
order is the JSON array order (manual, most-interesting-first).

* docs(code-review): require integrations.json entry + doc page for integrations

Add a review rule: every added/released integration must have an entry in
hindsight-docs/src/data/integrations.json (single source of truth for the
gallery + sidebar) and a docs-integrations/<slug> page, enforced by
check-integrations.mjs. Also note the changelog generator keeps its own
INTEGRATIONS list that must be updated for releases.

* docs(integrations): sidebar on (unversioned) integration pages + alphabetical order

- Give the integration doc pages their own sidebar without versioning them:
  point the unversioned `integrations` plugin at sidebars-integrations.ts,
  generated from integrations.json (doc items so each page associates with the
  sidebar and renders it). Previously these pages had sidebarPath: false (no
  sidebar at all).
- Sort integrations alphabetically by name in all three surfaces — the
  Integrations Hub gallery, the main docs sidebar, and the new integration-page
  sidebar — via a shared src/lib/integrations.ts helper (gallery + swizzle) and
  an inline sort in the config-loaded integration sidebar. JSON array order is
  no longer significant for display.
- The swizzle now only fills the main-docs placeholder category, leaving the
  generated integration-page sidebar untouched.

* docs(integrations): replace placeholder/wrong icons with official brand icons

Fetch real brand icons from each integration's official site (apple-touch-icon
/ high-res favicon) and point integrations.json at them, replacing
self-generated, generic, or reused placeholders:

- New brand icons for claude-agent-sdk, superagent, paperclip, codex, grok-build,
  ai-sdk, chat, local-mcp, openclaw, langgraph, autogen, opencode, n8n, pipecat,
  smolagents, dify, strands, outsystems, pydantic-ai, and refreshed many others
  (litellm, crewai, perplexity, llamaindex, vapi, flowise, hindclaw, agno,
  hermes, agentcore, google-adk, openai-agents, roo-code, skills, claude-code).
- claude-agent-sdk now uses the Claude/Anthropic brand (was reused claude-code
  icon); context-forge uses the MCP logo (it's an MCP gateway); superagent uses
  its pyramid logo (was generic package icon); paperclip its paperclip mark.
- Kept the existing real marks for nemoclaw (NVIDIA NeMo) and right-agent — no
  official brand favicon exists for those, and the auto-fetched candidates were
  wrong (a letter favicon / the repo author's avatar).
- Removed 7 now-orphaned icon files.

* ci(docs): add explicit integrations check step to build-docs

Run scripts/check-integrations.mjs as a named, fail-fast step before the docs
build (the build runs it too, but this surfaces it clearly and fails before the
slow build). Pure Node, no npm install; uses the tags already fetched via
fetch-depth: 0.

* ci(docs): trigger build-docs (integrations check) on integration changes

Add hindsight-integrations/** to the docs path filter so the integrations
single-source check runs on integration-only PRs (which can add/rename an
integration without touching hindsight-docs/**).
… tags (vectorize-io#2051)

Forum report (related to vectorize-ioGH-1558): a user configures an 'application' entity
label (map type, tag=True) with multi-value 'id' and 'name' fields, marks up
source text with [[Matched Text (name, id)]] notation, and expects a consistent
{application:name:X, application:id:Y} pair per tagged element. They observe
inconsistent results: often only one half of the pair, sometimes neither, worse
when several tags share a chunk.

Adds a focused reproduction harness in test_entity_labels.py:
- two deterministic tests pinning the map post-processing mechanics (emits the
  full pair when the LLM returns both fields; faithfully drops half when it
  doesn't -- there is no backfill, so pairing must come from the model)
- one map-config end-to-end test (hs_llm_core): three tags in one chunk with
  non-canonical surface forms, asserting every element yields a complete pair

Finding: on gemini-2.5-flash the map config is robust -- complete pairs across
all runs (including denser/larger documents tried during investigation). The
reported inconsistency did not reproduce on this model, pointing to model
capability / much larger real documents as the likely driver. The harness is
parameterized so a weaker model can be plugged in to reproduce.
nicoloboschi and others added 26 commits June 17, 2026 14:43
…ectorize-io#2260)

Extend local device detection so sentence-transformers can use an Intel
XPU (e.g. Arc A770) when a torch XPU build is loaded, falling back to CPU
otherwise. Split out from vectorize-io#2233.

Co-authored-by: Sveinbjörn Geirsson <raudbjorn@github>
…em message (vectorize-io#2266)

vectorize-io#1968 moved routing metadata out of the transcript (it no longer prepends a
'[context]' system message) and into the retain API context field, but left
three agent_end integration assertions on the old shape:
- transcript no longer starts with a {role:'system', '[context]...'} entry
- message_count reflects the structured turn length without the system pad
  (1 for a single-user turn, 2 for the last user+assistant turn)

Updates index.test.ts's sibling integration tests to match.
…nfig mock (vectorize-io#2265)

extract_facts_from_text reads config.retain_structured_chunk_size (passed
to chunk_text), but the test's SimpleNamespace mock only set
retain_chunk_size, so the test raised AttributeError instead of exercising
the quota-defer path. Add the field (None = plain chunking) to fix it.
…ulti-chunk sub-batches (vectorize-io#2269)

Ingesting a large single document (~88k chars) dropped most of its body — and
any fact past the first slice — when retained. Two bugs, both only triggered when
an oversized item is split into sequential sub-batches whose slices each re-chunk
into several extraction chunks (the default config: batch tokens 10k → ~30k-char
slices, re-chunked at 3k → ~10 chunks/slice):

1. chunk_index offset (sync + async). retain_batch_async advanced the
   per-document chunk_index cursor by re-chunking item["content"] AFTER the
   orchestrator had consumed (popped) it. chunk_text("") returns [""] (count 1),
   so the cursor moved by 1 per sub-batch instead of by the real chunk count;
   later slices restarted ~1 slot in, colliding chunk_id = {bank}_{doc}_{index}
   and overwriting earlier chunks via upsert. Fix: count the slice's chunks
   before handing it to the orchestrator, while content is still present.

2. whole-document recovery skip (async only). All sub-batches of one submitted
   operation share one operation_id; the first slice stamps the document into
   result_metadata.facts_committed_document_ids. The crash-recovery fast-path
   then saw every later slice's document already "committed" and skipped
   extraction entirely, so only the first slice survived. Fix: only take the
   whole-document skip when the call starts the document at chunk 0
   (chunk_index_offset == 0); a non-zero offset means this call continues a
   document another sub-batch already started. Per-chunk hash recovery
   (existing_chunk_hashes) still provides crash-safety for those chunks.

The existing vectorize-io#1888 coverage tests use RETAIN_BATCH_TOKENS=100 (a ~300-char
budget, under the chunk size) so every slice collapses to ONE chunk, which masks
both bugs. New test_subbatch_multichunk_coverage.py sizes the body so each slice
fans out to ~6 chunks, with globally-unique tokens (no chunk-hash dedup), and
asserts full coverage + contiguous chunk_index + a needle planted in a late slice
across BOTH the sync (retain_batch_async) and async (submit_async_retain) paths.
Add agrasandhany integration
…ectorize-io#2271)

The "computed freshness trend (stable/strengthening/weakening/new/stale)"
described across the developer docs maps to code in
reflect/observations.py that is unreferenced — not wired into recall,
reflect, or the API, and absent from the OpenAPI schema. It is not a
surfaced feature, so the docs overstated it.

Replace those claims with the freshness behavior that IS shipped: when
newer memories haven't been consolidated yet, reflect treats the affected
observations as stale and verifies them against raw facts. Touches
developer/index, observations, configuration, api/recall, and
best-practices, plus the regenerated skills/hindsight-docs mirror.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tain rule) (vectorize-io#2153)

MCP-only Zed integration: hindsight-zed init wires the Hindsight MCP server into Zed's settings.json (via mcp-remote) plus a recall/retain rule in AGENTS.md. Validated end-to-end in real Zed.
These new integrations were added to VALID_INTEGRATIONS / CI but not to the
generate-changelog registry, so release-integration.sh failed at the changelog
step. Add their package names so releases can be cut.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The last entry had a trailing comma, so build-docs' 'Check integrations'
step (strict JSON.parse) failed on main and every PR. Drop it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…one stale (vectorize-io#2267)

* blog(freshness): add freshness-aware memory post

Concept deep-dive on how Hindsight tracks belief currency: the per-observation
freshness trend (new/strengthening/stable/weakening/stale, computed from evidence
timestamps over 30/90-day windows by density ratio) and the consolidation-lag
signal (up_to_date/slightly_stale/stale from pending memories), plus how the
reflect loop uses both to verify stale beliefs against raw facts.
…l/retain rule) (vectorize-io#2276)

Long-term memory for OpenHands via native Streamable-HTTP MCP: hindsight-openhands init wires the Hindsight MCP server into config.toml + a recall/retain rule in AGENTS.md.
Resolve all 50 fixable critical/high Dependabot alerts across the monorepo.

Python (uv.lock):
- starlette 1.0.1 -> 1.3.1, python-multipart -> 0.0.32, pyjwt -> 2.13.0,
  tornado -> 6.5.7, urllib3 -> 2.7.0 across root + integration projects.
- cryptography -> 49.0.0 (GHSA-537c-gmf6-5ccf, bundled-OpenSSL OOB read).
  Lifted the hindsight-api-slim <47 cap: 47/48/49 verified importing and
  running RSA sign/verify cleanly on linux/arm64 (Docker on Apple Silicon)
  and native arm64 macOS; the SIGILL of pyca/cryptography#14733 does not
  reproduce on current tooling (upstream issue closed unconfirmed).
- Root and haystack uv.lock pick up uv lockfile revision 3 (the format the
  rest of the repo's locks and CI's setup-uv@v7 already use).

npm:
- shell-quote -> 1.8.4 (critical); ws -> 7.5.11 / 8.21.0; vite -> 8.0.16
  across root + integrations; embed control-center UI vite ^5 -> ^6.4.3
  (build verified); n8n form-data override -> ^4.0.6.
- zapier: overrides for form-data, serialize-javascript, tar, tmp,
  yeoman-environment (dev-only zapier-platform-cli tree); npm audit clean.

Not fixed (no safe path):
- nltk (llamaindex, pipecat): no patched release exists upstream (<=3.9.4).
- pipecat-ai (pipecat): fix needs 1.2.0 but the integration is pinned <1.0
  pending a module-restructure migration.
Restore HINDSIGHT_API_MCP_INSTRUCTIONS for HTTP MCP servers.
Append the extra guidance only to retain and recall tool
descriptions, matching the original local MCP behavior without
changing reflect or other management tools.
… metrics (vectorize-io#2253)

* Instrument async worker completion path with operation metrics

The async worker never emitted hindsight_operation_operations_total /
_duration_seconds — record_operation() was only called from the synchronous
API layer. In prod, retain/reflect/consolidation run through the async worker,
so the Operations dashboard showed no retain activity and there was no
Prometheus signal for async throughput, latency or success/failure.

Emit operation metrics from the worker on terminal outcomes:
- Add MetricsCollector.record_operation_result(): direct (non-context-manager)
  recording with an explicit success label, for paths that need success control
  rather than the exception-based record_operation() CM. The CM now delegates
  to it (no behaviour change, no duplication).
- In WorkerPoller._execute_task_inner, record source="worker" with success=true
  on normal completion and success=false on failure. Deferrals (DeferOperation)
  and retries (RetryTaskAt) are not terminal and are deliberately not counted.
- Normalise the retain operation_type variants (batch_retain,
  file_convert_retain) onto operation="retain" so worker completions share the
  API path's series, which the Operations dashboard keys off.

This makes async retain visible on the dashboard and gives a Prometheus signal
for async worker throughput and success/failure.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Harden worker metric: record outside executor scope + cover defer/retry

W1: recording the success metric inside the executor try meant a metrics
failure could be caught by the broad except Exception and mark a completed
task as failed. Record on terminal outcomes outside the exception scope and
guard the call so instrumentation can never flip terminal task state.

W2: add no-DB tests for _execute_task_inner asserting completion/failure emit
the metric (with success true/false, retain normalised) and that
DeferOperation/RetryTaskAt do not.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style: apply ruff format

Satisfy verify-generated-files: blank line after _metric_operation_label and
single-line record_operation_result test call, per ruff format.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: correct reflect coverage in worker metric comment

reflect runs only on the synchronous API path (execute_task has no reflect
branch), so operation="reflect" never emits with source="worker". Reword the
comment to list retain/consolidation and the other worker task types instead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* metrics(worker): scope success label to completion-throughput, not failure-rate

Address review: the worker success label infers success from raise/no-raise, but
memory_engine.execute_task swallows deterministic failures (file_convert_retain,
non-retryable errors) — it marks the op failed and returns normally — so those
record success=true. Rather than re-engineer execute_task to thread status back,
narrow this metric's documented meaning to a completion-throughput signal and
defer authoritative failure visibility to the now-merged
hindsight_async_operations{status="failed"} gauge (vectorize-io#1987), which reads each
operation's final DB status.

- Reword the poller comment: success=false means the task raised to the poller
  (unexpected / retry-exhausted); deterministic self-handled failures are not
  captured here — point operators at the failed gauge.
- Add test_executor_self_handled_failure_records_success_by_design to lock the
  intentional behavior so any future change to the inference is deliberate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(worker): fold self-handled-failure case into the completion test

The separate test_executor_self_handled_failure_records_success_by_design
asserted nothing the completion test didn't: at the poller boundary a
self-handled failure is indistinguishable from a clean completion (both return
normally), and with the executor mocked there is no real mark-failed / DB status
to observe. Remove the duplicate and document the intentional scoping in the
renamed test_executor_returning_normally_records_success docstring instead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Update version to 0.8.3 in all components
- Regenerate OpenAPI spec and client SDKs
- Python packages: hindsight-api, hindsight-dev, hindsight-all, hindsight-embed
- Python client: hindsight-clients/python
- TypeScript client: hindsight-clients/typescript
- hindsight-all npm wrapper: hindsight-all-npm
- Rust CLI: hindsight-cli
- Control Plane: hindsight-control-plane
- Helm chart
- Sync documentation to version-0.8
* docs: changelog and blog post for v0.8.3

* docs: drop Richer MCP Tools section from 0.8.3 blog
…tests xdist group (vectorize-io#2272)

* test(retain): serialize multichunk sub-batch coverage test on worker_tests xdist group

test_subbatch_multichunk_coverage.py's async case submits via
submit_async_retain, which inserts parent/child rows into async_operations.
test_worker.py drives its own WorkerPoller.claim_batch() against the same pool,
so on different xdist workers the two files steal each other's pending rows.
Add the shared xdist_group("worker_tests") marker (matching
test_async_batch_retain.py and the other async-queue tests) so they serialize
on one xdist process. Follow-up to vectorize-io#2269.

* test(worker): scope claim_batch count assertions to the test's own bank

The xdist_group("worker_tests") marker only serializes the tagged
async-queue test files among themselves. It cannot stop test_retain.py
(not tagged) from scheduling a 'consolidation' async_operation in the
public schema while a worker poller test runs — WorkerPoller.claim_batch()
scans the whole schema, so that stray op gets claimed and the global
'assert len(claimed) == N' counts it (observed: assert 3 == 2 in
test_poller_discovers_tenants_dynamically).

Filter claimed tasks to the test's own bank_id before counting, matching
the existing 'my_claims' convention already used by ~10 tests in this
file. Covers the remaining global-count assertions in the public-schema
poller tests; the max_slots cap test and the isolated custom-schema test
are unaffected (their global counts are robust by construction).
…rize-io#2294)

* fix(docs): use GitHub icon for agrasandhany integration

The agrasandhany gallery entry reused the Obsidian logo. Its repo lives
on GitHub, so point it at a GitHub mark instead.

* fix(docs): mark agrasandhany as community integration

It's authored by external contributor yugandhar-maram, not the Hindsight
team — switch type official->community and credit the author.
… Agent Framework (vectorize-io#2293)

* blog(agent-framework): add Microsoft Agent Framework persistent memory post
Resync of vectorize-io/hindsight main (0.7.2 → 0.8.3, 100 commits/1377 files).
Source-mirror only: the deploy is pinned (install.sh CHART_REF=v0.6.2 + upstream
:0.6.2 image), so this merge does NOT deploy 0.8.3 and runs none of the 11 new
alembic migrations. The worker slot model (vectorize-io#809/vectorize-io#820 values.yaml fix) is
byte-identical at 0.8.3 and unaffected.

3 conflicts resolved:
- engine/providers/openai_compatible_llm.py: took upstream. Upstream's rate-limit
  handling (_rate_limit_retry_at + _raise_provider_quota_defer: parses Retry-After
  AND body reset-at timestamps AND windows, and DEFERS the task on cooldowns
  longer than max_backoff instead of blocking the worker slot) is a superset of
  the Blockcast-only _retry_after_seconds (header-only, cap-and-sleep). The
  auto-merge had left BOTH coexisting (two retry paths); --theirs makes it
  coherent.
- tests/test_liveness_backpressure.py (Blockcast-only): dropped the now-removed
  _retry_after_seconds import + test (and its unused _Response/_StatusError
  helpers); kept the 2 unrelated tests (health-degrade, retain-truncation). Both
  pass.
- paperclip/README.md: union — upstream's dynamicBankId/bankId rows + Blockcast's
  seedBank* fork rows (tagged "(Blockcast fork)").
- paperclip/package-lock.json: took upstream (matches upstream package.json).

Supersedes #12 (whose head is upstream's branch, unpushable). The actual
0.6.2→0.8.3 UPGRADE-DEPLOY is a separate, higher-risk change: backup + planned
run of the 11 migrations (widen bank_id→text, split history tables, drop archive
embedding, invalidated_memory_units, maintenance routines), bump install.sh ref +
image tag, review config.py changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@kkroo kkroo mentioned this pull request Jun 18, 2026
@kkroo
kkroo merged commit d5e7e17 into main Jun 18, 2026
82 of 83 checks passed
@kkroo
kkroo deleted the sync/upstream-0.8.3 branch June 18, 2026 21:23

@allyblockcast allyblockcast Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ally — Consolidated PR Review

Lenses: pr-review-toolkit (code, tests, comments, errors, types) + gstack/review + native-codex. 1,378 changed files — review scoped to the 3 conflict-resolved files, key new upstream modules, and migration SQL safety.


Critical Issues (0)

Important Issues (0)

Suggestions (3)

  • [pr-review-toolkit/comments] PR body — Conflict count mismatch: body says "3 conflicts resolved" but lists 4 items (openai_compatible_llm.py, test_liveness_backpressure.py, paperclip/README.md, paperclip/package-lock.json).

    • Minor — no action required, but worth fixing to prevent confusion when future reviewers count the list against the claim.
  • [pr-review-toolkit/comments] PR body — Path shortening may mislead: paperclip/README.md and paperclip/package-lock.json in the body are actually hindsight-integrations/paperclip/README.md / …/package-lock.json.

    • Cosmetic; the actual files and their conflict resolution are correct. Future maintainers searching for those paths will need to know the full location.
  • [native-codex] hindsight-api-slim/hindsight_api/api/disconnect.py — _should_monitor hardcodes endpoint suffixes /memories/recall and /reflect. Any future long-running endpoint that needs disconnect cancellation won't get it automatically.

    • Add a brief comment flagging this as a maintenance point (# Add new long-running endpoints here), or consider a config set. Not a bug today, just atrophies silently as new endpoints land.

Strengths

  • Conflict resolutions are all correct. openai_compatible_llm.py → took upstream superset (_rate_limit_retry_at + _raise_provider_quota_defer defers on long cooldowns instead of blocking the worker slot — strictly better under rpm:2). test_liveness_backpressure.py → _retry_after_seconds import and test cleanly dropped; both retained tests (test_health_check_degrades_quickly…, test_retain_fact_prompt_truncates…) are correct and unaffected. hindsight-integrations/paperclip/README.md → union merge confirmed correct: upstream dynamicBankId/bankId rows present, Blockcast seedBank* rows present with (Blockcast fork) tags.
  • Migration SQL is injection-safe. b2d4f6a8c1e3's EXECUTE format() uses %I for all identifier interpolations (schema names from pg_namespace, p_table, p_ts_col) and %L for literals. Schema names come from pg_class/pg_namespace (system catalog, not user input). Migration chain is idempotent (IF NOT EXISTS, CREATE OR REPLACE).
  • disconnect.py correctly solves the BaseHTTPMiddleware ASGI receive channel problem. The pump task lifecycle (finally: pump_task.cancel() + suppress(CancelledError)) is correct; cancelling an already-finished task is a no-op, and cancelling mid-await is properly caught.
  • Memory Defense extension is well-designed. _fingerprint_value is length-aware and keeps provider-identifying prefix + tail without leaking raw secrets. Regex extension follows the sensitive_data detector contract. Migration c9a1b2d3e4f5 correctly separates live (memory_units) from cold (invalidated_memory_units) storage — no embedding column in the archive prevents dimension-mismatch on model upgrades.
  • PR description is exemplary for a 1,378-file sync: explicit risk rating, per-conflict rationale, and a clear ⚠️ callout separating "merge this sync" from "deploy 0.8.3" — making the deploy path a deliberate, separate decision rather than an implicit consequence.

Recommended Action

  1. No Critical or Important issues — merge-ready from Ally's perspective.
  2. Address Suggestions opportunistically (conflict count / path precision in the body, maintenance comment in disconnect.py).
  3. Deploy upgrade (0.6.2 → 0.8.3 with 11 Alembic migrations) as a separate PR with maintenance window, per PR body guidance.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.