Conversation
…eration status endpoint (#2037) PR #2013 added a durable progress snapshot (OperationProgress: stage/at/ processed/total/detail) plus an updated_at heartbeat and an include_payload query param yielding task_payload to GET .../operations/{operation_id}, but the 'Get operation status' docs had no response-field prose for any of them (the example even passes include_payload without explaining it). Added a response-fields subsection sourced from http.py. Regenerated the skills mirror.
… table (#2036) FireworksLLM overrides supports_batch_api()->True (fireworks_llm.py:106), and provider=="fireworks" dispatches to FireworksLLM (llm_wrapper.py:424), but the base OpenAICompatibleLLM grants batch only to openai/groq (openai_compatible_llm.py:1236) so the override is load-bearing. The capabilities matrix in llmProviders.json was missing the fireworks batchApi flag, rendering it as '-' (not supported) and understating the provider. Regenerated the CI-enforced skills mirror (models.md).
…aliases (#2031) * docs(configuration): document shared HINDSIGHT_API_COHERE_API_KEY / LITELLM_API_BASE / LITELLM_API_KEY fallback aliases * docs(configuration): document shared HINDSIGHT_API_COHERE_API_KEY / LITELLM_API_BASE / LITELLM_API_KEY fallback aliases
…onfig.py (#2030) * docs(models): sync gemini + vertexai default models to 3.x matching config.py * docs(models): regenerate skills mirror default-model table (gemini+vertexai 3.x) * docs(models): sync Vertex AI walkthrough + gemini examples to 3.x (complete #2030 scope) The defaults table fix (#2030) left the env-var examples and Vertex AI setup walkthrough still handing users the retired gemini-2.0-flash-001 (404 on Vertex) and stale gemini-2.0-flash. Sync the prose surface: - Vertex AI examples + google/ prefix note -> gemini-3.1-flash-lite (vertexai default) - Gemini AI Studio example -> gemini-3.5-flash (gemini default) Regenerated the CI-enforced skills mirror.
The prompt template used **what**, **when** etc. as field labels. This Markdown bold syntax leaked into LLM outputs causing non-JSON responses across all tested models (GPT-4, Ollama models: gemma4, kimi-k2, llama3.2, qwen3.5, glm-5.1). Replaced **field** with "field" — same visual emphasis for the model but no Markdown syntax to confuse JSON output parsing. Fixes #1138 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…2027) --- updated-dependencies: - dependency-name: pyarrow dependency-version: 23.0.1 dependency-type: indirect dependency-group: uv - dependency-name: aiohttp dependency-version: 3.14.0 dependency-type: indirect dependency-group: uv - dependency-name: idna dependency-version: '3.15' dependency-type: indirect dependency-group: uv - dependency-name: starlette dependency-version: 1.0.1 dependency-type: indirect dependency-group: uv - dependency-name: starlette dependency-version: 1.0.1 dependency-type: indirect dependency-group: uv - dependency-name: urllib3 dependency-version: 2.7.0 dependency-type: indirect dependency-group: uv - dependency-name: python-multipart dependency-version: 0.0.27 dependency-type: indirect dependency-group: uv - dependency-name: aiohttp dependency-version: 3.14.0 dependency-type: indirect dependency-group: uv - dependency-name: idna dependency-version: '3.15' dependency-type: indirect dependency-group: uv ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
* fix(retain): expose retain outcome metadata * fix(retain): avoid double-counting batch extraction errors; drop dup json parse - _write_batch_extraction_errors overwrites extraction_errors_* instead of folding in stored counters, which double-counted on batch crash recovery (resumed batch reprocesses all results and recomputes errors from scratch). - Remove now-unused _parse_result_metadata helper and merge_errors method. - Log retain-outcome-metadata write failures at warning (not debug): a missing write silently regresses clients to the ambiguous pre-fix behaviour. --------- Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
…sql|oracle) (#2024) * docs(configuration): document HINDSIGHT_API_DATABASE_BACKEND (postgresql|oracle) * docs(configuration): document HINDSIGHT_API_DATABASE_BACKEND (postgresql|oracle)
…2039) * feat(recall): make semantic threshold configurable * refactor(recall): rename semantic_threshold to semantic_min_similarity Align the new semantic gate with its sibling BM25_MIN_SCORE: per-strategy prefix, and 'min_similarity' since the value is a cosine similarity. Renames the env var (HINDSIGHT_API_SEMANTIC_MIN_SIMILARITY), config field, and the build_semantic_arm parameter (min_similarity). --------- Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
…maintenance loop (#1969) (#2019) * feat(consolidation): periodic reconcile + cross-tenant retention via maintenance loop (#1969) Add a single background MaintenanceLoop (engine/maintenance.py) started in MemoryEngine.initialize(), replacing the two per-recorder retention sweep tasks. One ~60s tick runs each job on its own interval: - Consolidation reconcile (HINDSIGHT_API_CONSOLIDATION_RECONCILE_INTERVAL_SECONDS, default 300, 0=off): re-schedules consolidation for banks with eligible-but- unscheduled facts and no in-flight consolidation, recovering facts stranded when a consolidation operation failed terminally (#1969). - Retention sweeps (hourly) for audit_log and llm_requests, now across ALL tenant schemas (the old sweeps only swept the base schema). Cross-tenant discovery uses server-side PL/pgSQL routines (migration e5f6a7b8c9d0): public.banks_needing_consolidation() and public.schemas_with_expired_rows(table, ts_col, days) — one round-trip each instead of a per-schema query storm at scale. Config gating resolves the full hierarchy per returned bank (global/tenant/bank); Tenant gains an optional tenant_id so tenant-layer overrides are honored. * fix(consolidation): gate maintenance loop to PostgreSQL The retention sweeps target PG-only tables and the reconcile relies on PG-only PL/pgSQL routines, so on Oracle every tick would call non-existent functions and spam warnings. Skip starting the loop when the backend is Oracle (mirrors the PG-only migration). * test(consolidation): 100-tenant maintenance loop targeting test Provisions 100 tenant schemas (cloning the five tables the loop touches) and verifies each job affects only the tenants it should: audit-log and llm-request retention purge expired rows only in schemas that have them (recent rows kept everywhere), and the consolidation reconcile enqueues only the eligible banks into their own schema — skipping auto-consolidation-disabled, in-flight, and already-consolidated banks. * fix(migration): chain maintenance routines after the split-history head After rebasing onto main, the maintenance-routines migration and #2007's split-history migration (a7b8c9d0e1f2) both pointed at d3e4f5a6b7c8, creating two alembic heads (test_single_head failed). Re-point down_revision to a7b8c9d0e1f2 so the tree is a single linear head again. * fix(maintenance): create public routines once + stop loop racing tests Two CI failures from the maintenance work: 1. Migration ran CREATE OR REPLACE FUNCTION public.* on every per-schema migration; concurrent tenant provisioning collided on the pg_proc catalog ('tuple concurrently updated'). Create the shared public routines only on the base-schema run (target_schema unset); tenant runs skip them. 2. The maintenance loop auto-starts in every test engine (llm-trace retention is on by default), and its background sweep deleted llm_requests rows that test_maintenance_multitenant had just inserted. Disable llm-trace retention in the test env too, so with reconcile already off and audit retention off by default no job is enabled and the loop never starts; tests drive it directly.
…int log, surfaced errors (#2047) * fix(opencode): observable logging — config-only debug, resolved-endpoint log, surfaced errors OpenCode users (notably on Windows) could see tool calls register but no memories land, with zero signal as to why: every retain/recall failure was swallowed via debugLog, the resolved API URL/bank was only logged when debug was on, and HINDSIGHT_DEBUG is unreliable to set for OpenCode's plugin runtime. - Add a Logger that routes through OpenCode's server log stream (client.app.log, service=hindsight) — TUI-safe, visible via --print-logs and the OpenCode log files. Falls back to console.error when no client. - error/warn/info are always emitted; debug is gated on config.debug. - Always log the resolved endpoint + bank at init (a common 'memories aren't saving' cause is silently defaulting to Hindsight Cloud). - Surface retain/recall/hook failures as errors instead of swallowing them; hooks still never throw, so OpenCode is not affected. - Drop the HINDSIGHT_DEBUG env override; 'debug' is now a config-only option (opencode.json plugin options or ~/.hindsight/opencode.json). - Tests for the logger; update config tests; document the change. Refs #1758 * style(opencode): prettier-format plugin.test.ts (pre-existing drift) * docs(opencode): document config-only debug + default error/endpoint logging
…2046) The existing recall suites only exercise the temporal retrieval arm incidentally. This adds a dedicated 'recall-temporal' suite that stamps all memories with one event_date and augments every query with a 1-day window on it, so the temporal entry-point scan matches (near-)all rows — the dense-temporal-zone regime from #1958 that #1983 bounded. - _populate_bank gains an optional event_date for the clustered regime - registered in SUITES; runs by default in the daily all-suites job - added to the workflow_dispatch suite choices for manual single runs Results flow to the perf dashboard automatically (publish script keeps the full suites[] array); a matching 'Recall + temporal' page has been added there.
…2045) * test(ci): de-flake TEI parallelism timing + disposition judge reruns Two pre-existing flaky tests that failed unrelated to their subject: - test_tei_cross_encoder::test_parallel_requests asserted absolute elapsed < 0.08s to prove parallelism; CI scheduling jitter pushed it to 0.10s. Widen the simulated latency and assert comfortably below the serial time (max_concurrent_observed > 1 remains the deterministic parallelism proof). - test_quality_integration::test_high_skepticism_response_is_more_hedged_than_low is a judge-evaluated disposition comparison that exhausted its 2 reruns in CI; bump to 3 (matching the heaviest LLM tests). * fix(ci): prettier-format opencode plugin.test.ts (verify-generated-files) CI runs prettier --write across all integrations and found opencode/src/ plugin.test.ts drifted from the shared .prettierrc.json (it was last hand-edited in #2038), failing verify-generated-files on every PR. Apply the formatting the generator expects (collapses a wrapped .toBe(...) to one line).
…works (#2049) * fix(opencode): call OpenCode app.log as a method so logging actually works 0.2.3 routed logs through client.app.log but extracted it to a detached reference (const log = client.app.log; log(...)). OpenCode's app.log is a class method that uses `this` internally, so the detached call threw 'this._client is undefined' — swallowed by the try/catch, and the console fallback was skipped because the reference was truthy. Net effect: 0.2.3 logged nothing in real OpenCode (no resolved-endpoint line, no surfaced errors). - Call app.log as a method on app so `this` is preserved. - On synchronous failure, fall through to the console.error fallback instead of swallowing. - Regression test with a this-dependent app.log (mirrors OpenCode's client). Verified live against OpenCode 1.16.2: 'service=hindsight ... Hindsight plugin initialized' and 'Injected recall context' now appear in the log stream. * chore(opencode): sync package-lock version to 0.2.3 * fix(opencode): make autoRecall independent of session.created ordering (#1758) autoRecall keyed off session.created marking recalledSessions and system.transform consuming it — which silently disabled recall if system.transform fired first (the relative order is an undocumented OpenCode detail that has differed across versions; #1758 item 2). Recall now runs on the first system.transform per session, using recalledSessions purely as a dedup marker for sessions already recalled into. session.created no longer participates. Behaviour is identical on 1.16.2 (where created fires first) but no longer breaks if the order flips. Verified order-independence with unit tests (recall before/after/without session.created) and a built-plugin harness.
Nearly all hs_llm_core flakiness comes from the judge: a single temperature-0 call to the judge model occasionally flips its verdict on borderline phrasing, failing a test whose system output was actually fine. Harden the shared judge (used by ~49 assertions across 24 files) so every judge-based test benefits at once: - When the primary (temp-0) verdict is 'not met', collect N independent higher-temperature second opinions and uphold the failure only if the majority still agrees. Verdicts that pass on the first call return immediately, so passing tests are unchanged in cost and behaviour, and genuine failures (all judges agree) still fail. Tunable via HINDSIGHT_TEST_JUDGE_CONFIRMATIONS / _CONFIRM_TEMPERATURE. - Retry transient judge-call errors (rate limits, 5xx) so judge-infra hiccups don't fail the test under evaluation (HINDSIGHT_TEST_JUDGE_CALL_ATTEMPTS). Also add the standard @pytest.mark.flaky backstop to the mental-model tag-security test, which lacked one.
…llery + sidebars (#2048) * docs(integrations): single source of truth for sidebar + guardrails Make src/data/integrations.json the single source for the Integrations sidebar across every docs version, and add build-time guardrails so it can't drift. - Inject the Integrations sidebar category at render time from integrations.json via a DocRoot/Layout/Sidebar swizzle. Every docs version (current + frozen 0.3-0.7) now shows the same list, and adding one JSON entry is all it takes - no per-version sidebar edits. The sidebar files keep only a positional placeholder category (a link to the gallery), which the swizzle replaces. - check-integrations.mjs, wired into `npm run build`: - forward: fail if a JSON entry has no docs-integrations/<slug> page (the injected sidebar isn't covered by Docusaurus link-checking). - reverse: fail if a released integration tag is missing from the JSON (skips gracefully without tags; excludes private cloudflare-oauth-proxy). - Add the released-but-undocumented integrations to the JSON so the gallery + sidebar show them: claude-agent-sdk and superagent (with new doc pages) and paperclip. - CI: fetch tags (fetch-depth: 0) in the docs build jobs so the reverse check can see them. One name + one icon per integration come straight from the JSON; display order is the JSON array order (manual, most-interesting-first). * docs(code-review): require integrations.json entry + doc page for integrations Add a review rule: every added/released integration must have an entry in hindsight-docs/src/data/integrations.json (single source of truth for the gallery + sidebar) and a docs-integrations/<slug> page, enforced by check-integrations.mjs. Also note the changelog generator keeps its own INTEGRATIONS list that must be updated for releases. * docs(integrations): sidebar on (unversioned) integration pages + alphabetical order - Give the integration doc pages their own sidebar without versioning them: point the unversioned `integrations` plugin at sidebars-integrations.ts, generated from integrations.json (doc items so each page associates with the sidebar and renders it). Previously these pages had sidebarPath: false (no sidebar at all). - Sort integrations alphabetically by name in all three surfaces — the Integrations Hub gallery, the main docs sidebar, and the new integration-page sidebar — via a shared src/lib/integrations.ts helper (gallery + swizzle) and an inline sort in the config-loaded integration sidebar. JSON array order is no longer significant for display. - The swizzle now only fills the main-docs placeholder category, leaving the generated integration-page sidebar untouched. * docs(integrations): replace placeholder/wrong icons with official brand icons Fetch real brand icons from each integration's official site (apple-touch-icon / high-res favicon) and point integrations.json at them, replacing self-generated, generic, or reused placeholders: - New brand icons for claude-agent-sdk, superagent, paperclip, codex, grok-build, ai-sdk, chat, local-mcp, openclaw, langgraph, autogen, opencode, n8n, pipecat, smolagents, dify, strands, outsystems, pydantic-ai, and refreshed many others (litellm, crewai, perplexity, llamaindex, vapi, flowise, hindclaw, agno, hermes, agentcore, google-adk, openai-agents, roo-code, skills, claude-code). - claude-agent-sdk now uses the Claude/Anthropic brand (was reused claude-code icon); context-forge uses the MCP logo (it's an MCP gateway); superagent uses its pyramid logo (was generic package icon); paperclip its paperclip mark. - Kept the existing real marks for nemoclaw (NVIDIA NeMo) and right-agent — no official brand favicon exists for those, and the auto-fetched candidates were wrong (a letter favicon / the repo author's avatar). - Removed 7 now-orphaned icon files. * ci(docs): add explicit integrations check step to build-docs Run scripts/check-integrations.mjs as a named, fail-fast step before the docs build (the build runs it too, but this surfaces it clearly and fails before the slow build). Pure Node, no npm install; uses the tags already fetched via fetch-depth: 0. * ci(docs): trigger build-docs (integrations check) on integration changes Add hindsight-integrations/** to the docs path filter so the integrations single-source check runs on integration-only PRs (which can add/rename an integration without touching hindsight-docs/**).
… tags (#2051) Forum report (related to GH-1558): a user configures an 'application' entity label (map type, tag=True) with multi-value 'id' and 'name' fields, marks up source text with [[Matched Text (name, id)]] notation, and expects a consistent {application:name:X, application:id:Y} pair per tagged element. They observe inconsistent results: often only one half of the pair, sometimes neither, worse when several tags share a chunk. Adds a focused reproduction harness in test_entity_labels.py: - two deterministic tests pinning the map post-processing mechanics (emits the full pair when the LLM returns both fields; faithfully drops half when it doesn't -- there is no backfill, so pairing must come from the model) - one map-config end-to-end test (hs_llm_core): three tags in one chunk with non-canonical surface forms, asserting every element yields a complete pair Finding: on gemini-2.5-flash the map config is robust -- complete pairs across all runs (including denser/larger documents tried during investigation). The reported inconsistency did not reproduce on this model, pointing to model capability / much larger real documents as the likely driver. The harness is parameterized so a weaker model can be plugged in to reproduce.
* Add Gemini service tier config * Format generated Gemini service tier files --------- Co-authored-by: r266-tech <r266-tech@users.noreply.github.com>
…em message (#2266) #1968 moved routing metadata out of the transcript (it no longer prepends a '[context]' system message) and into the retain API context field, but left three agent_end integration assertions on the old shape: - transcript no longer starts with a {role:'system', '[context]...'} entry - message_count reflects the structured turn length without the system pad (1 for a single-user turn, 2 for the last user+assistant turn) Updates index.test.ts's sibling integration tests to match.
…nfig mock (#2265) extract_facts_from_text reads config.retain_structured_chunk_size (passed to chunk_text), but the test's SimpleNamespace mock only set retain_chunk_size, so the test raised AttributeError instead of exercising the quota-defer path. Add the field (None = plain chunking) to fix it.
…ulti-chunk sub-batches (#2269) Ingesting a large single document (~88k chars) dropped most of its body — and any fact past the first slice — when retained. Two bugs, both only triggered when an oversized item is split into sequential sub-batches whose slices each re-chunk into several extraction chunks (the default config: batch tokens 10k → ~30k-char slices, re-chunked at 3k → ~10 chunks/slice): 1. chunk_index offset (sync + async). retain_batch_async advanced the per-document chunk_index cursor by re-chunking item["content"] AFTER the orchestrator had consumed (popped) it. chunk_text("") returns [""] (count 1), so the cursor moved by 1 per sub-batch instead of by the real chunk count; later slices restarted ~1 slot in, colliding chunk_id = {bank}_{doc}_{index} and overwriting earlier chunks via upsert. Fix: count the slice's chunks before handing it to the orchestrator, while content is still present. 2. whole-document recovery skip (async only). All sub-batches of one submitted operation share one operation_id; the first slice stamps the document into result_metadata.facts_committed_document_ids. The crash-recovery fast-path then saw every later slice's document already "committed" and skipped extraction entirely, so only the first slice survived. Fix: only take the whole-document skip when the call starts the document at chunk 0 (chunk_index_offset == 0); a non-zero offset means this call continues a document another sub-batch already started. Per-chunk hash recovery (existing_chunk_hashes) still provides crash-safety for those chunks. The existing #1888 coverage tests use RETAIN_BATCH_TOKENS=100 (a ~300-char budget, under the chunk size) so every slice collapses to ONE chunk, which masks both bugs. New test_subbatch_multichunk_coverage.py sizes the body so each slice fans out to ~6 chunks, with globally-unique tokens (no chunk-hash dedup), and asserts full coverage + contiguous chunk_index + a needle planted in a late slice across BOTH the sync (retain_batch_async) and async (submit_async_retain) paths.
Add agrasandhany integration
…2271) The "computed freshness trend (stable/strengthening/weakening/new/stale)" described across the developer docs maps to code in reflect/observations.py that is unreferenced — not wired into recall, reflect, or the API, and absent from the OpenAPI schema. It is not a surfaced feature, so the docs overstated it. Replace those claims with the freshness behavior that IS shipped: when newer memories haven't been consolidated yet, reflect treats the affected observations as stale and verifies them against raw facts. Touches developer/index, observations, configuration, api/recall, and best-practices, plus the regenerated skills/hindsight-docs mirror. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tain rule) (#2153) MCP-only Zed integration: hindsight-zed init wires the Hindsight MCP server into Zed's settings.json (via mcp-remote) plus a recall/retain rule in AGENTS.md. Validated end-to-end in real Zed.
These new integrations were added to VALID_INTEGRATIONS / CI but not to the generate-changelog registry, so release-integration.sh failed at the changelog step. Add their package names so releases can be cut. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The last entry had a trailing comma, so build-docs' 'Check integrations' step (strict JSON.parse) failed on main and every PR. Drop it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…one stale (#2267) * blog(freshness): add freshness-aware memory post Concept deep-dive on how Hindsight tracks belief currency: the per-observation freshness trend (new/strengthening/stable/weakening/stale, computed from evidence timestamps over 30/90-day windows by density ratio) and the consolidation-lag signal (up_to_date/slightly_stale/stale from pending memories), plus how the reflect loop uses both to verify stale beliefs against raw facts.
…l/retain rule) (#2276) Long-term memory for OpenHands via native Streamable-HTTP MCP: hindsight-openhands init wires the Hindsight MCP server into config.toml + a recall/retain rule in AGENTS.md.
Resolve all 50 fixable critical/high Dependabot alerts across the monorepo. Python (uv.lock): - starlette 1.0.1 -> 1.3.1, python-multipart -> 0.0.32, pyjwt -> 2.13.0, tornado -> 6.5.7, urllib3 -> 2.7.0 across root + integration projects. - cryptography -> 49.0.0 (GHSA-537c-gmf6-5ccf, bundled-OpenSSL OOB read). Lifted the hindsight-api-slim <47 cap: 47/48/49 verified importing and running RSA sign/verify cleanly on linux/arm64 (Docker on Apple Silicon) and native arm64 macOS; the SIGILL of pyca/cryptography#14733 does not reproduce on current tooling (upstream issue closed unconfirmed). - Root and haystack uv.lock pick up uv lockfile revision 3 (the format the rest of the repo's locks and CI's setup-uv@v7 already use). npm: - shell-quote -> 1.8.4 (critical); ws -> 7.5.11 / 8.21.0; vite -> 8.0.16 across root + integrations; embed control-center UI vite ^5 -> ^6.4.3 (build verified); n8n form-data override -> ^4.0.6. - zapier: overrides for form-data, serialize-javascript, tar, tmp, yeoman-environment (dev-only zapier-platform-cli tree); npm audit clean. Not fixed (no safe path): - nltk (llamaindex, pipecat): no patched release exists upstream (<=3.9.4). - pipecat-ai (pipecat): fix needs 1.2.0 but the integration is pinned <1.0 pending a module-restructure migration.
Restore HINDSIGHT_API_MCP_INSTRUCTIONS for HTTP MCP servers. Append the extra guidance only to retain and recall tool descriptions, matching the original local MCP behavior without changing reflect or other management tools.
… metrics (#2253) * Instrument async worker completion path with operation metrics The async worker never emitted hindsight_operation_operations_total / _duration_seconds — record_operation() was only called from the synchronous API layer. In prod, retain/reflect/consolidation run through the async worker, so the Operations dashboard showed no retain activity and there was no Prometheus signal for async throughput, latency or success/failure. Emit operation metrics from the worker on terminal outcomes: - Add MetricsCollector.record_operation_result(): direct (non-context-manager) recording with an explicit success label, for paths that need success control rather than the exception-based record_operation() CM. The CM now delegates to it (no behaviour change, no duplication). - In WorkerPoller._execute_task_inner, record source="worker" with success=true on normal completion and success=false on failure. Deferrals (DeferOperation) and retries (RetryTaskAt) are not terminal and are deliberately not counted. - Normalise the retain operation_type variants (batch_retain, file_convert_retain) onto operation="retain" so worker completions share the API path's series, which the Operations dashboard keys off. This makes async retain visible on the dashboard and gives a Prometheus signal for async worker throughput and success/failure. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Harden worker metric: record outside executor scope + cover defer/retry W1: recording the success metric inside the executor try meant a metrics failure could be caught by the broad except Exception and mark a completed task as failed. Record on terminal outcomes outside the exception scope and guard the call so instrumentation can never flip terminal task state. W2: add no-DB tests for _execute_task_inner asserting completion/failure emit the metric (with success true/false, retain normalised) and that DeferOperation/RetryTaskAt do not. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style: apply ruff format Satisfy verify-generated-files: blank line after _metric_operation_label and single-line record_operation_result test call, per ruff format. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: correct reflect coverage in worker metric comment reflect runs only on the synchronous API path (execute_task has no reflect branch), so operation="reflect" never emits with source="worker". Reword the comment to list retain/consolidation and the other worker task types instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * metrics(worker): scope success label to completion-throughput, not failure-rate Address review: the worker success label infers success from raise/no-raise, but memory_engine.execute_task swallows deterministic failures (file_convert_retain, non-retryable errors) — it marks the op failed and returns normally — so those record success=true. Rather than re-engineer execute_task to thread status back, narrow this metric's documented meaning to a completion-throughput signal and defer authoritative failure visibility to the now-merged hindsight_async_operations{status="failed"} gauge (#1987), which reads each operation's final DB status. - Reword the poller comment: success=false means the task raised to the poller (unexpected / retry-exhausted); deterministic self-handled failures are not captured here — point operators at the failed gauge. - Add test_executor_self_handled_failure_records_success_by_design to lock the intentional behavior so any future change to the inference is deliberate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(worker): fold self-handled-failure case into the completion test The separate test_executor_self_handled_failure_records_success_by_design asserted nothing the completion test didn't: at the poller boundary a self-handled failure is indistinguishable from a clean completion (both return normally), and with the executor mocked there is no real mark-failed / DB status to observe. Remove the duplicate and document the intentional scoping in the renamed test_executor_returning_normally_records_success docstring instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Update version to 0.8.3 in all components - Regenerate OpenAPI spec and client SDKs - Python packages: hindsight-api, hindsight-dev, hindsight-all, hindsight-embed - Python client: hindsight-clients/python - TypeScript client: hindsight-clients/typescript - hindsight-all npm wrapper: hindsight-all-npm - Rust CLI: hindsight-cli - Control Plane: hindsight-control-plane - Helm chart - Sync documentation to version-0.8
* docs: changelog and blog post for v0.8.3 * docs: drop Richer MCP Tools section from 0.8.3 blog
…tests xdist group (#2272) * test(retain): serialize multichunk sub-batch coverage test on worker_tests xdist group test_subbatch_multichunk_coverage.py's async case submits via submit_async_retain, which inserts parent/child rows into async_operations. test_worker.py drives its own WorkerPoller.claim_batch() against the same pool, so on different xdist workers the two files steal each other's pending rows. Add the shared xdist_group("worker_tests") marker (matching test_async_batch_retain.py and the other async-queue tests) so they serialize on one xdist process. Follow-up to #2269. * test(worker): scope claim_batch count assertions to the test's own bank The xdist_group("worker_tests") marker only serializes the tagged async-queue test files among themselves. It cannot stop test_retain.py (not tagged) from scheduling a 'consolidation' async_operation in the public schema while a worker poller test runs — WorkerPoller.claim_batch() scans the whole schema, so that stray op gets claimed and the global 'assert len(claimed) == N' counts it (observed: assert 3 == 2 in test_poller_discovers_tenants_dynamically). Filter claimed tasks to the test's own bank_id before counting, matching the existing 'my_claims' convention already used by ~10 tests in this file. Covers the remaining global-count assertions in the public-schema poller tests; the max_slots cap test and the isolated custom-schema test are unaffected (their global counts are robust by construction).
* fix(docs): use GitHub icon for agrasandhany integration The agrasandhany gallery entry reused the Obsidian logo. Its repo lives on GitHub, so point it at a GitHub mark instead. * fix(docs): mark agrasandhany as community integration It's authored by external contributor yugandhar-maram, not the Hindsight team — switch type official->community and credit the author.
… Agent Framework (#2293) * blog(agent-framework): add Microsoft Agent Framework persistent memory post
Author
kkroo
added a commit
that referenced
this pull request
Jun 18, 2026
sync(hindsight): merge upstream vectorize-io 0.8.3 (resolves #12 conflicts)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.