Skip to content

feat(cache): stable per-project cache key + deterministic prefix for provider cache hits - #212

Merged
ltmoerdani merged 2 commits into
ltmoerdani:mainfrom
lecommander:feat/cache-parity-and-prefix-stabilization
Sep 11, 2026
Merged

ltmoerdani merged 2 commits into
ltmoerdani:mainfrom
lecommander:feat/cache-parity-and-prefix-stabilization

Conversation

@lecommander

@lecommander lecommander commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes two layers of cache misses between the opencode CLI and the VS Code extension.

Problem

  1. Routing misses: The CLI loads plugins from ~/.config/opencode/plugins/ (e.g. opencode-context-cache.mjs) on every chat.params call. The VS Code extension implements its own transport to opencode.ai/zen and never runs the CLI, so no plugin ever loads in VS Code. CLI and VS Code sessions used different affinity keys.

  2. Prefix misses: Even with a stable routing key, the provider compares prefix bytes. Tool registration order and JSON key order were non-deterministic, invalidating the prefix cache on every turn.

Changes

Layer 1 — Cache key parity (routing):

  • src/request/headers.ts: Replicate CLI plugin resolution (OPENCODE_PROMPT_CACHE_KEY -> OPENCODE_STICKY_SESSION_ID -> sha256(user@host:workspaceDir)). Sets x-session-id/conversation_id/session_id headers. Logs to context-cache-vscode.log when OPENCODE_CONTEXT_CACHE_DEBUG=1. Normalizes drive letter c:->C: and separators \->/ on Windows for hash parity. Extracts pure hashRawCacheKey(raw) for testability.

  • src/request/openai.ts: Include prompt_cache_key/promptCacheKey in chat-completions + responses bodies.

  • src/responsesRequest.ts: Add promptCacheKey field and passthrough to prompt_cache_key in envelope (was whitelisted away).

Layer 2 — Prefix stabilization (provider cache):

  • src/request/openai.ts / anthropic.ts / google.ts: Sort tools by name with plain code-unit comparison (a<b?-1:a>b?1:0) before mapping (VS Code registration order varies across reloads; localeCompare is locale/ICU-dependent).

  • src/request/schema.ts: Sort Object.entries keys with plain comparison for deterministic JSON prefix bytes.

  • src/provider/chatPrep.ts: Log [prefix-hash] per request (first 3 messages + sorted tool names), gated on OPENCODE_CONTEXT_CACHE_DEBUG, static crypto import, failures via deps.log.

Verifiability (added in fixup 94d93cd):

  • docs/references/opencode-context-cache-reference.md: Pins CLI plugin precedence + normalization + pinned hash, so parity is verifiable without local plugin file.
  • src/test/headers.test.ts: Pins hashRawCacheKey("testuser@testhost:C:/project") == 4f77e704edc190c1872f5ac84f42320d082a4cf4c52bfdacd2634db07de24120.

Verification

  • Routing: 356 requests, single hash f3410589..d63, byte-identical to CLI plugin for same workspace (after separator normalization).
  • Prefix: f94cacd57f2e stable across 10 turns with growing history (454->472 messages).
  • npm run compile / npm test (446 pass) / npm run lint all pass.

Parity scope

Parity holds when CLI is run from the first workspace root (process.cwd() === workspaceFolders[0].fsPath after normalization). Multi-root workspaces or CLI invoked from a subfolder will diverge — expected limitation, documented in reference doc.

Notes

  • prompt_cache_key only on chat-completions/responses (invalid on Anthropic/Google — headers cover those).
  • OPENCODE_CONTEXT_CACHE_DEBUG=1 requires full VS Code quit+relaunch (main process must inherit it, not just Reload Window).
  • Drive-letter + separator normalization is Windows-specific; no-op on other platforms.
  • x-session-id/conversation_id/session_id mirror CLI plugin headers; gateway primary is x-opencode-session -> x-session-affinity. Happy to drop the three if you confirm they are no-ops — no blocker.

@ltmoerdani

Copy link
Copy Markdown
Owner

Hey @lecommander, thanks for this. I looked into the caching docs (OpenAI + Anthropic) and tested a few things locally before reviewing.

The two-layer framing makes sense, and I verified your reasoning against the official specs: prompt_cache_key is supported on both chat-completions and responses, tool/schema ordering affecting the cached prefix is explicitly documented on both OpenAI and Anthropic sides, and Anthropic even calls out unstable key ordering breaking caches. So the core idea checks out.

A few things before we merge:

1. localeCompare isn't spec-guaranteed deterministic

In schema.ts and the three tool mappers, localeCompare() resolves the collator from the user's environment locale (I confirmed LANG=tr_TR.UTF-8 gives a resolved tr-TR collator in Node). For plain ASCII tool names it happens to sort the same across common locales on current ICU, but ICU/CLDR updates can change tailoring between Electron/Node releases, and nothing in the ECMAScript spec pins the output. Since the whole point is deterministic bytes, can we use the plain comparison? It's guaranteed by spec and slightly cheaper:

.sort((a, b) => (a.name < b.name ? -1 : a.name > b.name ? 1 : 0))

Same for the Object.entries sort in schema.ts.

2. Cache-key parity with the CLI plugin

The plugin (opencode-context-cache.mjs) isn't in this repo or upstream sst/opencode, it lives in your local ~/.config/opencode/plugins/. I can't check the byte-identical claim from the code alone, then. Two specific gaps I spotted:

  • VS Code uses workspaceFolders[0].fsPath, CLI uses its cwd. Multi-root workspaces or running the CLI from a subfolder will diverge.
  • On Windows, fsPath uses backslashes while the CLI likely hashes forward slashes. The PR normalizes the drive letter but not the separator.

Could you do one of these so the parity claim is verifiable for everyone, not just your setup?

  • Include the plugin file (or at least its key-resolution function) in this PR, maybe under docs/references/, and/or
  • Add a small unit test pinning the hash for a fixed input so future refactors can't silently break it.

And a line in the PR body about when parity holds (CLI run from the first workspace root) would help.

3. Minor

  • await import("crypto") inside prepareChatRequest runs per request. A static import at the top is simpler.
  • The [prefix-hash] block computes the hash on every request even when nobody reads the log. Fine to keep, but maybe gate it on the same debug env var? Also, that catch block swallows everything silently. A deps.log in there would match our no-swallowed-errors convention.

One question: do you know if the Zen gateway actually reads x-session-id / conversation_id / session_id? The existing code comment says it reads x-opencode-session first, and none of those three appear in any public API docs I could find. If they're no-ops, we could drop them and keep the surface smaller. No blocker either way, just curious what your 356-request run showed.

Happy to merge once 1 and 2 are in. Nice find on the responses envelope dropping prompt_cache_key in the whitelist, btw. That one would have been painful to track down.

@lecommander

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review — all points addressed in fixup 94d93cd (pushed):

1. localeCompare -> plain comparison — done

Replaced all 4 sites (openai.ts x2, anthropic.ts, google.ts, schema.ts) with (a < b ? -1 : a > b ? 1 : 0) — spec-guaranteed code-unit comparison, not locale/ICU-dependent. Tool names are ASCII [a-zA-Z0-9_-] so this is sufficient and cheaper.

2. Cache-key parity verifiability — done

  • Added docs/references/opencode-context-cache-reference.md pinning the CLI plugin's CacheKeyResolver precedence (OPENCODE_PROMPT_CACHE_KEY -> OPENCODE_STICKY_SESSION_ID -> user@host:dir -> headers -> sessionID), hashing (SHA256(raw)), and normalization — so parity is verifiable without the local plugin file.
  • Added src/test/headers.test.ts pinning hashRawCacheKey("testuser@testhost:C:/project") == 4f77e704edc190c1872f5ac84f42320d082a4cf4c52bfdacd2634db07de24120 — future refactors cannot silently break the hash.
  • Extracted pure hashRawCacheKey(raw) + normalizeDirForCacheKey(dir) (\->/, c:->C:) so C:\a\b and C:/a/b hash identically. You were right that fsPath vs process.cwd() both use \ on Windows today (verified byte-identical f3410589..d63 over 356 requests), but canonicalizing is more robust.
  • Documented parity scope in PR body + reference doc: parity holds when CLI is run from first workspace root; multi-root or CLI from subfolder will diverge (expected limitation).

3. Minor — done

  • await import("crypto") -> static import { createHash } from "crypto" at top of chatPrep.ts.
  • [prefix-hash] now gated on OPENCODE_CONTEXT_CACHE_DEBUG (1/true), tool names sorted deterministically, catch logs via deps.log instead of swallowing.

Question on x-session-id/conversation_id/session_id:

They mirror SESSION_ID_HEADER_NAMES from the CLI plugin — we kept parity. Gateway primary is x-opencode-session -> x-session-affinity (as the existing comment notes). No public Zen docs list the three; gateway accepted them in the 356-request run (no 400, hash logged consistently) but we have no gateway-side hit-rate log proving they improve cache vs x-opencode-session + prompt_cache_key alone. prompt_cache_key is the documented prefix-cache mechanism for chat-completions/responses; headers are the routing affinity for Anthropic/Google where prompt_cache_key is invalid. Happy to drop the three if you confirm they are no-ops and keep surface to x-opencode-session + prompt_cache_key — no blocker either way.

npm run compile / npm test (446 pass) / npm run lint all pass.

@ltmoerdani

Copy link
Copy Markdown
Owner

Nice work on the fixup @lecommander, thanks for moving fast on all three points. The reference doc plus the pinned hash test is exactly what I was hoping for, that closes the verifiability gap.

On the three headers: let's drop them. Your own run shows the honest answer, gateway accepted them but you have no hit-rate evidence they do anything beyond what x-opencode-session + prompt_cache_key already cover. Smaller surface, and if the Zen gateway ever documents these headers for real, adding them back is a tiny change. Please also update docs/references/opencode-context-cache-reference.md so the documented header set matches the code.

On CI: the failing check isn't your code, the editorconfig step failed downloading its binary from GitHub Releases (HTTP 500 + ECONNRESET, transient infra flake). Everything real passed: ESLint, Markdown, Prettier, Shell, TypeScript, tests. I'll re-run it on my side.

Once that's green and the headers are dropped, I'm good to merge.

…provider cache hits

- headers.ts: replicate CLI plugin cache-key resolution (env override ->
  sha256(user@host:workspaceDir)) so CLI and VS Code share the same
  affinity bucket. Sets x-session-id/conversation_id/session_id headers
  and logs to context-cache-vscode.log when OPENCODE_CONTEXT_CACHE_DEBUG=1.
  Normalizes drive letter c:->C: on Windows for hash parity.

- openai.ts: sort tools by name before mapping (VS Code registration order
  varies across reloads) and include prompt_cache_key / promptCacheKey in
  chat-completions + responses bodies.

- anthropic.ts / google.ts: same deterministic tool sort.

- schema.ts: sort Object.entries keys for deterministic JSON prefix bytes.

- responsesRequest.ts: add promptCacheKey field and passthrough to
  prompt_cache_key in envelope (was whitelisted away).

- chatPrep.ts: log [prefix-hash] per request (first 3 messages + sorted
  tool names) to measure prefix stability directly.

Verified: 356 requests single hash f3410589..d63 byte-identical to CLI
plugin; prefix hash f94cacd57f2e stable across 10 turns with growing
history.
…ebug-gated prefix hash

- Replace localeCompare with plain code-unit comparison (a<b) in
  openai/anthropic/google tool mappers and schema key sort — spec-
  guaranteed deterministic, not locale/ICU-dependent.
- Extract hashRawCacheKey + normalizeDirForCacheKey (backslash->slash
  + drive-letter upper-case) so C:\a\b and C:/a/b hash identically;
  add docs/references/opencode-context-cache-reference.md pinning the
  CLI plugin precedence and a pinned SHA256 regression test
  (src/test/headers.test.ts).
- Gate [prefix-hash] on OPENCODE_CONTEXT_CACHE_DEBUG, use static
  crypto import, sort tool names deterministically, log failures via
  deps.log instead of swallowing.

Parity holds when CLI is run from first workspace root; multi-root or
CLI from subfolder will diverge (documented).
@lecommander
lecommander force-pushed the feat/cache-parity-and-prefix-stabilization branch from 94d93cd to 80b4a3d Compare September 11, 2026 02:55
@lecommander

Copy link
Copy Markdown
Contributor Author

Rebased onto latest main and addressed the follow-up:

Dropped the three headers — done

  • src/request/headers.ts: SESSION_HEADER_NAMES (x-session-id / conversation_id / session_id) removed. buildOpenCodeRequestHeaders now sends only x-opencode-session + x-opencode-request + x-opencode-client + User-Agent. The project cache key flows via prompt_cache_key (chat-completions/responses bodies) only. Log line updated (no headers= field).
  • docs/references/opencode-context-cache-reference.md: already matches — your 7ed74d2 took the doc with the dropped-headers note, kept as-is through the rebase.

Merge conflicts resolved

npm run lint — Passed (all 7 steps). npm test result in next push check; npm run compile covered by lint's TypeScript step.

@ltmoerdani

Copy link
Copy Markdown
Owner

Checked the rebase, this looks clean. Header drop is exactly what we agreed, the conflict resolution kept the upstream fixes (#220 output channel, auxiliarySessionId, #206/#216) intact while layering the cache-key work on top, which is the right call. CI is green on my side too, all checks passing.

Nothing left from the review. Merging now with a merge commit so your branch history stays intact. Thanks for the patience through the two review rounds, and for pinning the hash test, that doc + test combo is the part I'd point other contributors to as a template.

@ltmoerdani
ltmoerdani merged commit b09d0b0 into ltmoerdani:main Sep 11, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants