Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions changelog.d/3207.added.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
- **Codex/ChatGPT-subscription chats keep their reasoning across turns again (#3207).** With
`store=false` the model's reasoning only survives a turn if its encrypted blob is threaded
back, and protoAgent never captured that blob — the streaming path drops the one event that
carries it — so every turn on `model.provider: openai-codex` started its reasoning from
scratch. It is now captured and replayed, and each item is stamped with the endpoint and
account that minted it: a blob is only ever sent back to the issuer that can decrypt it, so
switching a chat's model mid-thread (or signing in under a different ChatGPT account) costs
reasoning continuity from that point rather than failing the turn.
47 changes: 42 additions & 5 deletions docs/adr/0097-native-oauth-subscription-providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,13 +210,50 @@ Hermes's Codex adapter (`agent/codex_responses_adapter.py`) reached the same rul
independently — including the id strip and a session-wide replay kill switch — and its
`_issuer_kind` stamp is the model for the cross-issuer filter listed below.

## Encrypted-reasoning replay, delivered (2026-08-27, #3199 follow-up)

#3199 contained the damage — never send an item the backend can't verify. This wires the
capability the containment was standing in for, and closes the "contained, not delivered"
open item.

**Capture.** langchain-openai's streaming Responses path has no `response.output_item.done`
branch for reasoning (it has one for `compaction`, which carries the same kind of blob), and
the terminal `response.completed` event keeps only `parsed`/usage/`response_metadata`. So the
blob is visible in exactly one event, which the converter drops. `codex_client`
`_install_reasoning_capture` re-emits that event as a content-block delta that merges onto the
reasoning block already in flight, by `index`. The wrapper sits on the shared module-level
converter — there is no instance seam — but is **inert unless a contextvar this module's
client sets is present**, so every other `ChatOpenAI` in the process is untouched.

**`output_version` flipped to `responses/v1`.** `v0` collapses a turn's reasoning into ONE
`additional_kwargs` slot: later items overwrite earlier ones, and streamed fragments of two
different items merge into each other — so it structurally cannot carry per-item blobs. The
block format keeps each item separate and in order, and langchain replays it that way. The
rendering half of the v0 pin was already paid off (every answer site reads `AIMessage.text`,
which yields text blocks only); `text_of` now skips reasoning blocks outright rather than
writing a `_[reasoning]_` placeholder into exports/session memory/chat bundles, which is what
ADR 0021 asks for anyway. `PROTOAGENT_CODEX_OUTPUT_VERSION=v0` is the escape hatch.

**Issuer stamping.** `encrypted_content` is sealed to the endpoint *and account* that minted
it. Each captured item carries `issuer_fingerprint(base_url, account_id)` — a truncated
digest, so a checkpoint never stores a raw account id — and replay drops items stamped with a
different issuer. Unstamped items (checkpointed before this) still replay. This is the guard
that makes per-slot providers, per-tab model override and the fallback chain safe on a shared
thread; without it, the recovery middleware would be firing routinely instead of never.

**Not verified live.** The wire shape is tested end to end against the real converter, the
real merge and the real payload builder, but no turn has been driven against a real ChatGPT
subscription with this on. If the backend objects, `CodexReasoningReplayRecoveryMiddleware`
(#3199) strips the replay state and retries — the thread degrades to stateless continuity
rather than breaking, which is exactly why that half shipped first.

## Open items

- **Encrypted-reasoning replay is contained, not delivered.** Cross-turn reasoning
continuity on `openai-codex` is OFF: capturing the blob needs an `output_item.done`
handler for reasoning items that langchain-openai does not have (worth an upstream
issue). Once captured, replayed items should carry an issuer stamp (endpoint + account)
and be filtered when the current endpoint differs — Hermes's `_classify_responses_issuer`.
- **Encrypted-reasoning replay is unverified against a live subscription.** Capture,
issuer stamping and replay are wired (above) and covered by wire-shape tests, but no turn
has been driven against a real ChatGPT account with it on. Worth an upstream issue too:
langchain-openai should handle `output_item.done` for reasoning items the way it already
does for `compaction`, which would let protoAgent drop its converter wrapper.
- **Claude end-to-end still unproven on a real subscription** — the sign-in URL + PKCE +
refresh are unit-tested and the flow runs, but no Pro/Max approval has been driven here yet
(tool loop, streaming, `cache_control`).
Expand Down
9 changes: 9 additions & 0 deletions docs/reference/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -499,6 +499,15 @@ Any slot that takes a model name — `routing.fallback_models`, `routing.aux_mod

Hold a gateway key and both subscriptions and you can mix all of them at once — Claude for review, Codex for code, the gateway for cheap bulk work — whatever the main brain runs on. The qualified form is the one to reach for when two providers could plausibly serve the same model id.

::: tip Mixing providers mid-conversation costs reasoning continuity, not correctness
On `openai-codex`, the model's reasoning is threaded across turns as an encrypted blob that
is **sealed to the endpoint and account that minted it** — a blob replayed anywhere else is a
hard `400`. protoAgent stamps each captured item with its issuer and silently drops the ones
the current endpoint can't decrypt, so switching a chat's model mid-thread (or re-signing-in
under a different ChatGPT account) just restarts reasoning continuity from that point. The
conversation itself is unaffected.
:::

```yaml
model:
provider: anthropic-oauth # main brain on your Claude subscription
Expand Down
7 changes: 7 additions & 0 deletions docs/reference/environment-variables.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,13 @@ Every env var the template reads at runtime.
| `PROTOAGENT_MODEL` | (unset) | Overrides `model.name` on every config load — used by `evals/sweep.py` to run one agent against many models without editing YAML. |
| `PROTOAGENT_INSTANCE` | (unset) | Opt-in data-scoping key (ADR 0004): namespaces the knowledge/notes/tasks/checkpoint stores so several agents share a backend without colliding. Seeded from `instance.id` in config. |

## Native OAuth subscription providers (ADR 0097)

| Variable | Default | What |
|---|---|---|
| `PROTOAGENT_CODEX_BASE_URL` | `https://chatgpt.com/backend-api/codex` | Endpoint for `model.provider: openai-codex`. Changing it changes the **issuer** a captured reasoning blob is sealed to, so items minted against the old endpoint stop being replayed (by design — the new one can't decrypt them). |
| `PROTOAGENT_CODEX_OUTPUT_VERSION` | `responses/v1` | Escape hatch for the `openai-codex` content shape. `responses/v1` keeps each reasoning item as its own content block, which is what makes cross-turn encrypted-reasoning replay possible. Set to `v0` to fall back to the legacy string-content shape — the turn still works, but reasoning continuity across turns is off. |

## Deployment / UI tier (ADR 0010)

| Variable | Default | What |
Expand Down
8 changes: 8 additions & 0 deletions graph/message_blocks.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,14 @@ def text_of(message) -> str:
elif isinstance(block, dict):
if block.get("type") == "text" and block.get("text"):
parts.append(str(block["text"]))
elif block.get("type") == "reasoning":
# Skipped outright, not placeholdered. ADR 0021 says reasoning is
# never persisted, and every caller here WRITES what it returns
# (exports, session memory, chat bundles). A `_[reasoning]_` marker
# would be noise in all three — and on the Responses providers
# (openai-codex, ADR 0097) reasoning is a block on EVERY assistant
# turn, so it would be noise on every line.
continue
elif block.get("type"):
parts.append(f"_[{block['type']}]_")
return "\n\n".join(parts)
Expand Down
Loading
Loading