Skip to content

Modernize OpenAI execution and Claude agent model support - #2

Draft
ComputerByte wants to merge 2 commits into
mainfrom
feat/modern-model-support
Draft

ComputerByte wants to merge 2 commits into
mainfrom
feat/modern-model-support

Conversation

@ComputerByte

@ComputerByte ComputerByte commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Acceptance status: draft / NO-GO (native transport gap remains)

Independent adversarial review verified that all requested modern OpenAI request/routing paths and Claude CLI agent delegation paths—including Claude Opus 5.5, Sonnet 5.5, and Haiku 5.5—are working and verified live. The full initial acceptance criteria included an interactive main-agent native transport: Claude currently executes through the external Claude Code CLI agent integration, and there is no native Anthropic Messages transport for interactive main-agent sessions. That transport remains an external architecture gap.

Claude Opus 5.5 resolution: The version discrepancy between /Users/byte/.npm-global/bin/claude (2.1.296) and /opt/homebrew/bin/claude (2.1.191) was resolved through explicit agent command configuration (config.toml [[agents]]), avoiding hardcoded operator paths in portable application defaults (cli = "claude"). Authenticated Opus 5.5 execution, streaming and tool calling, and live child delegation from GPT-6.1 Sol were verified successfully; Opus 5.5 is enabled in the operator's agent configuration.

Behavior

GPT-6.1 Sol OAuth execution failed with the bundled 0.153 client protocol. This change preserves the exact model ID and uses the verified released 0.162.1 contract only for official Codex OAuth requests/discovery. It adds the seven documented current OpenAI IDs, exact effort serialization including max/none, public versus OAuth limits, and provider-aware picker/catalog behavior. Invalid model/parameter 400 errors terminate instead of entering outer retry loops. The existing Responses-Lite fix remains intact across HTTP and WebSocket serialization.

Configured agent roles preserve their names and exact model overrides. Auto Drive retains unavailable selections and reports diagnostics instead of replacing routes; max survives the editor, decision schema and parser. Claude CLI agents use exact 5.5 IDs and supported effort flags, with older exact selections retained.

Models Verification
gpt-6.1-sol Official docs and current account discovery; real HTTP/WS Lite serialization mocks; authenticated ModelClient, final built CLI and actual child-agent delegation passed
gpt-6-astra, gpt-6-sol, gpt-6-luna Official docs and account discovery; HTTP/WS Lite and public Responses serialization mocked; live inference untested
gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna Official docs and account discovery; HTTP/WS Lite and public Responses serialization mocked; live inference untested
claude-opus-5-5 Official ID/metadata and CLI arguments verified; authenticated modern CLI (2.1.296 via explicit agent command) returned MODEL_OK; streaming + tool calling passed (OPUS_TEST_OK); live child delegation from GPT-6.1 Sol returned OPUS_DELEGATION_OK; enabled in agent configuration
claude-sonnet-5-5 Authenticated exact CLI execution and streaming + Read tool passed; live child delegation from GPT-6.1 Sol returned SONNET_DELEGATION_OK
claude-haiku-5-5 Authenticated exact CLI execution passed; live child delegation from GPT-6.1 Sol returned HAIKU_DELEGATION_OK

Full support matrix, authoritative sources, architecture/file ownership, upstream just-every#615/just-every#618/just-every#621/just-every#625 audit, role configuration and macOS Apple Silicon build/install instructions: docs/model-support.md. The literal audit inventory is docs/model-support-inventory.json.

Validation

  • Required build-fast dev build passed without warnings, including Rust type checking, on Apple Silicon / Rust 1.96.
  • 1,165 unit tests passed: auto-drive 84, common 29, core 808, TUI 232, version 12. Three core tests ignored; the opt-in live test additionally passed when invoked with existing authorized OAuth credentials.
  • Seven integration tests passed across provider discovery/cache isolation, delegation wake and tool selection.
  • Real serialized request regressions cover modern HTTP/WS Lite, public Responses, custom provider isolation, explicit overrides, unsupported models, bounded HTTP/WS 400s and WS 404 transport fallback. Claude subprocess mocks cover actual argv/effort/tool launch behavior.
  • git diff --check passed. rustfmt was not run because AGENTS.md prohibits it.
  • Strict Clippy failed on two unchanged redundant closures in git-tooling/src/ghost_commits.rs:87/92, reproduced on the exact base. It is not reported as green.
  • Existing environment-mutating unit fixtures were isolated with serial execution and the global Git URL rewrite excluded for HTTPS fixtures. JS-specific suites were not run; JS source is unchanged.
  • Public API-key live inference, live OAuth WebSocket/tool round trips, parallel cross-provider handoffs, and context saturation remain unverified.

Base: e7d4a5d
Final commit: ef5d054

Only ComputerByte/Every-code was written. The codex-rs mirror, upstream repository, production credentials/account configurations and installed binaries remain unchanged. No merge or deployment was performed or authorized by this review.

Brent Duarte added 2 commits October 10, 2026 00:00
Preserve exact GPT-6.1 selection, isolate public/OAuth/proxy metadata, serialize current reasoning effort, and bound invalid-request retries. Modernize existing Claude CLI delegation and gate unverified Opus 5.5 by explicit configuration. Document validation and the native interactive Claude acceptance gap.
…mmand

Configure modern Claude CLI explicitly through existing agent command
configuration without embedding operator paths into portable defaults.
Verify authenticated Opus 5.5 execution, streaming and tool use, and
live delegation from GPT-6.1 Sol alongside Sonnet and Haiku 5.5.
Update support matrix, regression tests, and audit status.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant