Skip to content

新增 OrcaReplay:录下 agent 整次运行并离线复现 - #29

Open
xizhuomengcontin wants to merge 2 commits into
modelscope:mainfrom
xizhuomengcontin:orcareplay-repro
Open

xizhuomengcontin wants to merge 2 commits into
modelscope:mainfrom
xizhuomengcontin:orcareplay-repro

Conversation

@xizhuomengcontin

Copy link
Copy Markdown

在 8 📦 复现、发布与归档 的表格里追加一行。

为什么放这一节

这一节已有 Manage AI Research Projects,它做的是「可复现的项目结构 + metadata 审计 + 记录 AI
workflow 资产」。OrcaReplay 做的是它记录的那份资产本身:把 agent 那次运行逐字节留下来,并且能
再跑一遍。

$ orca record generic-openai -- python experiment.py
info recorded run=run_8f21c3 events=41 exit=0

$ orca replay last
info replaying exchanges=6 egress=blocked
info replay.done reused=6/6 exact=6 divergences=0 exit=0

第二条命令里没有任何模型被调用——模型回复从录像里取回,agent 自己的代码、控制流和工具调用照常
真跑一遍。对科研场景的意义很直接:审稿人或同行不需要你的 API key、不花钱、不联网,就能把你那次
「AI 跑出来的结果」原样复现;而不是只拿到一份日志和一句「当时它是这么做的」。

agent 侧不需要任何改动:orca record 把 agent 当子进程启动,只给它改模型服务的 origin,所以不绑定
框架——LangGraph/LangChain、CrewAI、OpenAI Agents SDK、OpenHands、Strands 以及 Claude Code / Codex /
opencode 这些 CLI 都实测过。

两个必须先说清楚的边界

  • egress=blocked 只挡模型侧出网,不是沙箱。 录下的工具调用会真的再执行一次——那次跑过 curl
    或写过数据库的,重放会再动一次。
  • 重放通过 ≠ 确定性结论。 模型并没有被重新提问,是把录下的回复喂回去;「重新跑一次是不是还会
    得到同样结果」是另一个问题,重放答不了。对科研复现来说这条区别很重要,所以写在前面。

Apache-2.0,Node 20+,npm i -g orcareplay。我是维护者,按自荐处理即可。

@VoyagerXvoyagerx

Copy link
Copy Markdown
Collaborator

Thanks for the contribution! This adds OrcaReplay — it byte-records an agent run and replays it without an API key, network, or cost, though recorded tool calls still re-execute.

One request: could you also add the same entry to the English README at https://github.com/modelscope/Awesome-Vibe-Research/blob/main/README_en.md so the two stay in sync? Thanks again!

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@xizhuomengcontin

Copy link
Copy Markdown
Author

Done — README_en.md now carries the same entry, in section 8 (Reproduction, Release & Archiving), right after Manage AI Research Projects, which is the position matching the Chinese README.

The PR is now README.md +1/-0 and README_en.md +1/-0.

The English wording is a translation of the Chinese row rather than a looser rewrite, so the two stay genuinely in sync: same claim about what is recorded (model requests and responses byte for byte, tool calls, shell exit codes, file changes), same claim about replay (no network, no tokens, no API key needed to reproduce), same tool type and the same two links.

One thing I kept from your summary because it is the honest caveat: a replay serves the model replies from the recording, but recorded tool calls still re-execute. That is why the row says the model replies come from the recording rather than claiming the whole run is inert.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants