Skip to content

feat(cli): stream text and tool calls as they complete - #1405

Merged
RemiliaForever (RemiliaForever) merged 3 commits into
mainfrom
feat/toolcall-stream
Aug 31, 2026
Merged

RemiliaForever (RemiliaForever) merged 3 commits into
mainfrom
feat/toolcall-stream

Conversation

@RemiliaForever

@RemiliaForever RemiliaForever (RemiliaForever) commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Closes qcom-ai-hub/geniex#1524 — a request carrying tools buffered the whole generation, so the client saw nothing until it finished. Closes qcom-ai-hub/geniex#1527 — only the first tool call of a turn survived, and a matched call also dropped the prose around it.

Approach

One incremental ToolCallScanner replaces the two implementations each format used to need (a stateless Boundary rescan plus a stateful Watch/Feed pair that had to agree). A format is now just parse + feed, and adding one is a constructor plus a parser.

  • feed(all, from) reports (-1,-1) / (n,-1) / (n,m) in absolute offsets; the earliest format wins, so what a later one finds inside another's region cannot hijack it.
  • Push returns the text safe to emit plus every call that finished, stepping until nothing closes so one chunk can carry several calls.
  • Tail recovers a call the model stopped short of closing, and tries every format because the region holding the tail may contain another's syntax.
  • tool_calls is a slice on both the streaming and blocking paths, with a delta index that rises across chunks.
  • A matched call no longer costs the text around it: content and tool_calls are returned together.
  • reasoning_format now applies to tool-call requests too, streaming included — it used to be ignored, which left the thinking block in content. Defaults are unchanged (none keeps it inline).

Verification

bazel test //cli/cmd/... //cli/internal/... //cli/server/... passes, and both formats were driven end to end against real models on Linux/Vulkan.

Real-model results — --compute hybrid, gemma-4-E2B-it-Q4_0 and Qwen3-4B-Q4_0
scenario result
gemma4 blocking + tools finish_reason=tool_calls, get_weather({"city":"Beijing"})
gemma4 streaming, parallel calls index 0 / 1 at +388ms / +452ms — 64ms apart, not batched at Tail
content kept beside tool_calls Qwen3 inline mode returns both the <think>… text and the call (previously the text was dropped)
reasoning_format=deepseek + tools, blocking content='\n\n', reasoning_content holds the chain of thought, call intact
same, streaming 72 reasoning_content deltas from +199ms, no <think> in any content delta, call at +712ms
plain prose while tools is set 7 content deltas 14-18ms apart, finish_reason=stop
Test coverage
  • TestScannerStream — 30 responses × chunk sizes 1-8: every byte is either emitted as text or consumed by a call, whatever the token boundaries.
  • TestParseMatchesStream / TestStreamAgreesWithJSONParse / TestGemma4StreamAgreesWithParse — the streaming path and the whole-text parsers check each other.
  • TestFeedOffsetsStayAhead — 1005 near-miss marker triples driven the way Push drives them, so no format can report an offset behind from.
  • TestStreamToolCallSeparatesReasoning / TestStreamToolCallIndexes — SSE frames of the handler itself, which had no coverage before.
  • BenchmarkToolCallScanner — prose stays the fast path at ~240 MB/s; every mutation of the new logic was checked to break at least one test.
Known limitations, each pinned by a test
  • markerFormat is not string-aware, so an end marker quoted inside a call's own arguments cuts the region short and the call streams as text. Flagged BUG: in the code; a fix needs a per-format string hook.
  • Tail takes the first format that finds a call, so a truncated turn mixing two syntaxes can lose the other one.
  • gemma4's parse searches for { without a bound, so prose that literally spells <|tool_call>call: can fabricate a call — pre-existing, and now less costly since text before the hold point still streams.
  • thinkfsm matches whole tokens, so a closing think marker that arrives split or fused leaves the call inside reasoning_content.

@RemiliaForever
RemiliaForever (RemiliaForever) force-pushed the feat/toolcall-stream branch 6 times, most recently from f7c150c to 79f19d6 Compare August 28, 2026 13:06
Closes qcom-ai-hub/geniex#1524

Signed-off-by: RemiliaForever <remilia@koumakan.cc>
A request with tools skipped reasoning separation entirely, so the thinking
block stayed in content. That was invisible while a matched tool call dropped
content, and became visible once content was kept alongside tool_calls.
Tool-call syntax never appears inside the thinking block, so the scanner does
not need to see it.

Signed-off-by: RemiliaForever <remilia@koumakan.cc>
The three formats are only ever reached through NewToolCallScanner, and
contentChunk was tokenChunk's content case written twice.

Signed-off-by: RemiliaForever <remilia@koumakan.cc>
@mintlify

mintlify Bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated (UTC)
qualcomm-0801e48b 🟢 Ready View Preview Aug 31, 2026, 12:24 PM

💡 Tip: Enable Workflows to automatically generate PRs for you.

@RemiliaForever
RemiliaForever (RemiliaForever) merged commit 47a3648 into main Aug 31, 2026
52 of 54 checks passed
@RemiliaForever
RemiliaForever (RemiliaForever) deleted the feat/toolcall-stream branch August 31, 2026 14:07

This branch was successfully deployed

1 active deployment
staging - docs — f418fb93 Deployed Aug 31, 2026 by mintlify[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants