feat(cli): stream text and tool calls as they complete - #1405
Merged
Merged
Conversation
RemiliaForever (RemiliaForever)
force-pushed
the
feat/toolcall-stream
branch
6 times, most recently
from
August 28, 2026 13:06
f7c150c to
79f19d6
Compare
Closes qcom-ai-hub/geniex#1524
Signed-off-by: RemiliaForever <remilia@koumakan.cc>
A request with tools skipped reasoning separation entirely, so the thinking block stayed in content. That was invisible while a matched tool call dropped content, and became visible once content was kept alongside tool_calls. Tool-call syntax never appears inside the thinking block, so the scanner does not need to see it. Signed-off-by: RemiliaForever <remilia@koumakan.cc>
The three formats are only ever reached through NewToolCallScanner, and contentChunk was tokenChunk's content case written twice. Signed-off-by: RemiliaForever <remilia@koumakan.cc>
RemiliaForever (RemiliaForever)
force-pushed
the
feat/toolcall-stream
branch
from
August 31, 2026 12:22
79f19d6 to
f418fb9
Compare
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Workflows to automatically generate PRs for you. |
RemiliaForever (RemiliaForever)
marked this pull request as ready for review
August 31, 2026 12:22
Mengsheng Wu (mengshengwu)
approved these changes
Aug 31, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Closes qcom-ai-hub/geniex#1524 — a request carrying
toolsbuffered the whole generation, so the client saw nothing until it finished. Closes qcom-ai-hub/geniex#1527 — only the first tool call of a turn survived, and a matched call also dropped the prose around it.Approach
One incremental
ToolCallScannerreplaces the two implementations each format used to need (a statelessBoundaryrescan plus a statefulWatch/Feedpair that had to agree). A format is now justparse+feed, and adding one is a constructor plus a parser.feed(all, from)reports(-1,-1)/(n,-1)/(n,m)in absolute offsets; the earliest format wins, so what a later one finds inside another's region cannot hijack it.Pushreturns the text safe to emit plus every call that finished, stepping until nothing closes so one chunk can carry several calls.Tailrecovers a call the model stopped short of closing, and tries every format because the region holding the tail may contain another's syntax.tool_callsis a slice on both the streaming and blocking paths, with a deltaindexthat rises across chunks.contentandtool_callsare returned together.reasoning_formatnow applies to tool-call requests too, streaming included — it used to be ignored, which left the thinking block incontent. Defaults are unchanged (nonekeeps it inline).Verification
bazel test //cli/cmd/... //cli/internal/... //cli/server/...passes, and both formats were driven end to end against real models on Linux/Vulkan.Real-model results —
--compute hybrid,gemma-4-E2B-it-Q4_0andQwen3-4B-Q4_0finish_reason=tool_calls,get_weather({"city":"Beijing"})0/1at +388ms / +452ms — 64ms apart, not batched atTailtool_calls<think>…text and the call (previously the text was dropped)reasoning_format=deepseek+ tools, blockingcontent='\n\n',reasoning_contentholds the chain of thought, call intactreasoning_contentdeltas from +199ms, no<think>in any content delta, call at +712mstoolsis setfinish_reason=stopTest coverage
TestScannerStream— 30 responses × chunk sizes 1-8: every byte is either emitted as text or consumed by a call, whatever the token boundaries.TestParseMatchesStream/TestStreamAgreesWithJSONParse/TestGemma4StreamAgreesWithParse— the streaming path and the whole-text parsers check each other.TestFeedOffsetsStayAhead— 1005 near-miss marker triples driven the wayPushdrives them, so no format can report an offset behindfrom.TestStreamToolCallSeparatesReasoning/TestStreamToolCallIndexes— SSE frames of the handler itself, which had no coverage before.BenchmarkToolCallScanner— prose stays the fast path at ~240 MB/s; every mutation of the new logic was checked to break at least one test.Known limitations, each pinned by a test
markerFormatis not string-aware, so an end marker quoted inside a call's own arguments cuts the region short and the call streams as text. FlaggedBUG:in the code; a fix needs a per-format string hook.Tailtakes the first format that finds a call, so a truncated turn mixing two syntaxes can lose the other one.parsesearches for{without a bound, so prose that literally spells<|tool_call>call:can fabricate a call — pre-existing, and now less costly since text before the hold point still streams.thinkfsmmatches whole tokens, so a closing think marker that arrives split or fused leaves the call insidereasoning_content.