feat(pipeline): Jev classifier front-runner for typed-decision calls - #7
Conversation
Ports the Jev integration from argus-private (#301): TypeSafe's System One eval API front-runs five typed-decision surfaces — addressed judge, convention relations, intent verification, scoring FP pre-filter, and an observe-only triage shadow — with the LLM as the fail-safe fallback on any error, timeout, missing answer, or mid-band probability. Credentials resolve per installation: a stored 'typesafe' provider key (repo-level, then org-level; BYOK — the key IS the consent) else the env TYPESAFE_API_KEY gated by the jev_classifier feature flag (default OFF). Settings → Providers gains a TypeSafe Jev card beside Embeddings; providers/settings pages get section landmarks, consistent card states, and a whitelisted LLM provider count.
There was a problem hiding this comment.
6 issues found across 31 files
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="backend/internal/pipeline/triage.go">
<violation number="1" location="backend/internal/pipeline/triage.go:77">
P2: When a BYOK Jev key is configured, `startJevTriageShadow` resolves it synchronously before creating the shadow goroutine. The database lookup therefore blocks the LLM leg despite this path being documented as parallel; resolve the evaluator asynchronously or start both operations concurrently.</violation>
</file>
<file name="backend/internal/pipeline/jev_conventions.go">
<violation number="1" location="backend/internal/pipeline/jev_conventions.go:66">
P1: When stored convention text contains an inline prompt directive, `sanitizeUserInput` leaves it intact because its patterns require the start of a line. Sanitize both convention fields with the untrusted-memory sanitizer, or explicitly instruct Jev to treat the wrapped values as data, before allowing a confident relation to skip the LLM.</violation>
</file>
<file name="backend/internal/pipeline/jev_scoring.go">
<violation number="1" location="backend/internal/pipeline/jev_scoring.go:61">
P2: When Jev marks a finding as dropped, this map eventually removes it from `run.FileReviews`, so `indexComments` never creates its `review_comments` row. Preserve the finding as a suppressed row while excluding it from the GitHub inline output.</violation>
<violation number="2" location="backend/internal/pipeline/jev_scoring.go:77">
P1: A PR-controlled finding description or suggestion can inject instructions into Jev and make a real finding satisfy the drop condition, silently suppressing it. Sanitize and delimiter-wrap each finding before adding it to the classifier state.</violation>
</file>
<file name="backend/internal/jev/jev.go">
<violation number="1" location="backend/internal/jev/jev.go:275">
P2: An HTTP 200 response with negative usage counts passes `json.Unmarshal`, returns no error, and produces negative cost; malformed `answers` is also accepted. Validate the response schema and non-negative usage before logging the call as completed or billing it.</violation>
</file>
<file name="backend/internal/jev/jev_test.go">
<violation number="1" location="backend/internal/jev/jev_test.go:14">
P3: Two documented, safety-relevant client behaviors have no unit tests in this package: (1) the `maxStateBytes` (120KB) fail-fast — `Evaluate` must error on oversized state instead of issuing a doomed round-trip; (2) `Result.Noul`/`Result.Choice` rejecting NaN/±Inf/out-of-range probabilities — `encoding/json` decodes `1e400` to `+Inf` (range errors are swallowed for floats) and `1.7` decodes fine, so these guards are reachable and exist precisely because an out-of-range value must never clear a caller's threshold. Neither path is exercised anywhere in this PR (`rg maxStateBytes internal --glob '*_test.go'` and `rg IsInf internal --glob '*_test.go'` both return nothing; only the 1.7 case is covered indirectly via pipeline tests). Add table tests for the cap and for NaN/±Inf/out-of-range probabilities returning nil / ok=false.</violation>
</file>
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
| // conventions are user-controlled text persisted across reviews. | ||
| func jevConventionState(candidate string, neighbors []memory.PatternMatch) map[string]any { | ||
| state := map[string]any{ | ||
| "candidate_convention": conventionPromptField("candidate_convention", util.Truncate(candidate, jevConventionFieldCap, false)), |
There was a problem hiding this comment.
P1: When stored convention text contains an inline prompt directive, sanitizeUserInput leaves it intact because its patterns require the start of a line. Sanitize both convention fields with the untrusted-memory sanitizer, or explicitly instruct Jev to treat the wrapped values as data, before allowing a confident relation to skip the LLM.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/jev_conventions.go, line 66:
<comment>When stored convention text contains an inline prompt directive, `sanitizeUserInput` leaves it intact because its patterns require the start of a line. Sanitize both convention fields with the untrusted-memory sanitizer, or explicitly instruct Jev to treat the wrapped values as data, before allowing a confident relation to skip the LLM.</comment>
<file context>
@@ -0,0 +1,102 @@
+// conventions are user-controlled text persisted across reviews.
+func jevConventionState(candidate string, neighbors []memory.PatternMatch) map[string]any {
+ state := map[string]any{
+ "candidate_convention": conventionPromptField("candidate_convention", util.Truncate(candidate, jevConventionFieldCap, false)),
+ }
+ for i, n := range neighbors {
</file context>
| findings := make([]string, 0, len(allComments)) | ||
| for i, ic := range allComments { | ||
| c := run.FileReviews[ic.fileIdx].Comments[ic.commentIdx] | ||
| findings = append(findings, scoringFindingText(i, run.FileReviews[ic.fileIdx].Path, c)) |
There was a problem hiding this comment.
P1: A PR-controlled finding description or suggestion can inject instructions into Jev and make a real finding satisfy the drop condition, silently suppressing it. Sanitize and delimiter-wrap each finding before adding it to the classifier state.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/jev_scoring.go, line 77:
<comment>A PR-controlled finding description or suggestion can inject instructions into Jev and make a real finding satisfy the drop condition, silently suppressing it. Sanitize and delimiter-wrap each finding before adding it to the classifier state.</comment>
<file context>
@@ -0,0 +1,131 @@
+ findings := make([]string, 0, len(allComments))
+ for i, ic := range allComments {
+ c := run.FileReviews[ic.fileIdx].Comments[ic.commentIdx]
+ findings = append(findings, scoringFindingText(i, run.FileReviews[ic.fileIdx].Path, c))
+ }
+ var pr strings.Builder
</file context>
|
|
||
| // Shadow Jev eval — fires in parallel with the LLM leg, joined below. | ||
| // Observe-only: its answers are logged, never used for routing. | ||
| shadowCh := ts.startJevTriageShadow(ctx, run) |
There was a problem hiding this comment.
P2: When a BYOK Jev key is configured, startJevTriageShadow resolves it synchronously before creating the shadow goroutine. The database lookup therefore blocks the LLM leg despite this path being documented as parallel; resolve the evaluator asynchronously or start both operations concurrently.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/triage.go, line 77:
<comment>When a BYOK Jev key is configured, `startJevTriageShadow` resolves it synchronously before creating the shadow goroutine. The database lookup therefore blocks the LLM leg despite this path being documented as parallel; resolve the evaluator asynchronously or start both operations concurrently.</comment>
<file context>
@@ -69,6 +72,10 @@ func (ts *TriageStage) Execute(ctx context.Context, run *PipelineRun) (err error
+ // Shadow Jev eval — fires in parallel with the LLM leg, joined below.
+ // Observe-only: its answers are logged, never used for routing.
+ shadowCh := ts.startJevTriageShadow(ctx, run)
+
// Phase 2: LLM refinement — only for manageable file counts
</file context>
Cubic review remediation on the Jev port: - sanitize + delimiter-wrap all untrusted text entering Jev state (filenames, finding descriptions, suggestions, conventions via retrieved-memory sanitizer); escalate when convention text would be truncated rather than classify partial rules - StageTokens.Aux: per-leg spend ledger so mixed Jev+LLM buckets keep per-model attribution; stats/models aggregates aux legs and now covers every scalar stage (intent, auto-resolve, acceptance, cross_pr, reply, simulations) - triage shadow: async BYOK resolution off the caller path, cancel in-flight eval on join-budget expiry, drain raced results, never bill skipped/abandoned calls - jev client: single state marshal, bounded response read, clamp negative usage, oversized-state telemetry - resolver: validate stored base_url (http(s) + host), fall back to default endpoint with a warning - providers page: show repo-scoped typesafe keys read-only, surface delete errors - .env.example: opt-in UPDATE gains WHERE clause Generated with [Devin](https://devin.ai)
… channel Generated with [Devin](https://devin.ai)
There was a problem hiding this comment.
1 existing issue remains and 3 new issues found across 20 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="backend/internal/pipeline/jev_conventions.go">
<violation number="1" location="backend/internal/pipeline/jev_conventions.go:82">
P3: The new oversize-escalation branch has no test coverage, unlike every other escalation path in this file (TestJevConventions_OneUncertainEscalatesWholeBatch, _ErrorEscalates, _InvalidChoiceEscalates). Add a test asserting that an oversized candidate or neighbor returns ok=false, escalates to the LLM with full text, and bills zero Jev spend.</violation>
</file>
<file name="backend/internal/pipeline/jev_resolver.go">
<violation number="1" location="backend/internal/pipeline/jev_resolver.go:82">
P2: A stored endpoint with a query or fragment passes this check, but `jev.Client.evaluate` appends `/v1/systemone` to the raw URL. The request then targets the wrong path and every BYOK evaluation escalates; reject query and fragment components or join URL paths structurally.</violation>
<violation number="2" location="backend/internal/pipeline/jev_resolver.go:82">
P3: `validJevBaseURL` accepts plaintext `http://` endpoints even though this URL receives the TypeSafe API key (Bearer header) and tenant PR/finding content. For the egress of credentials plus data, restrict the accepted scheme to `https` (self-hosted plaintext relays can be terminated behind an explicit TLS proxy instead).</violation>
</file>
Requires human review: Auto-approval blocked because this review re-detected 1 unresolved issue already reported by Cubic.
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| // hosts, paths without scheme, non-http schemes) is rejected. | ||
| func validJevBaseURL(raw string) bool { | ||
| u, err := url.Parse(raw) | ||
| return err == nil && (u.Scheme == "https" || u.Scheme == "http") && u.Host != "" |
There was a problem hiding this comment.
P2: A stored endpoint with a query or fragment passes this check, but jev.Client.evaluate appends /v1/systemone to the raw URL. The request then targets the wrong path and every BYOK evaluation escalates; reject query and fragment components or join URL paths structurally.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/jev_resolver.go, line 82:
<comment>A stored endpoint with a query or fragment passes this check, but `jev.Client.evaluate` appends `/v1/systemone` to the raw URL. The request then targets the wrong path and every BYOK evaluation escalates; reject query and fragment components or join URL paths structurally.</comment>
<file context>
@@ -66,6 +74,14 @@ func resolveJevEvaluator(ctx context.Context, keys jevKeyResolver, env jevEvalua
+// hosts, paths without scheme, non-http schemes) is rejected.
+func validJevBaseURL(raw string) bool {
+ u, err := url.Parse(raw)
+ return err == nil && (u.Scheme == "https" || u.Scheme == "http") && u.Host != ""
+}
+
</file context>
| // A convention longer than the field cap would be truncated in state — | ||
| // a confident verdict on partial text could supersede or file a conflict | ||
| // on an incomplete rule. Escalate instead; the LLM prompt sees full text. | ||
| if len(candidate) > jevConventionFieldCap { |
There was a problem hiding this comment.
P3: The new oversize-escalation branch has no test coverage, unlike every other escalation path in this file (TestJevConventions_OneUncertainEscalatesWholeBatch, _ErrorEscalates, _InvalidChoiceEscalates). Add a test asserting that an oversized candidate or neighbor returns ok=false, escalates to the LLM with full text, and bills zero Jev spend.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/jev_conventions.go, line 82:
<comment>The new oversize-escalation branch has no test coverage, unlike every other escalation path in this file (TestJevConventions_OneUncertainEscalatesWholeBatch, _ErrorEscalates, _InvalidChoiceEscalates). Add a test asserting that an oversized candidate or neighbor returns ok=false, escalates to the LLM with full text, and bills zero Jev spend.</comment>
<file context>
@@ -76,6 +76,17 @@ func jevConventionState(candidate string, neighbors []memory.PatternMatch) map[s
+ // A convention longer than the field cap would be truncated in state —
+ // a confident verdict on partial text could supersede or file a conflict
+ // on an incomplete rule. Escalate instead; the LLM prompt sees full text.
+ if len(candidate) > jevConventionFieldCap {
+ return nil, StageTokens{}, false
+ }
</file context>
| // hosts, paths without scheme, non-http schemes) is rejected. | ||
| func validJevBaseURL(raw string) bool { | ||
| u, err := url.Parse(raw) | ||
| return err == nil && (u.Scheme == "https" || u.Scheme == "http") && u.Host != "" |
There was a problem hiding this comment.
P3: validJevBaseURL accepts plaintext http:// endpoints even though this URL receives the TypeSafe API key (Bearer header) and tenant PR/finding content. For the egress of credentials plus data, restrict the accepted scheme to https (self-hosted plaintext relays can be terminated behind an explicit TLS proxy instead).
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/jev_resolver.go, line 82:
<comment>`validJevBaseURL` accepts plaintext `http://` endpoints even though this URL receives the TypeSafe API key (Bearer header) and tenant PR/finding content. For the egress of credentials plus data, restrict the accepted scheme to `https` (self-hosted plaintext relays can be terminated behind an explicit TLS proxy instead).</comment>
<file context>
@@ -66,6 +74,14 @@ func resolveJevEvaluator(ctx context.Context, keys jevKeyResolver, env jevEvalua
+// hosts, paths without scheme, non-http schemes) is rejected.
+func validJevBaseURL(raw string) bool {
+ u, err := url.Parse(raw)
+ return err == nil && (u.Scheme == "https" || u.Scheme == "http") && u.Host != ""
+}
+
</file context>
model_pricing misses previously costed $0. Now the lookup chain falls back to OpenRouter's public /models catalog (no auth): manual rows still win, then exact id → unique suffix → prefix match. Catalog is cached 6h, failed fetches serve stale and throttle retries per TTL. Ambiguous suffix matches never guess a price. Generated with [Devin](https://devin.ai)
There was a problem hiding this comment.
1 issue found across 3 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="backend/internal/app/app.go">
<violation number="1" location="backend/internal/app/app.go:167">
P2: When a review uses a non-OpenRouter or custom OpenAI-compatible provider without a manual pricing row, this fallback still assigns an OpenRouter catalog price because the callback cannot see the provider. Restrict the catalog fallback to OpenRouter calls or make cost lookup provider-aware; otherwise cost telemetry and billing are incorrect.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| if in, out, ok := pricingCache.Lookup(ctx, model); ok { | ||
| return in, out, true | ||
| } | ||
| return openRouterPricing.Lookup(model) |
There was a problem hiding this comment.
P2: When a review uses a non-OpenRouter or custom OpenAI-compatible provider without a manual pricing row, this fallback still assigns an OpenRouter catalog price because the callback cannot see the provider. Restrict the catalog fallback to OpenRouter calls or make cost lookup provider-aware; otherwise cost telemetry and billing are incorrect.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/app/app.go, line 167:
<comment>When a review uses a non-OpenRouter or custom OpenAI-compatible provider without a manual pricing row, this fallback still assigns an OpenRouter catalog price because the callback cannot see the provider. Restrict the catalog fallback to OpenRouter calls or make cost lookup provider-aware; otherwise cost telemetry and billing are incorrect.</comment>
<file context>
@@ -154,11 +154,17 @@ func Run() error {
+ if in, out, ok := pricingCache.Lookup(ctx, model); ok {
+ return in, out, true
+ }
+ return openRouterPricing.Lookup(model)
})
logger.InfoContext(ctx, "pricing cache initialization completed", "cache_ttl", 10*time.Minute)
</file context>
- handlers_config: reject malformed typesafe base_url with 400 - jev: export ValidBaseURL shared by resolver + config handler - jev: emit jev.call.failed on oversized state for failure parity - config: firstNonBlankEnv trims TYPESAFE_API_KEY before fallback - triage: injectable join/drain budgets — tests use ms windows - tests: conventions missing-answer escalation, intent Diff fixture, scoring defect_2 below veto, NaN/Inf answer guards, resolver atomics
There was a problem hiding this comment.
1 existing issue remains and 3 new issues found across 12 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="backend/internal/jev/jev_test.go">
<violation number="1" location="backend/internal/jev/jev_test.go:176">
P3: The comment's premise is incorrect: encoding/json returns an UnmarshalTypeError whenever strconv.ParseFloat overflows float64, so `1e400` would make Evaluate error-out and escalate to the LLM fallback — the Noul/Choice range guard never sees it. The added Inf cases still validly exercise the guards (defense in depth for directly constructed Results), but the comment misstates the threat model and should be corrected so readers don't believe a hostile endpoint can deliver +Inf probabilities untouched.</violation>
</file>
<file name="backend/internal/pipeline/jev_resolver_test.go">
<violation number="1" location="backend/internal/pipeline/jev_resolver_test.go:230">
P3: In TestResolveCandidates_BYOKSelectsJevCascade, the failure message formats the atomic.Value itself instead of its payload. %q on an atomic.Value prints the unexported struct contents, so when this test fails the message shows no actual auth header — the one value the test exists to verify. Change %q's operand to auth.Load().</violation>
</file>
<file name="backend/internal/pipeline/jev_triage_shadow_test.go">
<violation number="1" location="backend/internal/pipeline/jev_triage_shadow_test.go:193">
P3: The millisecond budgets leave ~20ms of wall-clock headroom, which can flake under a loaded `-race`/CI runner. `finishJevTriageShadow` realistically waits join+drainBudget (≈40ms), but the cap is `2*join` (60ms) — two back-to-back timer overshoots of ~10ms each fail the assertion. The delay (2*join=60ms) is also only 20ms past join+drain (40ms), so the abandon ordering relies on timer accuracy. Since the fake's Evaluate is aborted by the cancel, a much longer delay adds no wall time: widen the margins so only an order-of-magnitude skew can flake.</violation>
</file>
Requires human review: Auto-approval blocked because this review re-detected 1 unresolved issue already reported by Cubic.
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| // 1e400 decodes to +Inf without a json error — the range guard is | ||
| // the only thing stopping it clearing a caller's threshold. |
There was a problem hiding this comment.
P3: The comment's premise is incorrect: encoding/json returns an UnmarshalTypeError whenever strconv.ParseFloat overflows float64, so 1e400 would make Evaluate error-out and escalate to the LLM fallback — the Noul/Choice range guard never sees it. The added Inf cases still validly exercise the guards (defense in depth for directly constructed Results), but the comment misstates the threat model and should be corrected so readers don't believe a hostile endpoint can deliver +Inf probabilities untouched.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/jev/jev_test.go, line 176:
<comment>The comment's premise is incorrect: encoding/json returns an UnmarshalTypeError whenever strconv.ParseFloat overflows float64, so `1e400` would make Evaluate error-out and escalate to the LLM fallback — the Noul/Choice range guard never sees it. The added Inf cases still validly exercise the guards (defense in depth for directly constructed Results), but the comment misstates the threat model and should be corrected so readers don't believe a hostile endpoint can deliver +Inf probabilities untouched.</comment>
<file context>
@@ -166,28 +166,35 @@ func TestEvaluate_OversizedResponseErrors(t *testing.T) {
+ "out_high": {Type: TypeNoul, Noul: ptr(1.01)},
+ "out_neg": {Type: TypeNoul, Noul: ptr(-0.01)},
+ "nan": {Type: TypeNoul, Noul: &nan},
+ // 1e400 decodes to +Inf without a json error — the range guard is
+ // the only thing stopping it clearing a caller's threshold.
+ "pos_inf": {Type: TypeNoul, Noul: &posInf},
</file context>
| // 1e400 decodes to +Inf without a json error — the range guard is | |
| // the only thing stopping it clearing a caller's threshold. | |
| // Defense in depth: encoding/json already errors on floats that | |
| // overflow float64 (ParseFloat ErrRange), so +Inf cannot reach the | |
| // guard from the wire; it still rejects directly-constructed results. |
| if hits.Load() != 1 { | ||
| t.Fatalf("BYOK eval must hit the stored endpoint once, got %d", hits.Load()) | ||
| } | ||
| if auth.Load() != "Bearer ts-byok" { |
There was a problem hiding this comment.
P3: In TestResolveCandidates_BYOKSelectsJevCascade, the failure message formats the atomic.Value itself instead of its payload. %q on an atomic.Value prints the unexported struct contents, so when this test fails the message shows no actual auth header — the one value the test exists to verify. Change %q's operand to auth.Load().
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/jev_resolver_test.go, line 230:
<comment>In TestResolveCandidates_BYOKSelectsJevCascade, the failure message formats the atomic.Value itself instead of its payload. %q on an atomic.Value prints the unexported struct contents, so when this test fails the message shows no actual auth header — the one value the test exists to verify. Change %q's operand to auth.Load().</comment>
<file context>
@@ -222,10 +224,10 @@ func TestResolveCandidates_BYOKSelectsJevCascade(t *testing.T) {
+ t.Fatalf("BYOK eval must hit the stored endpoint once, got %d", hits.Load())
}
- if auth != "Bearer ts-byok" {
+ if auth.Load() != "Bearer ts-byok" {
t.Fatalf("stored key must auth the eval, got %q", auth)
}
</file context>
| start := time.Now() | ||
| ts.finishJevTriageShadow(context.Background(), run, s, | ||
| map[string]TriageResult{"f0.go": {File: "f0.go", Action: TriageDeep}}) | ||
| if elapsed := time.Since(start); elapsed >= 2*join { |
There was a problem hiding this comment.
P3: The millisecond budgets leave ~20ms of wall-clock headroom, which can flake under a loaded -race/CI runner. finishJevTriageShadow realistically waits join+drainBudget (≈40ms), but the cap is 2*join (60ms) — two back-to-back timer overshoots of ~10ms each fail the assertion. The delay (2*join=60ms) is also only 20ms past join+drain (40ms), so the abandon ordering relies on timer accuracy. Since the fake's Evaluate is aborted by the cancel, a much longer delay adds no wall time: widen the margins so only an order-of-magnitude skew can flake.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/jev_triage_shadow_test.go, line 193:
<comment>The millisecond budgets leave ~20ms of wall-clock headroom, which can flake under a loaded `-race`/CI runner. `finishJevTriageShadow` realistically waits join+drainBudget (≈40ms), but the cap is `2*join` (60ms) — two back-to-back timer overshoots of ~10ms each fail the assertion. The delay (2*join=60ms) is also only 20ms past join+drain (40ms), so the abandon ordering relies on timer accuracy. Since the fake's Evaluate is aborted by the cancel, a much longer delay adds no wall time: widen the margins so only an order-of-magnitude skew can flake.</comment>
<file context>
@@ -175,17 +175,22 @@ func TestJevTriageShadow_ErrorStampsNoProvider(t *testing.T) {
ts.finishJevTriageShadow(context.Background(), run, s,
map[string]TriageResult{"f0.go": {File: "f0.go", Action: TriageDeep}})
- if elapsed := time.Since(start); elapsed >= 2*jevShadowJoinBudget {
+ if elapsed := time.Since(start); elapsed >= 2*join {
t.Fatalf("join blocked %v, want <= join budget", elapsed)
}
</file context>
- jev.ValidBaseURL: reject query/fragment/userinfo; http only for loopback/private/link-local hosts - resolver: invalid stored base_url fails closed (no evaluator) instead of silently rerouting tenant data to the default endpoint - jev: check status before size so large error bodies surface the API's real message; cost-only stages counted in stats/models - conventions: fold Jev + LLM legs into Aux separately on escalation; foldAuxTokens flattens nested ledgers - openrouter pricing: async refresh (Lookup never blocks), cancellable ctx, 5min failure retry vs 6h TTL, NaN/Inf price guards, opt-out env OPENROUTER_PRICING_ENABLED - tests: oversize escalation w/ full-text assert, "too large" error pin, headline-after-fold pin, join-exact shadow timing, NaN/Inf rows
|
Disposition on the three review runs (244ef3cd re-review, 1111b7b1, 2a146647). Fixed in this push (9eb4071)Run 1111b7b1:
Run 2a146647 (pricing):
Stale — already fixed before filing (run 244ef3cd leftovers)
One intentional design decision (not a defect)
Outstanding
|
There was a problem hiding this comment.
2 existing issues remain and 3 new issues found across 15 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="backend/internal/api/handlers_org_stats.go">
<violation number="1" location="backend/internal/api/handlers_org_stats.go:222">
P3: The relaxed skip condition in addModel (stats/models) has no test pinning it, while the identical cost-only fix in the stage-costs aggregator is covered by TestAggregateStageCosts_KeepsCostOnlyStage. No test in the repo references statsModels/addModel/StatsModels, so a future change that reverts this condition (e.g., back to gating on TotalTokens alone) would silently drop cost-only spend from the dashboard's per-model stats with no failing test. Add a test that feeds a RunTokenUsage with {TotalTokens: 0, Cost > 0, Model set} and asserts the model appears with that cost, mirroring the existing stage-costs test.</violation>
</file>
<file name="backend/internal/llm/openrouter_pricing.go">
<violation number="1" location="backend/internal/llm/openrouter_pricing.go:67">
P2: When the first OpenRouter-only completion finishes before the background warm fetch completes, `LookupCtx` returns not-found and `EstimateCost` records a zero cost; the later refresh cannot repair that completed usage record. Keep a bounded initial fetch in the pricing path, or defer/recompute cost after the in-flight refresh completes.</violation>
</file>
<file name="backend/internal/pipeline/types.go">
<violation number="1" location="backend/internal/pipeline/types.go:536">
P2: The flatten only preserves per-leg provenance in memory; the persistence path for the AutoResolve bucket still collapses it. `resolveCandidates` passes the summed bucket (`stats.judgeTokens = tokens.AutoResolve`, now carrying one Aux entry per judge call — Jev legs + LLM escalation) to `persistAsyncStageTokens`, which marshals it as `entry` and merges via `MergeStageTokenEntry`. That SQL does `jsonb_build_array($2::jsonb - 'aux')` — it strips the entry's own aux and appends the bucket as one headline-stamped entry. Since auto-resolve persists with `run == nil` (`persistAsyncStageTokens(..., nil)`), the in-memory copy the fold feeds doesn't exist on that path, so the per-leg split the new fold exists to preserve never reaches the reviews.token_usage row that /stats and the per-review TokenPill read. Make the DB merge append the leg entries (or have the caller flatten before persisting) so the stored AutoResolve bucket keeps per-model attribution; the current head stores only the aggregate plus a single stamped entry.</violation>
</file>
Requires human review: Auto-approval blocked because this review re-detected 2 unresolved issues already reported by Cubic.
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
|
|
||
| // Warm kicks an asynchronous catalog fetch — startup warming only; the fetch | ||
| // itself is throttled and single-flighted by maybeRefresh. | ||
| func (o *OpenRouterPricing) Warm(ctx context.Context) { o.maybeRefresh(ctx) } |
There was a problem hiding this comment.
P2: When the first OpenRouter-only completion finishes before the background warm fetch completes, LookupCtx returns not-found and EstimateCost records a zero cost; the later refresh cannot repair that completed usage record. Keep a bounded initial fetch in the pricing path, or defer/recompute cost after the in-flight refresh completes.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/llm/openrouter_pricing.go, line 67:
<comment>When the first OpenRouter-only completion finishes before the background warm fetch completes, `LookupCtx` returns not-found and `EstimateCost` records a zero cost; the later refresh cannot repair that completed usage record. Keep a bounded initial fetch in the pricing path, or defer/recompute cost after the in-flight refresh completes.</comment>
<file context>
@@ -53,28 +62,14 @@ type openRouterCatalogEntry struct {
- defer cancel()
+// Warm kicks an asynchronous catalog fetch — startup warming only; the fetch
+// itself is throttled and single-flighted by maybeRefresh.
+func (o *OpenRouterPricing) Warm(ctx context.Context) { o.maybeRefresh(ctx) }
+
+// Refresh fetches the catalog synchronously and swaps the price map on
</file context>
| // A bucket folded into a bucket flattens its legs — a caller that already | ||
| // mixed providers (Jev leg + LLM escalation) returns per-leg provenance | ||
| // that must survive, not collapse into one headline-stamped entry. | ||
| if len(spend.Aux) > 0 { |
There was a problem hiding this comment.
P2: The flatten only preserves per-leg provenance in memory; the persistence path for the AutoResolve bucket still collapses it. resolveCandidates passes the summed bucket (stats.judgeTokens = tokens.AutoResolve, now carrying one Aux entry per judge call — Jev legs + LLM escalation) to persistAsyncStageTokens, which marshals it as entry and merges via MergeStageTokenEntry. That SQL does jsonb_build_array($2::jsonb - 'aux') — it strips the entry's own aux and appends the bucket as one headline-stamped entry. Since auto-resolve persists with run == nil (persistAsyncStageTokens(..., nil)), the in-memory copy the fold feeds doesn't exist on that path, so the per-leg split the new fold exists to preserve never reaches the reviews.token_usage row that /stats and the per-review TokenPill read. Make the DB merge append the leg entries (or have the caller flatten before persisting) so the stored AutoResolve bucket keeps per-model attribution; the current head stores only the aggregate plus a single stamped entry.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/pipeline/types.go, line 536:
<comment>The flatten only preserves per-leg provenance in memory; the persistence path for the AutoResolve bucket still collapses it. `resolveCandidates` passes the summed bucket (`stats.judgeTokens = tokens.AutoResolve`, now carrying one Aux entry per judge call — Jev legs + LLM escalation) to `persistAsyncStageTokens`, which marshals it as `entry` and merges via `MergeStageTokenEntry`. That SQL does `jsonb_build_array($2::jsonb - 'aux')` — it strips the entry's own aux and appends the bucket as one headline-stamped entry. Since auto-resolve persists with `run == nil` (`persistAsyncStageTokens(..., nil)`), the in-memory copy the fold feeds doesn't exist on that path, so the per-leg split the new fold exists to preserve never reaches the reviews.token_usage row that /stats and the per-review TokenPill read. Make the DB merge append the leg entries (or have the caller flatten before persisting) so the stored AutoResolve bucket keeps per-model attribution; the current head stores only the aggregate plus a single stamped entry.</comment>
<file context>
@@ -530,8 +530,14 @@ func foldAuxTokens(bucket *StageTokens, spend StageTokens) {
+ // A bucket folded into a bucket flattens its legs — a caller that already
+ // mixed providers (Jev leg + LLM escalation) returns per-leg provenance
+ // that must survive, not collapse into one headline-stamped entry.
+ if len(spend.Aux) > 0 {
+ bucket.Aux = append(bucket.Aux, spend.Aux...)
+ } else {
</file context>
| } | ||
| return | ||
| } | ||
| if st.Model == "" || (st.TotalTokens == 0 && st.Cost == 0) { |
There was a problem hiding this comment.
P3: The relaxed skip condition in addModel (stats/models) has no test pinning it, while the identical cost-only fix in the stage-costs aggregator is covered by TestAggregateStageCosts_KeepsCostOnlyStage. No test in the repo references statsModels/addModel/StatsModels, so a future change that reverts this condition (e.g., back to gating on TotalTokens alone) would silently drop cost-only spend from the dashboard's per-model stats with no failing test. Add a test that feeds a RunTokenUsage with {TotalTokens: 0, Cost > 0, Model set} and asserts the model appears with that cost, mirroring the existing stage-costs test.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At backend/internal/api/handlers_org_stats.go, line 222:
<comment>The relaxed skip condition in addModel (stats/models) has no test pinning it, while the identical cost-only fix in the stage-costs aggregator is covered by TestAggregateStageCosts_KeepsCostOnlyStage. No test in the repo references statsModels/addModel/StatsModels, so a future change that reverts this condition (e.g., back to gating on TotalTokens alone) would silently drop cost-only spend from the dashboard's per-model stats with no failing test. Add a test that feeds a RunTokenUsage with {TotalTokens: 0, Cost > 0, Model set} and asserts the model appears with that cost, mirroring the existing stage-costs test.</comment>
<file context>
@@ -219,7 +219,7 @@ func (s *Server) statsModels(w http.ResponseWriter, r *http.Request) {
return
}
- if st.Model == "" || st.TotalTokens == 0 {
+ if st.Model == "" || (st.TotalTokens == 0 && st.Cost == 0) {
return
}
</file context>
waitSiblingFanoutSettled read the lookup counter after OnReviewCompleted had already spawned the fanout goroutine. When the lookup landed first, the helper waited for a count that had already passed and burned the full deadline. Capture the baseline before triggering the fanout.
Summary
Ports the Jev integration (merged on argus-private as #301): TypeSafe Jev (
jev-1.13.0, System One eval API) front-runs Argus's typed classification calls, cutting LLM spend on the pipeline's highest-volume classifier decisions. LLM paths stay untouched as the fallback — Jev only decides when its probabilities clear conservative thresholds.backend/internal/jev— minimal client forPOST /v1/systemone: batched questions over one state, pinned model, input-only cost math,jev.call.*telemetry mirroringllm.call.*, 120KB state cap, probability range validation.Four cascades + one shadow, all fail-safe:
p≥0.97addressed +p≥0.5visible evidence → resolve;p≤0.05→ keep open; else → LLM judgechoiceper neighbor, all≥0.95→ skip LLM; any uncertainty escalates whole batchdelivers+ per-criterion + per-finding-scope nouls; confident all-clear onlyfp≥0.97ANDdefect≤0.5→ leaves judge prompt, re-enters as score-10 synthetic groupjev.shadow.triage, NEVER routedBYOK credentials: per-installation keys via the encrypted
provider_keysstore (providertypesafe, repo-level then org-level) — Settings → Providers → TypeSafe Jev card beside Embeddings (masked key, replace/delete, optional custom endpoint). A stored key IS the consent — no flag needed. The envTYPESAFE_API_KEYpath still requiresjev_classifieropt-in (default OFF). Resolution is lazy — candidate-free pushes pay zero lookups.Token accounting: Jev spend merges into each stage bucket (
AutoResolve,Intent,Scoring,Conventions,Triage); previously-unbilled convention classifier spend is now billed on both legs.Data egress: sanitized finding text, diff hunks, PR metadata/body, intent, stored conventions →
api.typesafe.ai(or stored endpoint), only per the rules above. Documented indocs/self-hosting.md+backend/.env.example.Test plan
go build ./... && go vet ./... && go test ./...— all packages greenpnpm lint && pnpm typecheck && pnpm build— greenSummary by cubic
Adds TypeSafe's Jev classifier as a front-runner for addressed-judge, convention-relation, intent-verification, and scoring false-positive decisions, plus an observe-only triage shadow. Previously every decision went to the LLM; now confident Jev answers skip it, anything uncertain, missing, oversized, or failed falls back unchanged, and a flaky cross-PR test is fixed by capturing the lookup count before review-completion fanout.
New Features
p>=0.97with visible evidence, keeps open atp<=0.05, and escalates the middle band to the LLM.>=0.95; any uncertain answer escalates the whole batch.delivers, acceptance criteria, and per-finding scope.fp>=0.97plusdefect<=0.5, then re-enters them as score-10 synthetic groups for the audit trail.jev.shadow.triageagreement on up to 40 files; slow evals are abandoned without billing.jev.call.failed.model_pricingrows still win; models without one now use OpenRouter's public catalog, refreshed in the background, instead of costing $0.Migration / Ops
typesafeprovider key (repo-level then org-level) or withTYPESAFE_API_KEYplus thejev_classifierfeature flag; no key means no Jev calls.typesafeendpoints are validated and normalized on save; malformed URLs return a 400, and an invalid stored URL disables Jev instead of rerouting traffic.Auxledger, andstats/modelsreports per-model totals for every scalar stage.OPENROUTER_PRICING_ENABLED=falsein restricted-egress deployments.Written for commit 8d2366a. Summary will update on new commits.