Skip to content

bug(sandbox): network broker logs expected socket races as warnings #3526

Description

@drew

This was generated by AI during triage.

Agent Diagnostic

  • Skills loaded: diagnose, OpenShell repository create-github-issue, and cluster inspection workflows
  • OpenShell version tested: 0.0.117-dev.211+g4cd5e5478
  • Latest release checked: v0.0.116; the issue reproduces on a newer development build containing the RFC 0012 sandbox runtime
  • Known fixes reviewed: release notes for v0.0.116, RFC 0012 implementation PR feat(isolation): implement the RFC 0012 sandbox architecture #2942, and current main network-broker handling
  • Possible duplicates reviewed: searched all OpenShell issues and merged PRs for the exact network notification denied text, Socket not connected, and related network/boundary warnings. No matching issue was found. bug(supervisor-network): boundary reconnect is followed by proxy exit and ControlSupervisorExited #3396 concerns a boundary reconnect that becomes fatal; the sandboxes here remained healthy.
  • Findings: six healthy Kubernetes workload pods emitted 4,183 sandbox network notification denied warnings in 12 hours. Nearly all were syscall 52 returning ENOTCONN; the remainder were syscall 62 returning ESRCH. All pods stayed Ready with zero restarts and there were no error-level runtime logs. Current main logs every dispatch_notification error at WARN, regardless of whether it represents an enforcement decision or an expected process/socket race.
  • Remaining reason for filing: these per-syscall warnings are high-volume and misleading, and obscure actionable policy or boundary failures.

Description

Actual behavior: The RFC 0012 sandbox network broker emits a WARN for every failed seccomp notification dispatch:

WARN openshell_sandbox::network_broker: sandbox network notification denied (tid=..., syscall=52): Socket not connected (os error 107)
WARN openshell_sandbox::network_broker: sandbox network notification denied (tid=..., syscall=62): No such process (os error 3)

During a 12-hour observation of six active agent workloads, this produced 4,183 network-broker warnings. The workloads remained Ready, had zero restarts, and recorded no error-level runtime events. A two-hour sample was dominated by 910 ENOTCONN warnings and 34 ESRCH warnings.

The current handler groups all dispatch_notification errors under the message sandbox network notification denied, even when the returned errno describes a socket/process race rather than an OpenShell policy denial. This makes benign application behavior look like a security or isolation failure and makes genuine broker failures difficult to find.

Expected behavior: Expected transient process/socket races should not create one warning per syscall. They should be handled at debug/trace level, aggregated, or rate-limited. Genuine OpenShell policy denials and broker-health failures must remain clearly observable and distinguishable from application-originated errnos.

Reproduction Steps

  1. Deploy the RFC 0012 sandbox runtime on Linux with seccomp network mediation enabled.
  2. Run a long-lived agent workload that performs ordinary concurrent HTTP/socket activity and periodically starts and exits subprocesses.
  3. Collect the workload pod's runtime logs for several polling cycles.
  4. Count sandbox network notification denied messages.
  5. Observe repeated syscall 52/ENOTCONN and syscall 62/ESRCH warnings while the sandbox remains Ready and functional.

Environment

  • OS: Ubuntu 24.04, Linux amd64
  • Kubernetes: v1.35.7-gke.1222000
  • Agent Sandbox controller: v0.5.0, v1beta1 API
  • Docker: not applicable; Kubernetes compute driver
  • OpenShell: 0.0.117-dev.211+g4cd5e5478
  • Deployment: Helm 0.0.0-dev, RFC 0012 separate workload and supervisor pods
  • Latest release checked: yes, v0.0.116; the tested development build is newer
  • Possible duplicates checked: yes; bug(supervisor-network): boundary reconnect is followed by proxy exit and ControlSupervisorExited #3396 is related to fatal boundary recovery but does not cover this non-fatal per-syscall warning storm

Logs

# Counts by workload over 12 hours:
workload-1 network_denied=190 errors=0 restarts=0
workload-2 network_denied=90  errors=0 restarts=0
workload-3 network_denied=1224 errors=0 restarts=0
workload-4 network_denied=897 errors=0 restarts=0
workload-5 network_denied=1046 errors=0 restarts=0
workload-6 network_denied=736 errors=0 restarts=0

# Dominant messages over a two-hour sample:
910 WARN ... syscall=52: Socket not connected (os error 107)
 34 WARN ... syscall=62: No such process (os error 3)

Acceptance Criteria

  • Expected ENOTCONN and ESRCH process/socket races do not emit an unbounded WARN per syscall.
  • Genuine policy denials remain observable as explicit policy decisions rather than being conflated with application errno handling.
  • Fatal listener or broker-health failures remain at error level and retain actionable context.
  • Telemetry provides aggregate counts or suitably rate-limited diagnostics for suppressed transient failures.
  • Tests cover notification targets disappearing and sockets closing between notification receipt and dispatch without producing warning storms.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:sandboxSandbox runtime and isolation workos:linuxIssue affects Linux hoststopic:observabilityLogging, metrics, and observability work

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions