Skip to content

feat(e2e): preserve Kubernetes clusters for debugging #3675

Description

@danehans

User Story

As an OpenShell contributor debugging a Kubernetes E2E failure, I want to preserve the ephemeral cluster and test artifacts after the run so that I can inspect the failed state directly.

Problem Statement

The Kubernetes E2E wrapper always deletes the k3d cluster and temporary work directory from its exit trap. Its failure diagnostics capture common resources, but they cannot anticipate every investigation and remove the live state needed for follow-up commands.

Impact / Why This Matters

Contributors must reproduce a potentially slow or intermittent failure, modify the wrapper locally, or run the setup manually. This increases investigation time and can hide evidence that exists only immediately after a failure.

Proposed Design

Add an opt-in OPENSHELL_E2E_KUBE_PRESERVE_CLUSTER=1 mode for ephemeral clusters created by the E2E wrapper. When enabled, the wrapper preserves its cluster and work directory and prints their locations plus explicit cleanup commands. Existing cleanup remains the default, and externally supplied clusters are not reclassified as wrapper-owned resources.

Preserved work directories can contain short-lived test credentials, so the output and documentation must tell contributors to remove them after investigation.

Acceptance Criteria

  • OPENSHELL_E2E_KUBE_PRESERVE_CLUSTER=1 preserves an ephemeral cluster created by the wrapper and its work directory after success or failure.
  • The wrapper prints the cluster name, kubeconfig, work directory, and explicit commands to delete both preserved resources.
  • Default behavior continues to delete wrapper-created clusters and temporary files.
  • The option does not change ownership or cleanup behavior for a caller-supplied Kubernetes context.
  • Invalid option values fail before cluster creation.
  • Contributor documentation describes the option and the temporary credential cleanup requirement.

Alternatives Considered

Editing the exit trap locally is error-prone and not reproducible. Reusing a manually managed cluster helps some investigations but does not preserve the exact state produced by the standard ephemeral E2E workflow.

Agent Investigation

The cleanup trap in e2e/with-kube-gateway.sh collects a fixed diagnostic bundle, uninstalls fixtures, deletes a wrapper-created k3d cluster, and removes its temporary work directory. An early, opt-in preservation branch can retain the live environment without changing default behavior.

Checklist

  • I've reviewed existing issues and the architecture docs
  • This is a design proposal, not a "please build this" request

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions