Tracekit is a signer that sits between an AI agent and its tools. Before each tool call runs, the signer checks it
against your policy and allows it, denies it, or holds it for a person to approve. Every call and decision is written
as a signed, hash-chained record that the agent can't edit. A run exports as one .tkb file that anyone can verify
offline.
Agents now run shell commands, edit code, query databases and send messages on real systems. Most agent logs are written by the agent's own process, so whatever controls the agent can also edit them or skip them. When something goes wrong, you need a record you can trust. You also need to show it to someone else: a security team, an auditor, a customer. Tracekit keeps that record outside the agent's control, and lets anyone check it without trusting you.
Record
- Every tool call, its decision and its result, plus model calls and approvals, as signed records in one chain per run.
- A missed or lost event becomes a signed gap record. Gaps are never silent.
Decide
- Policy packs for coding agents, server agents, browser agents, SQL, HTTP, cloud CLIs, and payments and messaging (policy reference).
- Each call is allowed, denied, or held until a person answers (
ask).
Approve
- A held call runs only after an approver says yes. The approval is bound to the exact arguments and used once.
- Approvers answer from the command line, from a web page with OIDC sign-in and passkeys, or from Slack (approvals).
Prove
tracekit verifychecks a bundle offline against keys you pinned.- Witnesses cosign checkpoints of the log, and a monitor watches it from outside (witnesses, monitor).
- Each report states an assurance level:
dev,local,witnessedorwitnessed+monitored(reading a verify report).
Watch
tracekit viewshows runs with their verdicts, blocked and held calls, approvals, gaps, a replay of each run, and a run review (viewer).
Explain (1.1, release candidate)
- Where each argument came from: the signer tracks values that entered a run through untrusted content (a fetched page, an email, a ticket) and can hold a call whose recipient, payee or host came only from there (provenance).
- Which input caused an action:
tracekit whyreplays a run with each suspect input removed, tools served from the recording, and reports the cause with a confidence interval; results on signed runs are signed findings the verifier rechecks (tracekit why).
Integrate
- Coding agents through their hooks: Claude Code, Codex CLI, Cursor, Gemini CLI (coding agents).
- Python and TypeScript SDKs, with integrations for LangChain and LangGraph, the OpenAI Agents SDK, the Claude Agent SDK, MCP clients, Browser Use (Python) and the Vercel AI SDK (TypeScript) (quickstarts).
- OpenTelemetry traces imported as evidence (OpenTelemetry).
- An LLM gateway for OpenAI-compatible APIs that records model calls outside the agent's process (gateway).
agent ──► client or hook ──────► signer ─────────────► witnesses
(SDK, adapter, - policy: allow, cosign each
coding-agent hook) deny or ask checkpoint
holds a run token, - keys: signs every
never a key record
- log: hash chain,
Merkle checkpoints
│
▼
run.tkb bundle ──► tracekit verify
(offline, with the keys
and witnesses you pinned)
- The agent's client asks the signer to decide each tool call before it runs.
- The signer decides with its own policy, signs the decision, and adds it to the log.
- After the call, the client reports the result, and the signer records that too.
- The signer signs checkpoints of the log. Witnesses cosign them.
- A run is exported as a bundle. Anyone can verify it offline.
The trust model:
- The agent never holds a signing key and never assigns sequence numbers.
- Nothing the agent reports about itself is trusted: not its isolation, its fail mode, gaps, or who approved a call. The signer measures or configures these itself.
- Who can rewrite the log decides what a bundle is worth. A dev signer runs as your own user, so an agent that runs commands as you can rewrite it. A signer under a separate OS user, in a container sidecar, or on a central host is out of the agent's reach.
- The operator who runs the signer could still rewrite its log with its own key. An independent witness and a monitor make that visible.
The full model: architecture, laptop and server threat models.
Three steps, with the agent you already have (OpenAI Agents SDK, LangGraph / LangChain, Claude Agent SDK, an MCP client, or plain OpenAI / Anthropic / Google Gen AI calls):
pip install tracekit-aiimport tracekit; tracekit.instrument() # the first line of your agent; then run it as usualtracekit last # Integrity: VERIFIED. / Assurance: dev; ...tracekit.instrument() starts a dev signer in the background (or reuses the running one) and gates and records every
tool and model call of the frameworks it finds. tracekit last exports the most recent finished run, pins the dev
signer's key, verifies the bundle and says where it is. Then tracekit view --dev browses every run: verdicts,
denials, approvals, replay. A call the policy holds for a person (ask) pauses the run or waits, depending on the
framework:
how each integration waits.
Per framework, with a 10-line example each:
OpenAI Agents SDK,
LangChain / LangGraph,
Claude Agent SDK,
MCP client,
a custom agent. The
v2 quickstart explains each step and what
dev assurance means; limits lists what this path
doesn't cover.
No agent at hand? From a checkout, python examples/v2/langchain/agent.py --scripted runs a LangGraph agent with a
mock model whose calls are allowed, denied and held for an approval. Without one, the same loop in a temp dir (with a
witness cosignature and a tampered copy that fails), and Claude Code's hooks wired to a dev signer:
tracekit demo --server # exit 0
tracekit demo --why # exit 0: a held email, its cause found by replay and cited by its signed record
tracekit init --dev --v2Or check the shipped sample bundles (v1) with the verifier alone (pip install --no-deps tracekit-ai is enough):
tracekit verify docs/sample/demo-run.tkb --key docs/sample/signer.pub # exit 0
tracekit verify docs/sample/demo-run-tampered.tkb --key docs/sample/signer.pub # exit 1- Python 3.9 to 3.13, and
pip install tracekit-ai. The framework integrations need Python 3.10 or newer. For the TypeScript SDK, Node 18 or newer and@cygnux/tracekit. - An OS. Dev mode runs on macOS, Linux and Windows. A signer under its own OS user runs on Linux (macOS is experimental). Containers and Kubernetes run on Linux (platforms).
- To try it: nothing more. Dev mode starts a signer as your own user.
- For protection against the agent: a signer the agent can't reach. That is a separate OS user, a container sidecar, or a central signer (deployment guide).
- For protection against the operator: a witness the operator doesn't control, and a monitor.
- For approvals by people: approvers listed in the signer's config. They answer from the command line, or through an OIDC identity provider (web page, passkeys) or Slack (identity).
- Dev mode protects nothing against an agent that runs as you. It can read the key, stop the signer and re-sign
the log; dev bundles say
isolation: same-user(FAQ). - A verified bundle proves the record is intact. It does not prove intent, that a reported tool result is true, or that every action was captured (FAQ).
- Actions outside the capture path leave no record, such as a subprocess a hook never sees or code that skips the adapter (adapters).
- Policy rules match text. A determined agent can reword a command so no rule matches. Rules are tripwires; the protection is the signer's isolation (policy reference).
- Rules match tool names. A shell tool named
run_commandis not covered by the shell rules until you map it (adapters). - Fail-open tool classes run while the signer is down, unrecorded; the signer then writes a signed gap for them. Every class fails closed by default (FAQ).
- The agent's own claims are never trusted. Isolation, fail modes and approvers come from the signer's config, so set them there (approvals).
- Without an independent witness, the operator can rewrite the log unseen. A system-mode signer with no witness
still verifies as
Assurance: dev(deployment). - Commitments hide content, not metadata. Timing, sizes and counts still show, and witnesses learn when checkpoints happen (privacy).
- The viewer's verdicts are operator-side. For evidence, verify the bundle yourself with keys you pinned (viewer).
Every limit: docs/limits.md
The agent's code is the same in every mode; TRACEKIT_SIGNER points it at the signer.
| Mode | Who can reach the keys and log | Recorded isolation |
|---|---|---|
| Laptop dev | the agent's own user | same-user |
| Laptop system mode (Linux; macOS experimental) | root and the signer's OS user | separate-user |
| Container or Kubernetes sidecar | the signer's uid and the node's administrator | separate-user |
| Central signer over HTTPS (or Docker Compose) | the signer host's administrators | remote |
Each mode end to end, with its commands, its tracekit doctor check and the assurance its bundles reach:
deployment guide. Moving from the v1 signer or from
a laptop to a server: migration.
For anyone evaluating the format itself, or writing another verifier:
- Specified. Evidence format v2 aims to be precise enough to write an independent verifier. Normative test vectors pin canonical JSON, record signatures and checkpoints.
- Built on standards. Canonical JSON is JCS (RFC 8785). Records are signed with Ed25519. Each log is an RFC 6962 Merkle tree. Checkpoints are C2SP signed notes, and witnesses cosign them per C2SP tlog-cosignature.
- Self-contained. A bundle is a zip of records, inclusion proofs, checkpoints and the policy snapshots that decided the calls. It holds no verification code. The verifier checks it offline against a trust file you pinned: log keys, witnesses and quorum, and optionally a monitor.
- Old bundles keep verifying. Formats only grow, and each format's verifier is frozen. v1 bundles go to the v1 verifier, which needs only the Python standard library.
- Forward-safe. A bundle names the minimum verifier version it needs. An older verifier reports it
UNVERIFIABLE (needs tracekit >= x), neverFAILED.
Verify a v1 bundle in a GitHub Actions workflow:
- uses: Cygnux-Labs/Tracekit@main
with:
bundle: run.tkb
key: keys/signer.pub # pin the signer; or witness: git:/path/to/clonerequire-anchor defaults to true, so an unanchored bundle fails the step; set require-anchor: "false" to only
report it.
- How it fits together: architecture, evidence format v2
- Running it: deployment, doctor, witnesses, monitor, viewer, observability, privacy
- Deciding calls: policy reference, approvals, LLM gateway, OpenTelemetry
- Checking evidence: reading a verify report, auditor guide, evaluation, technical report, FAQ and limits, known limits and trade-offs
- Asking why an action happened: tracekit why, its architecture
- Releasing: launch checklist for 1.0
- The v1 laptop signer (
tracekitd,tracekit init --dev,tracekit demo,tracekit observe) and its bundles keep working: coding agents, signing, platforms
1.0 is the first stable release, and 1.x is the supported line (SECURITY.md).
- Stable: evidence format v2 and its verifier. 1.x verifiers verify every 1.0 bundle. v1 bundles and the v1 verifier are unchanged.
- Versioned: the signer RPC (version 12). Clients and signers of different RPC versions refuse each other, so upgrade them together (changelog).
- Experimental: laptop system mode on macOS, the Rekor witness of the v1 signer, and the v1 OTLP receiver, ingest
gateway and model proxy (each needs
--experimental). - Not yet done: an external review of the threat models, and some release checks that wait on the owner (launch checklist).
CONTRIBUTING.md: make install, then make check.
Report vulnerabilities privately as described in SECURITY.md.
MIT, Copyright Cygnux Labs. See LICENSE.
