Skip to content
View HarperZ9's full-sized avatar

Block or report HarperZ9

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
HarperZ9/README.md
Zain Dana Harper. Tools and investigations that let anyone recheck what an AI system did, and who knew first.

Flywheel latest release Articulate latest release Telos latest release Writing Writing feed

I'm Zain Dana Harper. I build tools that let anyone recheck what an AI system did, and I publish investigations into who knew about AI incidents first and who pays the people who check. I work independently as a sole proprietor in Kent, Washington, and I take scoped evaluation work.

A result is worth trusting when an outside skeptic can rerun the check on their own machine and reach the same verdict. That holds whoever ran the model: any lab, open or closed, any company, any nation. My own verdicts get the same treatment.

Pick a door Go
Run an AI task with any model and keep a record you can recheck Flywheel and the flagships
Read the investigations Who Knew First and the series
See where the work is heading Watching the trace, checking before the action
Hire me for evaluation work Work with me
Get in touch Reach me

Flywheel and the flagships

Three verdicts: MATCH, the rerun agrees with the record; DRIFT, the rerun disagrees and says where; UNVERIFIABLE, the record cannot be checked.

Flywheel is a self-hostable, model-agnostic AI workstation and coding harness. It runs any model, frontier or local, behind one OpenAI-compatible surface, with your keys and data kept on your machine. Rowan, the desktop assistant, turns a plain request into a recorded run. Relay runs a permission-gated coding agent over your own folders. Lanes add research intake, workspace maps, memory, agent routing and writing. flywheel check-output checks answers against finance, medicine and law packs and can emit Lean 4 proofs. Every accepted run leaves a sealed receipt that an independent witness reruns offline, with no learned model deciding the verdict.

pip install flywheel-verify
flywheel up
flywheel lanes --probe

On the shipped benchmark the verified loop shows no measured accuracy gain over a single pass; the interval includes zero. The value is the workstation and a record you can check yourself. Flywheel is source-available under FSL-1.1-MIT.

Each flagship below also works on its own and plugs into Flywheel.

Current release of every flagship (refreshed daily from GitHub)
Tool What it does Release
Flywheel AI workstation and coding harness: any model, gated agent, rerunnable receipts v1.2.1 (2026-10-02)
Articulate Local writing checker and editor with content-free receipts v0.7.0 (2026-10-01)
Telos Accountable actuation: senses, actions and hardware control in permission tiers v0.7.0 (2026-10-02)
Accountable Surface Gates agent actions on explicit grants, with a hash-chained journal v0.2.0 (2026-09-18)
Forum Coordinates agent teams with a replayable ledger v1.16.0 (2026-10-01)
Relay Permission-checked coding agent for any model endpoint v0.6.0 (2026-10-01)
Gather Research intake from the web, papers, video, scans and audio, with provenance v2.1.0 (2026-10-01)
Index Offline repository and workspace maps with file and line evidence v2.15.0 (2026-10-01)
Mneme Agent memory where every recall can be rechecked v0.6.0 (2026-10-01)
Canon One memory and personality record shared across models and tools v0.6.0 (2026-10-01)
Crucible Tests falsifiable claims and records MATCH, DRIFT or UNVERIFIABLE v1.4.0 (2026-10-01)
EMET Checks that bytes reaching a model still match their source v1.3.0 (2026-09-13)
Learn Turns your own material into a course that never takes the test for you v2.1.0 (2026-10-01)
Plexus Finds and wires compatible tools in an agent toolchain v0.3.0 (2026-10-01)
Phantom Reversible hardware-identity privacy for owned Windows and Linux machines v1.1.1 (2026-09-10)
Work accepted upstream

Changes that survived another maintainer's review and merged:

The portfolio lists merged, open and closed contributions separately.

Who Knew First and the series

Nine marks for nine 2026 AI agent incidents. Six carry an outer mark: in those six, someone outside the organization that ran the model told the public first.

Who Knew First is a record of nine 2026 incidents in which an AI agent crossed a boundary. In six of them, someone outside the organization that ran the model told the public first. The argument is simple: whoever holds an incident's logs gets to name it, and the name decides how fast anyone else hears about it.

The series tests the questions that argument raises. Each piece stands alone, lists its sources and the confidence of each claim, and says what each claim does not prove.

Piece The question Status
Who Pays the Referees The people who check AI models depend on the labs they check. Which of those terms are public? Published 1 October 2026
The Terms for Telling The party that holds the records also writes the contracts of the people who could tell. Who got heard? Published 1 October 2026
Who Kept the Books In money cases from 1514 to Iran-Contra, what made the first account move? Planned
The Maker Is Part of the Story Three famous stories, read for who funds the work and who edits the record. Planned
A Check It Cannot Predict Does a check that is certain and outside the actor's control work on AI models too? Planned, no result yet

An Anthropic-built model helped compile these pieces, and Anthropic appears in the record, so each piece marks where Anthropic is a party and invites an outside check of those items.

Latest writing (refreshed daily from the site feed)
Date Piece
2026-10-01 Who Pays the Referees
2026-10-01 The Terms for Telling
2026-10-01 The Number Has a Vintage
2026-09-28 What the Formula Counts
2026-09-28 The Timestamp Is Not the Order
2026-09-28 The Scene the Song Did Not Tell You

Everything else, essays, briefings and papers, is on the writing page.

Watching the trace, checking before the action

A trace of observed steps runs into a check that sits before the action. One path continues to an action with a receipt; the other halts before the action runs.

A receipt tells you what happened after the fact. The next step is to watch an agent's trace as it runs and check each consequential action before it happens. If the check fails, the action halts on its own, with no person needing to step in, and the record shows why.

This is work in progress, and the pieces exist at different stages:

  • Accountable Surface lets an agent take only the action a person approved, then verifies the outcome and rolls back what it can.
  • Telos places sensing, actions and workstation hardware control in explicit permission tiers, each with confirmation points and receipts.
  • Rowan's monitor, in Flywheel 1.2.0 on PyPI, refuses to trust a check that rewrote its own grading files.

Trace observation uses what providers document: reasoning summaries, token counts, effort settings and ordinary outputs. It never tries to pull hidden reasoning out of a model through jailbreaks or prompt injection, and it never bypasses an access control.

How the pieces connect
flowchart LR
  T[Agent trace] --> M{Check before the action}
  M -- passes --> A[Action runs]
  M -- fails --> H[Halted, with the reason recorded]
  A --> R[Sealed receipt]
  H --> R
  R --> W[Independent rerun: MATCH, DRIFT or UNVERIFIABLE]
  W --> P[Published finding with its limits]
Loading

Work with me

I take scoped work on evaluation design review, harness integration, agent safety review before an audit, incident review, and conflict-of-interest review. Each engagement gets a quote built from the labor, time, compute and tooling it needs. I have no paid client today and no current sponsors.

Independence comes first, so the rules are public:

  • Income from this work is published with its exact source, API credits included.
  • When one source passes 15 percent of income over twelve months, I disclose it and give my findings about that party a second review.
  • At 50 percent, I decline new work evaluating that party. A first contract is most of the income by arithmetic, so it is disclosed in full and the decline rule waits for the second.
  • One standard for every lab. I build with Anthropic and OpenAI models and publish investigations that name both.
Roles and paths

I'm also open to technical and nontechnical roles in AI governance and evaluation, and willing to relocate to London or travel to San Francisco.

Path Where I fit Start here
Technical and evaluation Agent and model evaluation, developer tooling, CI, security testing, technical support and documentation. Engineering path
Public, union, and field Public service, facilities, parks and grounds, arboriculture, scheduling and safety judgment, from eleven years of field work. Public-service and field path
Education and research Fellowships, research operations and evidence-centered technical writing. Research

Resume · CV · Portfolio

Reach me

The art on this page is generated by scripts/profile_art.py. Motion stops when your system asks for reduced motion.

Pinned Loading

  1. flywheel flywheel Public

    Self-hostable, model-agnostic AI workstation and coding harness. Run any model (frontier or local) behind one OpenAI-compatible surface, run a gated coding agent over your folders, and keep a proof…

    Python 3

  2. crucible crucible Public

    Python and MCP tools for testing falsifiable claims and recording MATCH, DRIFT, or UNVERIFIABLE outcomes against supplied evidence.

    Python 3

  3. index index Public

    Offline repository and workspace maps, wikis, dependency graphs, and context artifacts with file:line provenance, zero runtime dependencies, and MCP.

    Python 4

  4. gather gather Public

    Research intake that reaches the hard places: web, video, papers, scanned PDFs, browser, OCR, and audio into structured research packets. DOM extraction and change tracking built in; provenance rid…

    Python 4

  5. buildlang buildlang Public

    A systems language with typed capability effects, sum types, C FFI, native binaries through its production C path, and HLSL/GLSL output. Linear types and non-C backends remain experimental. Ships a…

    Rust 5

  6. brender-archival brender-archival Public

    Public-safe BRender archival and revival tooling with native Win32 receipts, deterministic media, and bounded release evidence.

    Python 1