
claude-harness
A daily-driver Claude Code harness, mirrored public: incident-born guard hooks, enforced model routing, zero-dependency tools, and a nightly self-reflection…
Install with your AI
Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.
Install and set up claude-harness (git-clone project) into my current project. Found on https://claudeers.com/claude-harness Repo: https://github.com/ucsandman/claude-harness Homepage/docs: — Detected install method: git-clone → git clone https://github.com/ucsandman/claude-harness Category: devtools. Platforms: cli, api, desktop, web. Read the repo's README for exact setup and env vars, then install it and wire it into my project. Claudeers Health Verdict: active; community-verified: false. Confirm the source before running anything.
git clone https://github.com/ucsandman/claude-harness
// compatibility
| Platforms | cli, api, desktop, web |
|---|---|
| Operating systems | — |
| AI compatibility | claude |
| License | MIT |
| Pricing | open-source |
| Language | JavaScript |
claude-harness
The Claude Code setup I run every day, mirrored public. Incident-born guard hooks, a model-routing policy that is enforced rather than suggested, a dozen small zero-dependency tools, and a nightly self-reflection loop that promotes observations into rules only when the evidence earns it.
This is not a starter kit designed in an afternoon. It grew rule by rule out of real incidents over six months of daily use, and most of it exists because something broke first. The private repo that backs it also holds the agent's memory; that stays on the machine. Everything here was swept file by file before publishing (see Security).
Contents
- Why it exists
- How it fits together
- Guards
- Model routing and delegation
- Tools
- Scheduled jobs
- Subagents
- The meditation ladder
- Layout
- Stealing pieces
- What is templated, what is not here
- Security
- Contributing, license, support
Why it exists
Three examples of what "incident-born" means:
hooks/process-kill-guard.cjsexists because a subagent cleaned up its test window withStop-Process -Name notepadand killed a real Notepad session with about 40 tabs of unsaved work. Name-based process kills are now blocked at the tool layer. PID-based kills still work.hooks/agent-model-guard.cjsexists because one workflow spawned 110 agents on the most expensive model through inherited defaults and burned a full five-hour usage window. Every agent spawn now requires an explicit model, and the expensive tier is capped per session.git-hooks/pre-commitrunshooks/secret-guard.cjsover every staged file in every repo on the machine, whatever the language, because keys were once found sitting in plaintext on disk.- The four-round fix session of 2026-09-03 is written up in docs/postmortem-2026-09-03-declick-launch.md: what a file-scoped fix workflow, a capped advisor, an over-tight edit budget and a missing git guard cost on one launch day, and what each became.
hooks/git-tree-guard.cjsexists because a reviewer in a 17-agent fix workflow rangit stashto watch a test fail without its fix, the pop conflicted on a sibling's edit, and a 25-file fix pass sat silently reverted under six agents still working. Stash, path checkouts, restore, hard reset and clean are now denied in shell calls; baselines come from copies (git show HEAD:<path>into a scratch dir).workflows/fix-findings.jssays the same thing in every prompt it sends.
The design stance behind all of it: a rule written in prose is a hope. A rule that matters gets a hook, the hook gets a probe that makes it fail on purpose, and the probe runs on a schedule. A guard that has never been observed failing has been installed, not verified.
How it fits together
flowchart LR
subgraph session["Claude Code session"]
S[SessionStart] --> P[UserPromptSubmit]
P --> T{Tool call}
T -->|PreToolUse| G[Guards]
G -->|allow| X[Tool runs]
G -->|deny + reason| P
X -->|PostToolUse| A[Advisory hooks]
A --> P
P -->|Stop| L[Telemetry]
end
C[CLAUDE.md<br/>global agreement] -.loaded.-> S
SOUL[SOUL.md<br/>identity] -.read first.-> S
L --> M[Nightly meditation]
M -->|observation → fact → rule → trait| C
M --> SOUL
X -->|git commit| H[pre-commit chain<br/>secret scan, doc gates, lint]
Two layers do the work. Hooks sit on Claude Code's lifecycle events and either deny a call with a reason the model can act on, or inject context the model would otherwise forget. The ladder runs once a night, reads the telemetry and the day's work, and decides whether anything earned promotion into a standing rule. Everything else in the repo is tooling in service of those two.
Guards
Every guard is a single file, no dependencies, wired in settings.json. Each
carries an override marker so a deliberate exception is one comment away and
gets logged, rather than a reason to switch the guard off.
| Hook | Event | What it does | Override |
|---|---|---|---|
secret-guard.cjs | PreToolUse, pre-commit | Scans tool inputs and every staged file for key shapes, key=value secrets and env files. Placeholder-aware. | none |
output-secret-watch.cjs | MessageDisplay | The only output-side guard: watches what the model prints for the same shapes. | none |
process-kill-guard.cjs | PreToolUse | Denies name-based process termination (Stop-Process -Name, taskkill /IM, pkill, killall) and dynamic invocations carrying -Name. PID forms pass. | KILL_BY_NAME_OK |
agent-model-guard.cjs | PreToolUse | Every Agent, Task or Workflow spawn must name its model. Caps top-tier spawns per session. In Workflow scripts the top tier may only appear as a top-level synthesizer after the fan-out. | AGENT_GUARD_FABLE_CAP |
capability-graph-guard.cjs | PreToolUse, SubagentStart/Stop | Models delegate downward only (Fable → Opus → Sonnet → Haiku). Peers are not edges. The advisor agent is always placed one rung above its caller. | CAPABILITY_GRAPH_GUARD=off |
fable-delegate-guard.cjs | PreToolUse, SessionStart, UserPromptSubmit | When the main loop runs on the top model, budgets its direct edits and denies shell code-writing, so implementation goes to cheaper subagents. Decisions, review and synthesis stay. | # FABLE_OK: <why> |
batch-guard.cjs | PreToolUse | Denies the fourth consecutive single-statement shell call or single Read/Glob/Grep. Profiling showed one call per turn was the largest single cost. | # SEQ: <dependency> |
slow-command-guard.cjs | PreToolUse | Denies backgrounded finite test runs and recursive grep/find rooted at a projects dir, home or a drive. Both hang sessions. | BG_TEST_OK, SLOW_OK |
git-tree-guard.cjs | PreToolUse | Denies git stash, checkout/restore of paths, reset --hard, clean in shell calls: a working tree that several agents edit is read-only to git. Reads, branches and commits pass. Prose mentioning the words is not a hit. | # GIT_TREE_OK: <why> |
dev-server-guard.cjs | PreToolUse, PostToolUse | Denies dev servers piped through head/tail or backgrounded without an explicit opt-in; reminds that stopping a wrapper on Windows leaves the children alive. | DEV_SERVER_BG_OK |
scope-lock.cjs | PreToolUse, UserPromptSubmit | Confines Edit/Write to a directory for the session. Arm with scope-lock <dir> as a prompt. | scope-unlock |
repeat-tool-guard.cjs | PostToolUse | Counts identical consecutive calls and escalates a reminder at 3, 5 and 8. Advisory, never blocks. | REPEAT_GUARD_OFF=1 |
no-auto-compact.cjs | PreCompact | Turns "never auto-compact, ask at 80%" from prose into a hook. | none |
context-nudge.py | UserPromptSubmit | One nudge per high-context crossing to consider /compact or /clear. | none |
opus-handoff-inject.cjs | SessionStart, UserPromptSubmit | Detects an Opus session and injects the lower-cost operating notes once. | none |
creds-resolve.cjs | SessionStart | Fills .env from a local vault when .env.example exists, so the agent never asks for a key it already has. | none |
session-count.py | SessionStart | Warns when several sessions share one rate limit. | none |
guard-canary.ps1 | SessionStart, ~20h | Makes each guard fail on purpose and confirms it blocks. | none |
correction-tracker.ps1 | UserPromptSubmit | Buckets user corrections so a repeated one surfaces as a rule candidate instead of waiting for a human to notice. | none |
skill-telemetry.py | Stop | One JSONL record per turn: skills, agents, MCP servers, tools, tokens. The ladder's data layer. | none |
lsp-reaper.ps1 | scheduled | Kills orphaned TypeScript language servers (once found 663 of them holding 13.6 GB). | none |
hooks/tests/ holds the probes. hooks/adapters/ wires the same files into
Codex and Antigravity so there is one guard suite, not three copies
(docs/harness-parity.md). Mechanism and incident
history for each guard: docs/harness-guards.md.
Model routing and delegation
The global agreement routes work by cost and the hooks enforce it.
flowchart TD
F[Fable<br/>decisions, review, synthesis] --> O[Opus<br/>planning, orchestration, hard debugging]
O --> S[Sonnet<br/>implementation, exploration]
S --> H[Haiku<br/>lookups, mechanical edits]
S -. advisor .-> O
O -. advisor .-> F
- Downward only. A spawn that crosses a missing edge is denied. A fork inherits its caller's model, so it counts as a peer edge from any subagent.
- Upward is consultation, not delegation. A worker that hits an
architecture choice, a security boundary, or a second failed fix spawns
agents/advisor.md. The guard ignores any model it asks for and places it one rung up. Guidance comes back; ownership stays with the caller. - The economics are measured, not assumed. A subagent costs roughly 60k input tokens before its first tool call, then 2-4k per call. Under ten calls or eighty edited lines, doing it inline is cheaper on any model. The numbers and the method are in tools/tokflow/AUDIT-2026-09-02.md.
- Lean agent types by default.
haiku-scout,sonnet-implementerandopus-ownercarry restricted tool sets and cost about a third of a general-purpose spawn.
Tools
Zero-dependency, one directory each, each with its own README.
| Tool | What it does |
|---|---|
gates | Mechanical checks over the harness's own docs, hooks and skills: link rot, hook wiring, declared-vs-actual counts. Runs --staged in pre-commit. |
prove | Automates "a check never observed failing has been run, not verified": breaks the watched thing, confirms red, restores, confirms green. |
spend | Token and dollar ledger from local transcripts, per session and per day. |
tokflow | Transcript miner behind the token audit: where the fixed cost per turn actually goes. |
recall | One search across every institutional-memory store on the machine. |
skillfind | Finds any skill on the machine, including the ones no session can see. |
fleet | Live board of running Claude Code sessions. |
gitradar | Status board of every git repo on the machine: dirty trees, unpushed commits, stale branches. |
cronwatch | Health board for Windows Task Scheduler jobs: last run, last result, next due. |
envdoctor | Read-only checkup of secrets wiring. Reports names and locations, never values. |
procledger | Every process an agent starts gets a PID entry; cleanup is by PID, never by name. |
errorlog | Two daily error logs: one harvests itself from transcripts, one you type into. |
deskclaw | Read-only eye on the Windows desktop for native apps and dialogs; a hand only when a human arms it. Redacts before anything reaches a transcript. |
ears | Hears any audio or video file and returns a transcript. |
mouth | Minimal Windows text-to-speech so a long job can say it finished. |
harness-sync | Generates AGENTS.md and GEMINI.md from CLAUDE.md and reports parity across the three harnesses. |
Scheduled jobs
| Job | Cadence | What it does |
|---|---|---|
meditation | nightly | The reflection loop. See the ladder. |
fleet-briefing | daily, 7am | Collects overnight deploys, CI, revenue and error-tracker issues into one HTML board and emails it. |
errorlog | daily, before meditation | Harvests errors out of the day's transcripts so the meditation session can read them. |
harness-audit | weekly | A headless agent reads the harness itself in an isolated worktree and reports drift: registered hooks that do not dispatch, docs that lie, counts that are wrong. |
deploy-sentinel | every 30 min | Polls deploy states and CI, opens an incident exactly once per new failure. |
costclaw-watchdog | nightly | Diffs cumulative spend against last night, alerts on the incident shapes that once cost a real bill, renders a 30-night trend. |
launch-board | on demand | Local web console: per-project launch readiness with re-check buttons. |
harness-health.ps1 | on demand | Read-only check of hooks, MCP servers and plugins after any settings change. |
Subagents
| Agent | Model | Role |
|---|---|---|
haiku-scout | Haiku | Mechanical lookups: file searches, symbol hunting, inventory tables, git history. |
sonnet-implementer | Sonnet | Feature slices and refactors within a defined scope. Given files, acceptance criteria and a verify command. |
opus-owner | Opus | A large or risky task the main loop has scoped. May delegate downward. |
advisor | one rung above the caller | One focused decision. Read-only, guidance only, never capped: a blocked consultation becomes a guess, and a guess costs more than the advice. |
security-reviewer | Opus | Read-only review of anything touching auth, billing, secrets, webhooks or database access. Findings only, never edits. |
The meditation ladder
The part people ask about most. A scheduled session runs every morning, reflects on recent work, and appends dated observations. Ideas climb a ladder:
observation → fact (memory) → rule (CLAUDE.md) → trait (SOUL.md)
Each rung has explicit graduation gates. A rule needs three or more signals across two or more distinct sessions, with signals older than thirty days counting half. Every promotion cites the dated evidence that earned it. The ladder runs both ways: one contradiction is recorded, two demote. Failure lessons are written as evidence ("when X broke, Y fixed it"), not commands, so a hostile input cannot become a standing rule in one session.
The design goal is that refusing a promotion is the normal outcome. A gate that
has never once refused anything is not a gate. The gates, the demotion path
and the write rails for SOUL.md are in
meditations/MEDITATIONS.md.
Layout
| Path | What it is |
|---|---|
CLAUDE.md | The global working agreement, loaded into every session. Generated from a single source so Claude Code, Codex and Antigravity read the same text. |
SOUL.md | Who the agent is. Read before CLAUDE.md. Template here, see below. |
RTK.md | Notes for rtk, a Rust CLI proxy that compresses shell output 60 to 90 percent via a hook. |
settings.json | Hook wiring, permissions, env. Secrets live in a separate untracked file it points at. |
hooks/ | The guards above, their probes under tests/, and the Codex and Antigravity adapters. |
git-hooks/ | The global pre-commit chain (core.hooksPath): secret scan, staged doc gates, Python lint and dead-code gate. |
tools/ | The tools above. |
scripts/ | The scheduled jobs above. |
agents/ | The subagent definitions above. |
docs/ | Reference docs CLAUDE.md points at, plus reddit-claude-setup-share.md, a guided tour written to be pasted into Claude Code and adapted to your project. |
meditations/ | The nightly loop and the promotion ladder. Templates here, see below. |
Stealing pieces
Do not clone this expecting a turnkey install. Paths are Windows and specific to one machine. The useful move is taking one piece at a time.
One guard. Copy the file, then register it. A PreToolUse hook that prints a deny reason to stderr and exits 2 blocks the call and hands the model the reason:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash|PowerShell",
"hooks": [
{ "type": "command", "command": "node \"C:/path/to/hooks/process-kill-guard.cjs\"" }
]
}
]
}
}
The pre-commit chain. Point git at the directory once and every repo on the machine gets the secret scan:
git config --global core.hooksPath /path/to/git-hooks
The ladder. meditations/MEDITATIONS.md is self-contained. Start with an
empty CANDIDATES.md, run the nightly prompt in scripts/meditation/, and
refuse the first few promotions on purpose to see the gates work.
docs/reddit-claude-setup-share.md walks
the whole setup with a "steal this pattern" line per item.
What is templated, what is not here
SOUL.md and everything under meditations/ are structural templates in this
mirror. The real files accumulate personal and business context that stays
private. The mechanism, the gates and the write rails are here unchanged.
Not in the mirror:
projects/, the agent's memory store.skills/, personal skill definitions, some of them voice and identity material.- The profile block of the private
CLAUDE.md(memory wiring, project map, a remote devbox) and theautoModetrust-boundary block ofsettings.json. Both map my machines and business context. .secrets.envand anything else untracked.- The scheduled jobs that drive product repos and the autonomous company that runs on top of this harness. Different repos, private.
Security
The private repo never contained credentials. Before each sync this mirror is
swept file by file for key shapes, bearer tokens, credentialed URLs,
key=value secrets, emails and phone numbers, and the sweep prints the file
count beside its verdict so a clean result on zero files cannot pass as clean.
The last sync scanned 156 files; the only hits were fake keys inside
tools/deskclaw/tests/, which exist to prove the redaction works.
If you find something that should not be here, see SECURITY.md.
Contributing, license, support
Issues and pull requests are welcome, especially incident reports of the form "this guard let X through" with a probe that reproduces it. See CONTRIBUTING.md. Changes to the mirror are listed in CHANGELOG.md.
MIT. If these tools save you time:
// faq
What is claude-harness?
A daily-driver Claude Code harness, mirrored public: incident-born guard hooks, enforced model routing, zero-dependency tools, and a nightly self-reflection ladder.. It is open-source on GitHub.
Is claude-harness free to use?
claude-harness is open-source under the MIT license, so it is free to use.
What category does claude-harness belong to?
claude-harness is listed under devtools in the Claudeers registry of Claude-compatible tools.
// embed badge
[](https://claudeers.com/claude-harness)
// retro hit counter
[](https://claudeers.com/claude-harness)
// reviews
// guestbook
// related in Developer Tools
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Curs…
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs,…