claudeers.
// Developer Tools

claude-harness

A daily-driver Claude Code harness, mirrored public: incident-born guard hooks, enforced model routing, zero-dependency tools, and a nightly self-reflection…

// Developer Tools[ cli ][ api ][ desktop ][ web ][ claude ]#claude#agent-guardrails#ai-agents#anthropic#claude-code#developer-tools#hooks#llm-ops#devtools◷ MIT$open-sourceupdated about 1 month ago
Actively maintained
100/100
last commit 9 days ago
last release none
releases 0
open issues 0
// star history

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up claude-harness (git-clone project) into my current project.
Found on https://claudeers.com/claude-harness
Repo: https://github.com/ucsandman/claude-harness
Homepage/docs: —
Detected install method: git-clone → git clone https://github.com/ucsandman/claude-harness
Category: devtools. Platforms: cli, api, desktop, web.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
active; community-verified: false. Confirm the source before running anything.
// or clone
git clone https://github.com/ucsandman/claude-harness

// compatibility

Platformscli, api, desktop, web
Operating systems—
AI compatibilityclaude
LicenseMIT
Pricingopen-source
LanguageJavaScript

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

claude-harness

The Claude Code setup I run every day, mirrored public. Incident-born guard hooks, a model-routing policy that is enforced rather than suggested, a dozen small zero-dependency tools, and a nightly self-reflection loop that promotes observations into rules only when the evidence earns it.

This is not a starter kit designed in an afternoon. It grew rule by rule out of real incidents over six months of daily use, and most of it exists because something broke first. The private repo that backs it also holds the agent's memory; that stays on the machine. Everything here was swept file by file before publishing (see Security).


Contents


Why it exists

Three examples of what "incident-born" means:

  • hooks/process-kill-guard.cjs exists because a subagent cleaned up its test window with Stop-Process -Name notepad and killed a real Notepad session with about 40 tabs of unsaved work. Name-based process kills are now blocked at the tool layer. PID-based kills still work.
  • hooks/agent-model-guard.cjs exists because one workflow spawned 110 agents on the most expensive model through inherited defaults and burned a full five-hour usage window. Every agent spawn now requires an explicit model, and the expensive tier is capped per session.
  • git-hooks/pre-commit runs hooks/secret-guard.cjs over every staged file in every repo on the machine, whatever the language, because keys were once found sitting in plaintext on disk.
  • The four-round fix session of 2026-09-03 is written up in docs/postmortem-2026-09-03-declick-launch.md: what a file-scoped fix workflow, a capped advisor, an over-tight edit budget and a missing git guard cost on one launch day, and what each became.
  • hooks/git-tree-guard.cjs exists because a reviewer in a 17-agent fix workflow ran git stash to watch a test fail without its fix, the pop conflicted on a sibling's edit, and a 25-file fix pass sat silently reverted under six agents still working. Stash, path checkouts, restore, hard reset and clean are now denied in shell calls; baselines come from copies (git show HEAD:<path> into a scratch dir). workflows/fix-findings.js says the same thing in every prompt it sends.

The design stance behind all of it: a rule written in prose is a hope. A rule that matters gets a hook, the hook gets a probe that makes it fail on purpose, and the probe runs on a schedule. A guard that has never been observed failing has been installed, not verified.

How it fits together

flowchart LR
    subgraph session["Claude Code session"]
        S[SessionStart] --> P[UserPromptSubmit]
        P --> T{Tool call}
        T -->|PreToolUse| G[Guards]
        G -->|allow| X[Tool runs]
        G -->|deny + reason| P
        X -->|PostToolUse| A[Advisory hooks]
        A --> P
        P -->|Stop| L[Telemetry]
    end
    C[CLAUDE.md<br/>global agreement] -.loaded.-> S
    SOUL[SOUL.md<br/>identity] -.read first.-> S
    L --> M[Nightly meditation]
    M -->|observation → fact → rule → trait| C
    M --> SOUL
    X -->|git commit| H[pre-commit chain<br/>secret scan, doc gates, lint]

Two layers do the work. Hooks sit on Claude Code's lifecycle events and either deny a call with a reason the model can act on, or inject context the model would otherwise forget. The ladder runs once a night, reads the telemetry and the day's work, and decides whether anything earned promotion into a standing rule. Everything else in the repo is tooling in service of those two.

Guards

Every guard is a single file, no dependencies, wired in settings.json. Each carries an override marker so a deliberate exception is one comment away and gets logged, rather than a reason to switch the guard off.

HookEventWhat it doesOverride
secret-guard.cjsPreToolUse, pre-commitScans tool inputs and every staged file for key shapes, key=value secrets and env files. Placeholder-aware.none
output-secret-watch.cjsMessageDisplayThe only output-side guard: watches what the model prints for the same shapes.none
process-kill-guard.cjsPreToolUseDenies name-based process termination (Stop-Process -Name, taskkill /IM, pkill, killall) and dynamic invocations carrying -Name. PID forms pass.KILL_BY_NAME_OK
agent-model-guard.cjsPreToolUseEvery Agent, Task or Workflow spawn must name its model. Caps top-tier spawns per session. In Workflow scripts the top tier may only appear as a top-level synthesizer after the fan-out.AGENT_GUARD_FABLE_CAP
capability-graph-guard.cjsPreToolUse, SubagentStart/StopModels delegate downward only (Fable → Opus → Sonnet → Haiku). Peers are not edges. The advisor agent is always placed one rung above its caller.CAPABILITY_GRAPH_GUARD=off
fable-delegate-guard.cjsPreToolUse, SessionStart, UserPromptSubmitWhen the main loop runs on the top model, budgets its direct edits and denies shell code-writing, so implementation goes to cheaper subagents. Decisions, review and synthesis stay.# FABLE_OK: <why>
batch-guard.cjsPreToolUseDenies the fourth consecutive single-statement shell call or single Read/Glob/Grep. Profiling showed one call per turn was the largest single cost.# SEQ: <dependency>
slow-command-guard.cjsPreToolUseDenies backgrounded finite test runs and recursive grep/find rooted at a projects dir, home or a drive. Both hang sessions.BG_TEST_OK, SLOW_OK
git-tree-guard.cjsPreToolUseDenies git stash, checkout/restore of paths, reset --hard, clean in shell calls: a working tree that several agents edit is read-only to git. Reads, branches and commits pass. Prose mentioning the words is not a hit.# GIT_TREE_OK: <why>
dev-server-guard.cjsPreToolUse, PostToolUseDenies dev servers piped through head/tail or backgrounded without an explicit opt-in; reminds that stopping a wrapper on Windows leaves the children alive.DEV_SERVER_BG_OK
scope-lock.cjsPreToolUse, UserPromptSubmitConfines Edit/Write to a directory for the session. Arm with scope-lock <dir> as a prompt.scope-unlock
repeat-tool-guard.cjsPostToolUseCounts identical consecutive calls and escalates a reminder at 3, 5 and 8. Advisory, never blocks.REPEAT_GUARD_OFF=1
no-auto-compact.cjsPreCompactTurns "never auto-compact, ask at 80%" from prose into a hook.none
context-nudge.pyUserPromptSubmitOne nudge per high-context crossing to consider /compact or /clear.none
opus-handoff-inject.cjsSessionStart, UserPromptSubmitDetects an Opus session and injects the lower-cost operating notes once.none
creds-resolve.cjsSessionStartFills .env from a local vault when .env.example exists, so the agent never asks for a key it already has.none
session-count.pySessionStartWarns when several sessions share one rate limit.none
guard-canary.ps1SessionStart, ~20hMakes each guard fail on purpose and confirms it blocks.none
correction-tracker.ps1UserPromptSubmitBuckets user corrections so a repeated one surfaces as a rule candidate instead of waiting for a human to notice.none
skill-telemetry.pyStopOne JSONL record per turn: skills, agents, MCP servers, tools, tokens. The ladder's data layer.none
lsp-reaper.ps1scheduledKills orphaned TypeScript language servers (once found 663 of them holding 13.6 GB).none

hooks/tests/ holds the probes. hooks/adapters/ wires the same files into Codex and Antigravity so there is one guard suite, not three copies (docs/harness-parity.md). Mechanism and incident history for each guard: docs/harness-guards.md.

Model routing and delegation

The global agreement routes work by cost and the hooks enforce it.

flowchart TD
    F[Fable<br/>decisions, review, synthesis] --> O[Opus<br/>planning, orchestration, hard debugging]
    O --> S[Sonnet<br/>implementation, exploration]
    S --> H[Haiku<br/>lookups, mechanical edits]
    S -. advisor .-> O
    O -. advisor .-> F
  • Downward only. A spawn that crosses a missing edge is denied. A fork inherits its caller's model, so it counts as a peer edge from any subagent.
  • Upward is consultation, not delegation. A worker that hits an architecture choice, a security boundary, or a second failed fix spawns agents/advisor.md. The guard ignores any model it asks for and places it one rung up. Guidance comes back; ownership stays with the caller.
  • The economics are measured, not assumed. A subagent costs roughly 60k input tokens before its first tool call, then 2-4k per call. Under ten calls or eighty edited lines, doing it inline is cheaper on any model. The numbers and the method are in tools/tokflow/AUDIT-2026-09-02.md.
  • Lean agent types by default. haiku-scout, sonnet-implementer and opus-owner carry restricted tool sets and cost about a third of a general-purpose spawn.

Tools

Zero-dependency, one directory each, each with its own README.

ToolWhat it does
gatesMechanical checks over the harness's own docs, hooks and skills: link rot, hook wiring, declared-vs-actual counts. Runs --staged in pre-commit.
proveAutomates "a check never observed failing has been run, not verified": breaks the watched thing, confirms red, restores, confirms green.
spendToken and dollar ledger from local transcripts, per session and per day.
tokflowTranscript miner behind the token audit: where the fixed cost per turn actually goes.
recallOne search across every institutional-memory store on the machine.
skillfindFinds any skill on the machine, including the ones no session can see.
fleetLive board of running Claude Code sessions.
gitradarStatus board of every git repo on the machine: dirty trees, unpushed commits, stale branches.
cronwatchHealth board for Windows Task Scheduler jobs: last run, last result, next due.
envdoctorRead-only checkup of secrets wiring. Reports names and locations, never values.
procledgerEvery process an agent starts gets a PID entry; cleanup is by PID, never by name.
errorlogTwo daily error logs: one harvests itself from transcripts, one you type into.
deskclawRead-only eye on the Windows desktop for native apps and dialogs; a hand only when a human arms it. Redacts before anything reaches a transcript.
earsHears any audio or video file and returns a transcript.
mouthMinimal Windows text-to-speech so a long job can say it finished.
harness-syncGenerates AGENTS.md and GEMINI.md from CLAUDE.md and reports parity across the three harnesses.

Scheduled jobs

JobCadenceWhat it does
meditationnightlyThe reflection loop. See the ladder.
fleet-briefingdaily, 7amCollects overnight deploys, CI, revenue and error-tracker issues into one HTML board and emails it.
errorlogdaily, before meditationHarvests errors out of the day's transcripts so the meditation session can read them.
harness-auditweeklyA headless agent reads the harness itself in an isolated worktree and reports drift: registered hooks that do not dispatch, docs that lie, counts that are wrong.
deploy-sentinelevery 30 minPolls deploy states and CI, opens an incident exactly once per new failure.
costclaw-watchdognightlyDiffs cumulative spend against last night, alerts on the incident shapes that once cost a real bill, renders a 30-night trend.
launch-boardon demandLocal web console: per-project launch readiness with re-check buttons.
harness-health.ps1on demandRead-only check of hooks, MCP servers and plugins after any settings change.

Subagents

AgentModelRole
haiku-scoutHaikuMechanical lookups: file searches, symbol hunting, inventory tables, git history.
sonnet-implementerSonnetFeature slices and refactors within a defined scope. Given files, acceptance criteria and a verify command.
opus-ownerOpusA large or risky task the main loop has scoped. May delegate downward.
advisorone rung above the callerOne focused decision. Read-only, guidance only, never capped: a blocked consultation becomes a guess, and a guess costs more than the advice.
security-reviewerOpusRead-only review of anything touching auth, billing, secrets, webhooks or database access. Findings only, never edits.

The meditation ladder

The part people ask about most. A scheduled session runs every morning, reflects on recent work, and appends dated observations. Ideas climb a ladder:

observation  →  fact (memory)  →  rule (CLAUDE.md)  →  trait (SOUL.md)

Each rung has explicit graduation gates. A rule needs three or more signals across two or more distinct sessions, with signals older than thirty days counting half. Every promotion cites the dated evidence that earned it. The ladder runs both ways: one contradiction is recorded, two demote. Failure lessons are written as evidence ("when X broke, Y fixed it"), not commands, so a hostile input cannot become a standing rule in one session.

The design goal is that refusing a promotion is the normal outcome. A gate that has never once refused anything is not a gate. The gates, the demotion path and the write rails for SOUL.md are in meditations/MEDITATIONS.md.

Layout

PathWhat it is
CLAUDE.mdThe global working agreement, loaded into every session. Generated from a single source so Claude Code, Codex and Antigravity read the same text.
SOUL.mdWho the agent is. Read before CLAUDE.md. Template here, see below.
RTK.mdNotes for rtk, a Rust CLI proxy that compresses shell output 60 to 90 percent via a hook.
settings.jsonHook wiring, permissions, env. Secrets live in a separate untracked file it points at.
hooks/The guards above, their probes under tests/, and the Codex and Antigravity adapters.
git-hooks/The global pre-commit chain (core.hooksPath): secret scan, staged doc gates, Python lint and dead-code gate.
tools/The tools above.
scripts/The scheduled jobs above.
agents/The subagent definitions above.
docs/Reference docs CLAUDE.md points at, plus reddit-claude-setup-share.md, a guided tour written to be pasted into Claude Code and adapted to your project.
meditations/The nightly loop and the promotion ladder. Templates here, see below.

Stealing pieces

Do not clone this expecting a turnkey install. Paths are Windows and specific to one machine. The useful move is taking one piece at a time.

One guard. Copy the file, then register it. A PreToolUse hook that prints a deny reason to stderr and exits 2 blocks the call and hands the model the reason:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash|PowerShell",
        "hooks": [
          { "type": "command", "command": "node \"C:/path/to/hooks/process-kill-guard.cjs\"" }
        ]
      }
    ]
  }
}

The pre-commit chain. Point git at the directory once and every repo on the machine gets the secret scan:

git config --global core.hooksPath /path/to/git-hooks

The ladder. meditations/MEDITATIONS.md is self-contained. Start with an empty CANDIDATES.md, run the nightly prompt in scripts/meditation/, and refuse the first few promotions on purpose to see the gates work.

docs/reddit-claude-setup-share.md walks the whole setup with a "steal this pattern" line per item.

What is templated, what is not here

SOUL.md and everything under meditations/ are structural templates in this mirror. The real files accumulate personal and business context that stays private. The mechanism, the gates and the write rails are here unchanged.

Not in the mirror:

  • projects/, the agent's memory store.
  • skills/, personal skill definitions, some of them voice and identity material.
  • The profile block of the private CLAUDE.md (memory wiring, project map, a remote devbox) and the autoMode trust-boundary block of settings.json. Both map my machines and business context.
  • .secrets.env and anything else untracked.
  • The scheduled jobs that drive product repos and the autonomous company that runs on top of this harness. Different repos, private.

Security

The private repo never contained credentials. Before each sync this mirror is swept file by file for key shapes, bearer tokens, credentialed URLs, key=value secrets, emails and phone numbers, and the sweep prints the file count beside its verdict so a clean result on zero files cannot pass as clean. The last sync scanned 156 files; the only hits were fake keys inside tools/deskclaw/tests/, which exist to prove the redaction works.

If you find something that should not be here, see SECURITY.md.

Contributing, license, support

Issues and pull requests are welcome, especially incident reports of the form "this guard let X through" with a probe that reproduces it. See CONTRIBUTING.md. Changes to the mirror are listed in CHANGELOG.md.

MIT. If these tools save you time:

// faq

What is claude-harness?

A daily-driver Claude Code harness, mirrored public: incident-born guard hooks, enforced model routing, zero-dependency tools, and a nightly self-reflection ladder.. It is open-source on GitHub.

Is claude-harness free to use?

claude-harness is open-source under the MIT license, so it is free to use.

What category does claude-harness belong to?

claude-harness is listed under devtools in the Claudeers registry of Claude-compatible tools.

13 views
★ 24 stars
unclaimed
updated about 1 month ago

// embed badge

claude-harness on Claudeers
[![Claudeers](https://claudeers.com/api/badge/claude-harness.svg)](https://claudeers.com/claude-harness)

// retro hit counter

claude-harness hit counter
[![Hits](https://claudeers.com/api/counter/claude-harness.svg)](https://claudeers.com/claude-harness)

// reviews

// guestbook

0/500

// related in Developer Tools

🔓

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Curs…

// devtoolsaffaan-m/⟨JavaScript⟩★ 267,519◷ MIT[ claude ]
🔓

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

// devtoolsDietrichGebert/⟨JavaScript⟩★ 148,251◷ MIT[ claude ]
🔓

Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA

// devtoolsgarrytan/⟨TypeScript⟩★ 134,274◷ MIT[ claude ]
🔓

AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs,…

// devtoolssafishamsi/⟨Python⟩★ 123,348◷ MIT[ claude ]
→ see how claude-harness connects across the ecosystem