claudeers.
// MCP Servers

HarnessTrim

One token policy for Claude Code, Codex, OpenCode, Hermes Agent and Pi. A cross-harness token-economy control plane.

// MCP Servers[ cli ][ api ][ mobile ][ claude ]#claude#mcp-serversMIT$open-sourceupdated about 1 month ago
Actively maintained
100/100
last commit 16 days ago
last release 20 days ago
releases 7
open issues 0
// star history

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up HarnessTrim (git-clone project) into my current project.
Found on https://claudeers.com/harnesstrim
Repo: https://github.com/giuliastro/HarnessTrim
Homepage/docs: —
Detected install method: git-clone → git clone https://github.com/giuliastro/HarnessTrim
Category: mcp-servers. Platforms: cli, api, mobile.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
active; community-verified: false. Confirm the source before running anything.
// or clone
git clone https://github.com/giuliastro/HarnessTrim

// compatibility

Platformscli, api, mobile
Operating systems
AI compatibilityclaude
LicenseMIT
Pricingopen-source
LanguageTypeScript

HarnessTrim

One token policy for Claude Code, Codex, OpenCode, Hermes Agent and Pi.

HarnessTrim is a cross-harness control plane for coding agents: a portable skill pack, thin per-harness adapters, and a reproducible benchmark suite that together cut input tokens, output tokens, and noisy tool output, instead of optimizing just one of those layers the way existing tools do.

Full design rationale and phased roadmap: see PLAN.md.


The problem

In a coding agent, tokens are spent across several channels — and most tools only attack one of them:

ChannelWhat fills itWho attacks it today
Tool outputtest logs, git diff, grep, build output, big JSON, file readsRTK (shell only)
Model outputthe agent's own verbosityCaveman
Thinkingreasoning tokens, billed as outputmostly nobody
Fixed instructionsalways-loaded CLAUDE.md/AGENTS.mdskills (native)
Conversation historyeverything that survives compactioncompaction (native)

Each existing tool moves one lever. The waste is spread across all of them, so single-lever tools leave most of the budget on the table. HarnessTrim's thesis: coordinate all five levers behind one policy, using the deterministic hook/skill primitives every modern harness already exposes.

Strategy: skill-first, adapter-second, measured

Three principles, in priority order:

  1. Skill-first. The portable value is a pack of Agent Skills (the format every target harness already understands). Skills carry the policy; they cost almost nothing until invoked.
  2. Adapter-second. Thin per-harness adapters translate one shared policy into each harness's native dialect (hooks, plugins, compaction events). Adapters are where the real work is — and where fragility lives — so they stay deliberately small and delegate all logic to the shared core.
  3. Measured, not asserted. Every claim is backed by a reproducible benchmark. Competitors report self-measured numbers that don't compose; HarnessTrim ships the measurement harness itself.
flowchart TB
    subgraph Harnesses
        CC[Claude Code]
        CX[Codex]
        OC[OpenCode]
        HM[Hermes Agent]
        PI[Pi]
    end

    subgraph Adapters["Adapter layer (thin, per-harness)"]
        A1[hooks / plugins / compaction events]
    end

    subgraph Core["@harnesstrim/core (shared, deterministic)"]
        R[reducers]
        D[content dispatcher]
        P[policy presets]
        M[metrics / TrimEvent]
    end

    subgraph Skills["Portable skill pack"]
        S[delta-response · debug-log-slim · review-delta<br/>compact-handoff · scaffold-fast · delegate-bulk]
    end

    CC --- A1
    CX --- A1
    OC --- A1
    HM --- A1
    PI --- A1
    A1 --> Core
    Core --- Skills

    classDef done fill:#1f6f3f,stroke:#0d3,color:#fff;
    class CC,CX,OC,HM,PI,A1,R,D,P,M,S done;

Green = has a shipped adapter. All five targets (OpenCode, Codex, Claude Code, Hermes Agent, Pi) now have one, reusing the same core and skills.

The five levers

LeverMechanismHarnessTrim component
Progressive disclosurerecurring instructions live in on-demand skills, not always-loaded filesskill pack + doctor
Tool-output reductiona deterministic reducer slims noisy output before it reaches the modelreducers + adapter tool.execute.after
Thinking routingmatch reasoning effort to task type (low for mechanical, high for architecture)policy presets (advisory)
Subtask isolationisolate/handoff noisy work instead of polluting the main contextcompact-handoff + delegate-bulk skills, compaction hook
Observabilitynormalize what was actually saved into one schemaTrimEvent + metrics

How tool-output reduction works

The adapter intercepts tool results, the shared core decides what (if anything) to slim, and only the signal reaches the model. Reducers are deterministic and idempotent and never touch the cacheable prompt prefix — so they shrink cost without busting the prompt cache.

sequenceDiagram
    participant M as Model
    participant H as Harness
    participant A as HarnessTrim adapter
    participant C as core.reduceAuto
    M->>H: request tool call (e.g. run tests)
    H->>H: execute tool (1410 chars of output)
    H->>A: tool.execute.after(output)
    A->>C: reduceAuto(output)
    C-->>A: slimmed output + TrimEvent (124 chars)
    A-->>H: mutated output.output
    A->>A: append TrimEvent (telemetry, opt-in)
    H-->>M: slimmed output enters context

KPIs

What HarnessTrim optimizes for, and how each is measured:

KPIDefinitionTargetSource
Tool-output reduction1 − (chars out / chars in) per reduced tool call≥ 50% on noisy outputadapter telemetry, benchmark
Signal fidelity (recall)must-keep signal lines surviving reduction / total must-keep100% (bench fails otherwise)Tier A benchmark (measured now)
Blended session reductiontotal tokens saved / baseline session tokens30–50% (model)end-to-end benchmark (Tier B, planned)
Quality retentiontask-success parity vs the untrimmed baseline100% (no regressions)Tier B benchmark
Cache preservationshare of reductions that leave the cacheable prefix untouched100%design guarantee (reducers only touch volatile output)
Coverageshare of noisy tool calls that a reducer actually matchedgrow over timetelemetry (reducer: null = missed)
Overheadadded latency / tokens from the stack itselfnegligiblereducers run locally, no tokenizer in-process

Savings: measured vs hypothesized

Two honesty tiers. Keep them separate.

Measured (real numbers today)

The token number alone is not the point — a reducer that drops the one line you needed would post a great percentage and ruin the context. So the benchmark measures both: token reduction and signal fidelity — of the lines that must survive (the error, the failing test, the assertion, the changed files, the summary), how many are kept. It also audits any dropped line that looks like signal. Headline: −65% tokens at 100% signal recall across the seed fixtures (pnpm run bench, no LLM). The bench fails loudly if signal recall drops below 100% or a signal-looking line is dropped.

FixtureReducerTokensReductionSignal kept
jest, mostly-passtest-output-slim408 → 216−47.1%6/6
pytest, mostly-passtest-output-slim395 → 211−46.6%5/5
lockfile-heavy diffgit-diff-slim939 → 183−80.5%4/4
Overall1742 → 610−65%15/15 (100%)

Each fixture's must-keep lines are annotated in benchmarks/src/run.ts, so "what survives" is explicit and reproducible, not a claim.

  • One live OpenCode session: a real bash test run was reduced 1410 → 124 chars (−91.2%) in the actual pipeline, with the FAIL line and summary preserved (see PLAN.md §9, Phase 2 hardening).

These cover the tool-output lever only, on selected inputs. They are not a session-wide claim.

Hypothesized (illustrative model, not measured)

To reason about the blended win we model a "typical" medium debugging session. These percentages are an engineering hypothesis to be validated by the Tier B benchmark — not results.

Baseline budget of an illustrative session, by channel:

pie showData
    title Baseline session token budget (illustrative)
    "Tool output" : 45
    "Conversation history" : 15
    "Model output" : 15
    "Thinking" : 15
    "Instructions (fixed)" : 10

Applying a conservative per-lever reduction to each channel:

LeverChannel shareAssumed reduction of channelSaved (% of total)
Tool-output reduction45%65%29.3%
Thinking routing15%50%7.5%
Model-output discipline15%40%6.0%
Progressive disclosure10%50%5.0%
Subtask isolation15%30%4.5%
Blended≈ 52%
xychart-beta
    title "Hypothesized token savings by lever (% of total session budget)"
    x-axis ["Tool output", "Thinking", "Model output", "Instructions", "Subtask iso."]
    y-axis "Saved % of total" 0 --> 35
    bar [29.3, 7.5, 6.0, 5.0, 4.5]

Scenario range (blended reduction of total session tokens):

ScenarioAssumptionsBlended reduction
Conservativelow per-lever rates, tool output only partially matched~30%
Expectedthe table above~50%
Optimisticnoisy debugging session, high tool-output share~65%
xychart-beta
    title "Blended session reduction — hypothesized scenarios (% of total tokens)"
    x-axis ["Conservative", "Expected", "Optimistic"]
    y-axis "Reduction %" 0 --> 70
    bar [30, 50, 65]

Why the model is plausible but unproven: the tool-output lever (the largest slice) is already backed by the measured −65%/−91.2% numbers above. The other levers are extrapolated from vendor documentation on reasoning-token billing, prompt caching, and progressive disclosure. The Tier B end-to-end benchmark (planned) will replace this section's hypotheses with measured, quality-checked numbers comparing vanilla harness vs harness + HarnessTrim.


Status

Phases 0–4 in progress. Shipped: reducers + benchmark, the 6-skill pack, adapters for OpenCode (runtime plugin, hardened in a live session), Codex (skills + AGENTS.md reduce-pipe, live-validated via codex debug prompt-input), Claude Code (PostToolUse reducer hook), Hermes Agent (transform_tool_result plugin, verified in a live session), and Pi (tool_result extension), plus an MCP reduce server, the harnesstrim CLI (doctor / install / preset / metrics / reduce / hook / mcp / bench), telemetry, and policy presets. All five target harnesses now have an adapter. An end-to-end Tier B run on OpenCode confirmed the reducer cuts freshly-billed tool-output tokens ~60% without busting the prompt cache or breaking the task; broader multi-task Tier B runs are the main remaining work. 83 tests passing, typecheck clean on all packages.

Layout

packages/core/              deterministic, idempotent reducers + content dispatcher + presets + metrics
packages/adapter-opencode/  OpenCode plugin: slims tool output + injects compaction handoff + telemetry
packages/adapter-codex/     Codex: skill bundle + AGENTS.md reduce-pipe instruction
packages/adapter-claude/    Claude Code: PostToolUse reducer hook + skill bundle
packages/adapter-hermes/    Hermes Agent: transform_tool_result reducer plugin (Python)
packages/adapter-pi/        Pi: tool_result reducer extension (TypeScript)
packages/mcp/               MCP server exposing a `reduce` tool (Codex, Claude Code, any MCP client)
packages/cli/               harnesstrim CLI: doctor, install, preset, metrics, reduce, hook, mcp, bench
skills/                     portable Agent Skills (delta-response, debug-log-slim, review-delta,
                            compact-handoff, scaffold-fast, delegate-bulk)
benchmarks/                 Tier A micro-benchmarks: reducer token-reduction, no LLM involved
examples/opencode/          minimal opencode.json wiring the adapter (dry-run)

CLI

pnpm exec harnesstrim doctor [dir]            # diagnose token-waste signals in a project
pnpm exec harnesstrim install opencode [dir]  # OpenCode plugin -> opencode.json (dry-run)
pnpm exec harnesstrim install opencode --preset lean-debug --apply
pnpm exec harnesstrim install codex [dir]     # Codex: skills + AGENTS.md reduce-pipe (dry-run)
pnpm exec harnesstrim install claude [dir]    # Claude Code: skills + PostToolUse hook (dry-run)
pnpm exec harnesstrim install hermes [dir]    # Hermes Agent: transform_tool_result plugin (dry-run)
pnpm exec harnesstrim install pi [dir]        # Pi: tool_result extension (dry-run)
pnpm exec harnesstrim install hermes [dir]    # Hermes Agent: transform_tool_result plugin (dry-run)
pnpm exec harnesstrim preset list             # list policy presets
pnpm exec harnesstrim metrics [path]          # summarize adapter telemetry (JSONL)
npm test 2>&1 | pnpm exec harnesstrim reduce  # pipe: slim noisy output (Codex/shell)
pnpm exec harnesstrim bench                    # run the Tier A reducer micro-benchmark
  • doctor flags oversized always-loaded instruction files (CLAUDE.md/AGENTS.md/...), reports whether on-demand skills are used, and whether the OpenCode adapter is wired in.
  • install <harness> is dry-run until --apply. Each adapter uses that harness's native surface: OpenCode a tool.execute.after plugin, Claude Code a PostToolUse hook, Hermes a transform_tool_result plugin, Pi a tool_result extension, Codex an AGENTS.md reduce-pipe instruction. --preset (OpenCode) bakes a policy preset's adapter config in.
  • reduce is the pipe-friendly reducer (RTK-style) shared across harnesses.
  • metrics aggregates the telemetry the adapter emits (off by default) into chars saved per reducer.

Try it

pnpm install
pnpm run test        # unit tests (core reducers + dispatcher + adapter hooks)
pnpm run typecheck   # type-check every package against real dependency types
pnpm run bench       # Tier A micro-benchmark: token reduction on fixed fixtures

Using it in your harness

Each harness has a one-command installer (dry-run until --apply). First make the harnesstrim command available — until the package is published, either prefix commands with pnpm exec from this repo, or link it once:

pnpm install
pnpm --filter @harnesstrim/cli link --global   # exposes `harnesstrim` on PATH

The installer is a preview until you pass --apply: run it without --apply first to see exactly what files it would change. That is separate from each adapter's runtime reduction mode below.

Reduction mode & telemetry (per adapter)

Once installed, does the adapter actually slim output, and does it record metrics? This differs by harness. "dry-run mode" here means the adapter logs what it would slim without changing anything.

HarnessReduces after install?Make reduction permanentTelemetry (metrics)
OpenCodeYes — plugin mode defaults to activealready permanent in opencode.json; set "mode": "dryrun" there to only previewoff; set plugin option "telemetry": true (+ optional telemetryPath), read with harnesstrim metrics <path>
Claude CodeYes — the PostToolUse hook reduces once loaded (no dry-run mode)permanent once in .claude/settings.jsonnone (the hook does not emit metrics)
CodexNo automatic reduction — the model pipes through harnesstrim reduce or calls the MCP reduce toolpermanent (the AGENTS.md instruction / MCP registration persists)via the MCP/pipe path, not an adapter emitter
HermesNo — starts in dryrunset HARNESSTRIM_MODE=active persistently (your shell profile or Hermes' service environment, not a one-off export)off; HARNESSTRIM_TELEMETRY=1 writes ~/.hermes/harnesstrim-metrics.jsonl, read with harnesstrim metrics ~/.hermes/harnesstrim-metrics.jsonl
PiNo — starts in dryrunset HARNESSTRIM_MODE=active persistently in Pi's environmentnone yet (the extension only reduces)

Guidance: for the dry-run adapters (Hermes, Pi) keep the default while you confirm it slims the right things (watch stderr for [harnesstrim] dryrun ... lines), then flip to active persistently. Telemetry is off by default everywhere; enable it only where you want a metrics trail.

OpenCode

harnesstrim install opencode /path/to/project --apply

Wires the plugin into opencode.json. It reduces tool output automatically via tool.execute.after, no per-command action needed. The plugin defaults to "mode": "active" (reduces immediately); set "mode": "dryrun" in the plugin options first if you want to preview before enabling. Details: packages/adapter-opencode, example: examples/opencode.

Codex

harnesstrim install codex /path/to/project --apply

Copies the skill pack into .codex/skills and adds a reduce-pipe instruction to AGENTS.md. The agent then slims noisy output by piping it (pytest 2>&1 | harnesstrim reduce), so harnesstrim must be on PATH. For a first-class, native tool instead of a shell pipe, register the MCP reducer:

codex mcp add harnesstrim -- harnesstrim mcp

Details: packages/adapter-codex, packages/mcp.

Claude Code

harnesstrim install claude /path/to/project --apply

Copies the skill pack into .claude/skills and adds a PostToolUse hook (matched to Bash) to .claude/settings.json. The hook runs harnesstrim hook claude, so harnesstrim must be on PATH; reload Claude Code so the hook loads. It then slims noisy Bash output automatically before the model sees it. Details: packages/adapter-claude.

Hermes Agent

harnesstrim install hermes --apply                    # ~/.hermes/plugins/harnesstrim/
harnesstrim install hermes /path/to/project --apply   # project-local .hermes/plugins/

Copies a Python plugin that hooks Hermes' transform_tool_result and slims terminal output before it enters context (it shells out to harnesstrim reduce, so harnesstrim must be on PATH). After installing, enable it in ~/.hermes/config.yaml:

plugins:
  enabled:
    - harnesstrim

Restart Hermes. It starts in dryrun (logs to stderr what it would slim); set HARNESSTRIM_MODE=active in Hermes' environment to actually reduce. Details: packages/adapter-hermes.

Pi

harnesstrim install pi --apply             # <project>/.pi/extensions/harnesstrim/
harnesstrim install pi ~ --apply           # global: ~/.pi/... (pass your home dir)

Copies a TypeScript extension that hooks Pi's tool_result and slims noisy output via harnesstrim reduce (so harnesstrim must be on PATH). It starts in dryrun; set HARNESSTRIM_MODE=active in Pi's environment to reduce. Details: packages/adapter-pi.

Any MCP-capable harness

harnesstrim mcp starts a stdio MCP server exposing a reduce tool. Register it with any client that speaks MCP (Codex, Claude Code, …). See packages/mcp.

License

MIT — see LICENSE.

// faq

What is HarnessTrim?

One token policy for Claude Code, Codex, OpenCode, Hermes Agent and Pi. A cross-harness token-economy control plane.. It is open-source on GitHub.

Is HarnessTrim free to use?

HarnessTrim is open-source under the MIT license, so it is free to use.

What category does HarnessTrim belong to?

HarnessTrim is listed under mcp-servers in the Claudeers registry of Claude-compatible tools.

1 views
12 stars
unclaimed
updated about 1 month ago

// embed badge

HarnessTrim on Claudeers
[![Claudeers](https://claudeers.com/api/badge/harnesstrim.svg)](https://claudeers.com/harnesstrim)

// retro hit counter

HarnessTrim hit counter
[![Hits](https://claudeers.com/api/counter/harnesstrim.svg)](https://claudeers.com/harnesstrim)

// reviews

// guestbook

0/500

// related in MCP Servers

🔓

f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete…

// mcp-serversf/HTML167,135NOASSERTION[ claude ]
🔓

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Gemini CLI & Hermes Agent. Only official website: ccswitch.io

// mcp-serversfarion1231/Rust127,274MIT[ claude ]
🔓

An open-source AI agent that brings the power of Gemini directly into your terminal.

// mcp-serversgoogle-gemini/TypeScript106,524Apache-2.0[ claude ]
🔓

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

// mcp-serversJuliusBrussee/JavaScript100,343MIT[ claude ]
→ see how HarnessTrim connects across the ecosystem