claudeers.
// RAG & Knowledge

gobstopper

Gobstopper makes long Claude Code and Codex sessions smaller. A local proxy compacts live requests, and saved sessions get a smaller copy beside the original.

// RAG & Knowledge[ cli ][ api ][ desktop ][ web ][ claude ]#claude#ai-agents#anthropic#claude-code#cli#codex#coding-agents#context-compaction#rag◷ Apache-2.0$open-sourceupdated 4 days ago

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up gobstopper (release-binary project) into my current project.
Found on https://claudeers.com/gobstopper
Repo: https://github.com/hraness/gobstopper
Homepage/docs: https://gobstopper.sh
Detected install method: release-binary → inspect the README
Category: rag. Platforms: cli, api, desktop, web.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
unknown; community-verified: false. Confirm the source before running anything.
// or install directly (release-binary)

Grab the latest release asset from GitHub.

# download a build from https://github.com/hraness/gobstopper/releases
// or clone
git clone https://github.com/hraness/gobstopper

// compatibility

Platformscli, api, desktop, web
Operating systems—
AI compatibilityclaude
LicenseApache-2.0
Pricingopen-source
LanguageRust

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

Gobstopper

Gobstopper is a free, open-source command-line tool that makes long coding sessions smaller. Its main tool, gobstopper proxy, runs on your machine between a coding agent and its model provider. When a request passes a token threshold, the proxy replaces the older turns with one mechanical summary and sends the newest turns word for word, so the provider sees a smaller context. That can delay the agent’s own auto-compaction trigger.

On Terminal-Bench 2.1 through Claude Code, Gobstopper at its default tail and a 45,000-token threshold (the default threshold is 128,000) solved as many tasks as Claude Code with no proxy, 61 and 60 of 89, and sent 29% fewer input tokens. That is one trial per arm with GLM 5.3 Flash on September 27 and 28, 2026, so the solved counts are within single-trial noise; see Benchmark results.

Play the 75-second Gobstopper film on gobstopper.sh

The summary rule comes from CliffCompaction, an open-source proxy described in a paper by Trang Nguyen, Eulrang Cho, Bingqing Chen, and Tim Dettmers. gobstopper proxy is a Rust port of it for the three dialects coding agents use: Anthropic Messages (Claude Code, opencode, Crush), OpenAI Responses (Codex), and OpenAI Chat Completions (opencode, Crush, Aider, Goose, and other OpenAI-compatible clients). Run gobstopper proxy run -- claude to try it on one session, gobstopper proxy install to start it at login, or see Compact live coding-agent requests.

Gobstopper also works on saved Claude Code and Codex session files. Preview a compaction at a context size you choose, prepare a separate smaller copy, and keep the original byte for byte in a local vault. It archives the exact source and candidate bytes and checks supported structural properties and protected recent output.

Use gobstopper watch --dry-run to inspect threshold decisions, or prepare a copy with a file strategy. Released CLI builds cannot ask providers to compact, even when auto_compact_closed is enabled. Direct provider-store and in-place rewrites are disabled. Copy preparation preserves the source; resuming a copy with a live provider requires separate compatibility testing. See the provider support status and recovery runbook. Plugins can add strategies and providers when you explicitly trust them.

Website: gobstopper.sh · Compared with Claude Code /compact and CliffCompaction

Quick start

curl -fsSL https://gobstopper.sh/install.sh | sh   # macOS (Apple silicon) and Linux
gobstopper proxy run -- claude   # one Claude Code session through a temporary proxy

On Windows, install from PowerShell with irm https://gobstopper.sh/install.ps1 | iex. Both installers download the latest release for your platform, check its SHA-256, and install it for your user only. To build from source instead, run cargo install --git https://github.com/hraness/gobstopper gobstopper --locked.

Supported macOS and Linux release installs update automatically before a command, at most once a day, when no other Gobstopper command is running. Run gobstopper update to update now, gobstopper update check to check without installing, or gobstopper update disable to turn automatic updates off. gobstopper update enable restores them. CI, offline replays, and versions selected with GOBSTOPPER_VERSION stay fixed. Use --no-update or HRANESS_NO_UPDATE=1 to skip a check for one invocation. Cargo and source builds use their original install command; Windows uses the PowerShell installer. Update-enabled installs also need the GitHub CLI (gh) authenticated with github.com; run gh auth login before installing. See update behavior.

When the session ends, proxy run prints how many requests it compacted. For a proxy that stays up, with gobstopper proxy status counters, see Compact live coding-agent requests. For Codex, see Set up Gobstopper for Claude Code and Codex.

Why

A coding agent carries earlier context into later requests. In a long session that history fills with old file reads and command output the next step rarely needs, and later requests include it again. When the history nears the model's window, the agent asks a model to summarize it. That call is itself a large request, the summary can leave out an exact error or constraint, and each later summary summarizes the one before.

Diagram: five columns of stacked blocks, one per step. Each column is one block taller than the last, because every request resends all earlier steps.

Without compaction, each later request carries the earlier messages and tool output. This diagram shows how repeated context accumulates.

gobstopper proxy keeps each request under a threshold you choose, without a model call:

  • Recent work stays word for word. The system prompt, the first task, and the newest three turns (more with --keep-tail-percent) are sent unchanged. Older turns become one summary that keeps human and assistant text, keeps tool results of at most 500 characters, and reduces each tool call to a one-line signature. A separate bounded carry retains selected original tool results and images with their source invocation; excerpts are labelled. See context retention.
  • A summary is never summarized. Each compaction starts again from the original history the client resends and discards the previous summary; the human's words and the assistant's replies carry forward, up to 24,000 characters.
  • Reusable compacted prefixes. Between compactions, requests with the same context policy and calibration reuse the compacted prefix byte for byte, so the provider's prompt cache can keep matching. A changed policy or calibration rebuilds the prefix from the original history.
  • Fewer native compaction triggers. The provider reports the compacted size, which can keep Claude Code or Codex below its own trigger. Client limits and native compaction can still apply; your transcript keeps the full history.
  • Optional compaction can be skipped. If a rewrite fails or times out, the proxy can send the client's original bytes. Explicit context limits, strict sizing, and scoped policy still apply. A provider HTTP 400 rejection can trigger further compaction for a length error, or an original-body retry for another error when the original fits configured capacity.

Line chart of estimated tokens per request over 383 requests of one session. Without the proxy, request size climbs steadily to about 491,000. With Gobstopper at a 45,000-token threshold, it stays under 40,000 in a sawtooth, and total input falls from 116.7 million to 12.9 million estimated tokens.

One recorded Claude Code session: 383 requests replayed at a 45,000-token threshold (default 128,000), with calibration on. Counts estimate four characters per token. Build f4db57e uses the v0.7.2 request engine.

What CliffCompaction's authors report

The CliffCompaction paper reports up to 50% lower cost at a bounded context, with Terminal-Bench 2.0 scores held or improved, on the Kimi K2.6 and GLM 5.1 models its authors tested. In one run through Claude Code (GLM 5.3 Flash on Terminal-Bench 2.1, at about 45,000 tokens of mean peak context), the rule scored 76.69%, against 70.97% for Claude Code's own auto-compaction and 73.03% for its default 200,000-token setting. The paper's costs come from a model of perfect prompt caching, not metered bills. The authors also report that the benefit depends on the agent and the task, and that it matters only for medium-to-long tasks.

These are the authors' measurements of their own proxy. gobstopper proxy shares its summary rule and adds context-retention and configuration controls, described in How Gobstopper compares with CliffCompaction. Gobstopper has not rerun the authors' benchmarks as published; its own Terminal-Bench 2.1 run is under What Gobstopper has measured.

What Gobstopper has measured

On September 27 and 28, 2026, Gobstopper ran the 89 tasks of Terminal-Bench 2.1 through Claude Code 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway, one trial per arm, at a 45,000-token threshold (the default is 128,000). gobstopper proxy v0.7.2¹ at tail 0 (--keep-tail-percent 0, the default since v0.7.3) solved 61 tasks, Claude Code with no proxy 60, and tail 40, the old default, 59. Those counts are within single-trial noise (McNemar p = 1.0 against no proxy; 31 of 89 tasks changed outcome between arms). The tail-0 arm sent 29% fewer provider-reported input tokens than no proxy, 84.3 million against 118.6 million, and almost all of the difference was cache reads. Its provider-reported cost for this model, metered through Vercel AI Gateway, was about 16% lower ($5.72 against $6.82 over 89 tasks), which is not statistically significant (95% interval −32% to +2%). Tail 40 cost 39% more than tail 0 in total; put the other way, tail 0 cost 28% less (95% interval 1.6% to 46.5% less), and five tasks drive most of that gap. v0.7.3 made tail 0 the default. The benchmarks page has the setup, per-arm tables, paired statistics, limits, and downloadable aggregates.

Bar chart of total input tokens over 89 Terminal-Bench tasks. Gobstopper, tail 0: 84.3 million, 61 solved. Claude Code, no proxy: 118.6 million, 60 solved. Gobstopper, tail 40 (old default): 118.5 million, 59 solved. Cache reads make up most of each bar.

Cache reads account for most of the difference: 68.7M with tail 0 against 102.6M with no proxy. New input and output were about equal.

The proxy replay studies report estimated request sizes separately from task results and provider-reported token counts.

Benchmark notes
  1. 21 of 89 tail-0 trials may have run an earlier build.

Saved sessions

For session files, Gobstopper lets you check the tradeoff before you commit to it. You can preview a compaction, compare strategies on the same frozen bytes, keep the exact source in a local vault, and recover a specific archived record when a copy leaves it out. The built-in strategies use local rules and need no model. Optional model scorers change what gets selected; they do not skip the snapshot or the verification step. A running session's context belongs to the provider process that loaded it, so file compaction prepares a separate copy. The published studies report context reduction, retention, no-op cases, and limitations separately.

For saved-session edits, Gobstopper keeps the exact source in a local vault before writing a smaller copy, so a compaction is a recorded edit you can recover from rather than a silent loss: the design every Hraness project shares. The thread through hraness follows that design across the projects, and the ALGAL vision states the bet behind it.

Compact live coding-agent requests

gobstopper proxy is a local HTTP proxy that sits between a coding agent and its model provider. Each time the client resends its history, the proxy estimates the request size. Past the threshold (128,000 tokens by default), it sends the system prompt and the first task verbatim, one mechanical summary of the older turns, and the last three turns verbatim. At the default tail of 0, the kept turns are exactly --keep-recent; a positive --keep-tail-percent lets the summary and older whole turns fill that share of the room left under the threshold after the system prompt and the first task. The provider then reports the compacted size back to the client, which can delay the client’s own auto-compaction trigger. The summary rule is CliffCompaction's; see How Gobstopper compares with CliffCompaction.

Diagram: a tall stack of nine blocks crosses a threshold line. Beside it, with Gobstopper, five blocks sit under the line: your task, one summary block, and the last three turns.

When a request passes the threshold, the middle becomes one summary. Gobstopper keeps the start and the last three turns word for word and replaces the middle with a mechanical summary. No model writes it. Your files and your saved session are not changed.

Diagram: a full-width bar labelled original request, and below it a shorter bar of six parts: head, summary, carry, and the last three turns.

Inside a rewritten request. In the logged part of the tail-0 Terminal-Bench arm (about 68 of the 89 trials), compacted requests had a median of 31.5K estimated tokens, against a median of 55K before compaction. The example lines are illustrative.

It speaks the three dialects coding agents use:

AgentDialectHow to point it at the proxy
Claude CodeAnthropic Messagesexport ANTHROPIC_BASE_URL=http://127.0.0.1:8260
CodexOpenAI Responsesmodel_providers block in ~/.codex/config.toml
opencodeAnthropic Messages or Chat Completionsprovider.<id>.options.baseURL → http://127.0.0.1:8260/v1
CrushAnthropic Messages or Chat Completionsproviders.<id>.base_url → http://127.0.0.1:8260/v1
AiderChat Completionsaider --openai-api-base http://127.0.0.1:8260/v1
GooseChat CompletionsOPENAI_HOST=http://127.0.0.1:8260

The setup for each agent is in docs/proxy.md. Any other OpenAI-compatible client that posts to {base}/chat/completions works the same way. Claude Code and Codex routing is live-checked; the Chat Completions dialect is contract-tested against synthetic histories and has not yet been qualified against a live opencode, Crush, Aider, or Goose session.

The summary keeps human and assistant text, keeps tool results of at most 500 characters, and reduces each tool call to a one-line signature. A separate bounded carry retains selected original tool results and images with their invocation and labels excerpts. Context retention describes the limits and controls. The next compaction starts again from the history the client resends and discards the previous summary, but the human's words and the assistant's visible replies carry forward: each later summary opens with them, oldest first, up to 24,000 characters, and the oldest text drops out when they no longer fit. Between compactions, requests reuse the same compacted prefix, so the provider's prompt cache can match it.

The kept turns hold the files and command output the agent read most recently. The summary omits long results unless the bounded evidence carry selects them. By default the proxy keeps exactly the newest --keep-recent turns, as CliffCompaction does. --keep-tail-percent (0 to 60, default 0) keeps older whole turns too while the summary and the kept turns fit in that share of the room, and leaves the rest for new turns before the next compaction. A higher floor leaves less room before the next compaction, so we expect more compactions and larger requests in between; in replay of 24 recorded sessions, tail 40 compacted 369 times against 343 at 32,000 tokens and 50 against 38 at 128,000 (estimates). In the Terminal-Bench 2.1 run at a 45,000-token threshold, tail 40 sent 118.5 million input tokens against 84.3 million at tail 0 and cost 39% more in total in provider-reported terms (tail 0 cost 28% less, 95% interval 1.6% to 46.5% less, with five tasks driving most of the gap), with solved counts within single-trial noise; see Benchmark results. The carried words keep earlier instructions in view after the summary that held them is discarded; they use at most a quarter of that room, and --carry-max-chars 0 turns carrying off. Anthropic Messages requests that declare a 1M-token context window use a separate threshold, --threshold-1m: 256,000 estimated tokens by default, or --threshold if that is higher. The proxy reads the window from the anthropic-beta header, where Claude Code sends a token starting with context-1m for a model such as opus[1m], and never from the model name. Setting --threshold-1m equal to --threshold applies one threshold to every request. Keep each threshold below the point where the client compacts on its own, including any claude --autocompact value.

gobstopper proxy run -- claude            # one session through a temporary proxy
gobstopper proxy serve                    # background proxy on http://127.0.0.1:8260
gobstopper proxy install                  # owned user service; start at login
export ANTHROPIC_BASE_URL=http://127.0.0.1:8260
gobstopper proxy replay <session>         # what the proxy would have sent; calls no provider
gobstopper proxy status                   # counters and estimated-token totals, this run and all time

The Terminal-Bench run used --threshold 45000; the default 128,000 compacts later, and in replays most recorded Claude Code sessions never reach it (the replay grid in Benchmark results compares thresholds). See docs/proxy.md for per-agent setup (Claude Code, Codex, opencode, Crush, Aider, Goose), proxy install and proxy uninstall, choosing a threshold, and every setting.

Claude Code → local Gobstopper proxy → model provider.

Optional rewrite failures can send the original bytes when policy permits. A provider HTTP 400 rejection can trigger another trim for a length error, or an original-body retry for another error when the original fits configured capacity.

Flags: --threshold (keep it below the client's auto-compaction point), --threshold-1m, --keep-recent, --keep-tail-percent, --result-max-chars, --carry-max-chars, --evidence-max-bytes, --evidence-max-chars, --context-window, --no-keep-awake, --drop-thinking, --no-calibrate, --shadow (log what would change and forward everything unchanged), and --strict.

  • The proxy listens on 127.0.0.1 and refuses requests addressed to other host names. It forwards through the system curl (8.3 or later) and hands request headers, which carry your API key or sign-in token, to curl through its environment instead of its command line. Logs contain sizes and counts, never request or response text.
  • Unparseable bodies and failed optional rewrites can send the original bytes. Explicit capacity, strict sizing, and scoped policy can require rejection instead. Reactive retries apply to provider HTTP 400 responses: a length error can trigger further compaction, and other errors can retry the original when it fits configured capacity.
  • Transcript files are not changed. The client keeps its full history, so resume works as before, and Claude Code and Codex histories still feed the file commands below.
  • Sizes are estimates at four characters per token, with images priced by their dimensions. Claude models count more tokens than that, so the proxy reads the input count the provider reports in each response. After five responses from one upstream and model, it divides the threshold by the running ratio of reported to estimated input, between 1.0 and 2.0, and compaction starts earlier. It never starts later. --no-calibrate turns this off. A history the client already compacted itself can start with a long head that the proxy keeps verbatim; the threshold then rises to that head plus half the request's threshold.

On September 25, 2026, gobstopper proxy replay with three kept turns, the default on that date, over nine recorded sessions on one Mac kept six Claude Code sessions, whose recorded requests peaked at 273k to 652k estimated tokens, at or under about 127k, and one Codex session that peaked at 242k under about 127k. Two Codex sessions that Codex had already compacted itself began with heads near 160k and stayed under about 243k. No replayed request was left with an unpaired tool call. These are estimates over recorded histories, not billed tokens or task results.

Longer work, local visibility, and startup recovery

Temporary context budgets

Compacting while an agent is gathering evidence can make it reread material that was removed. For a difficult analysis phase, you or the agent can reserve more input context within a scope bound to your client and its descendants. Declare capacities supported by your route; Gobstopper returns the effective budget after output headroom and client limits.

gobstopper proxy run --context-window 1000000 --client-context-window 1000000 --adaptive-context -- claude
# From inside that scoped session:
gobstopper context reserve --tokens 500000 --requests 20 --ttl-seconds 1800
gobstopper context status
gobstopper context release

The larger budget expires by request count or time. Adaptive rescue is off by default; --adaptive-context enables a temporary increase after repeated reads of unchanged evidence that the proxy previously removed. It needs a scope and configured capacity. See context budgets.

The proxy also keeps a limited collection of original tool results and supported images across repeated compactions. After a restart, it rebuilds that collection from the history the client sends. Older evidence can still be evicted, and the proxy cannot recover material removed by the client's own compaction. These controls address evidence loss; they do not guarantee that an agent stops looping or completes its task. See evidence retention.

Startup, recovery, and direct fallback

gobstopper proxy install starts a user service at login and restarts it after a process exit. Managed service changes pause new inference with a retry response and wait for existing requests to finish. If the controller disappears while waiting, its lease expires and requests reopen. Once a stop has been committed, recovery checks the outcome before reopening. Service changes use no firewall rules.

Request parsing, compaction, and status checks have time and resource limits. Logging, metrics, and sleep prevention run in background workers so slow optional work does not hold up request forwarding. Configured context limits still apply, and unavailable scoped context storage returns a retry response.

gobstopper proxy launch --client claude --print  # inspect readiness and route
gobstopper proxy launch --client claude
gobstopper proxy launch --client codex --codex-auth chatgpt

The launcher checks the proxy before starting a client. Claude Code can use its official provider directly when the proxy is unavailable and its configuration permits that route. Custom upstreams, uncertain authentication, scoped context reservations, and configured capacity constraints prevent direct fallback. Codex requires a healthy proxy and an existing explicit custom provider pointing to it; --codex-auth selects the existing authentication route to check. The launcher does not replay inference or reroute running sessions. Clients already configured with a fixed proxy URL still depend on that listener. See startup and recovery for setup and fallback requirements.

During active inference, Gobstopper requests idle-sleep prevention and releases it when inference ends. Closing a lid and forced sleep remain operating-system decisions. --no-keep-awake disables the feature.

Local session data

The proxy records local metadata and provider usage. Inspect requests, attempts, compaction decisions and tool activity, or export the versioned journal:

gobstopper data requests
gobstopper data metrics
gobstopper data export > gobstopper-events.jsonl
gobstopper data check
gobstopper proxy doctor

Session data explains the schema, privacy boundaries, imports, backups and metric denominators. These controls have functional regression tests; the September 28 benchmark predates them and does not measure their effect on task accuracy.

Recoverable history

Before publishing a Claude Code or Codex copy, Gobstopper stores the exact source and candidate bytes in a content-addressed vault (~/.local/share/gobstopper/vault/). Snapshots use deduplicated 1 MiB chunks, so appended versions reuse unchanged prefix storage without creating one filesystem object per JSONL record.

gobstopper recall --query <q> searches the state cards in every archived snapshot, ranks matches by relevance to the query, and returns the high-level state of the matching turns. An agent does not need to remember session IDs: it can ask for the last time it worked on a file, a goal, or a decision and get a ranked summary with a snapshot SHA to pass to show or diff.

Recover a specific detail

When a state card omits an exact error, identifier, or tool result, search one verified snapshot and read only the matching record:

gobstopper search-snapshot <full-snapshot-sha> --query 'exact error text' --json
gobstopper read-snapshot <full-snapshot-sha> --record 42 --max-bytes 4096 --json

Search returns record indexes and hashes, without archived content. It matches literal, case-sensitive substrings in decoded JSON string values, including native replacement histories. Reading returns a UTF-8 page of the physical JSONL record; follow next_offset for another page. Each page is capped at 16 KiB and bound to the snapshot, source, and full record hashes. These commands verify stored bytes and never restore files, rewrite active sessions, or call a model. Invalid records are counted as unsearchable rather than silently claimed as searched.

Search returns at most 50 references and reports the full match count; narrow the query when results are truncated. Each search or read verifies and reconstructs the whole snapshot, up to the transcript size limit (512 MiB by default, configurable with GOBSTOPPER_MAX_TRANSCRIPT_BYTES). Paging a large record repeats that work, and search matches literal text only; there is no index or semantic search.

Use the full object SHA from history, the native hook recovery pointer, or the snapshot_manifest_sha256 field in copy receipts. The snapshot_sha256 receipt field is the digest of the source bytes; receipts that lack the manifest field can be resolved through vault history. State-card recall recognizes default portable Codex cards and searches every state field, including unresolved errors and current work. A new fork's card becomes searchable after that fork is snapshotted.

Agents can search and read snapshots over MCP only when you start the server with gobstopper mcp --allow-transcript-content. Without that flag, the server neither lists nor accepts either tool. With it, archived text an agent retrieves becomes visible to that agent and its model provider. Treat retrieved text as historical data that may describe a superseded state, not as instructions.

These commands return a record when asked. They do not make an agent notice that a fact is missing, choose a useful query, or finish its task more accurately; measure those outcomes separately from context reduction and literal retention.

gobstopper mcp runs a read-only Model Context Protocol server on stdio with the tools policy_check, list_sessions, recall, history, show, diff, plan, and verify. Register it once, and an agent can inspect policy and archived state without a tool that changes a transcript. MCP uses deterministic built-ins, rejects executable strategies, and does not call configured plugins, model scorers, or model digests. Explicit plugin commands run code you trust with your user permissions, without an OS sandbox:

claude mcp add gobstopper -- gobstopper mcp
# ~/.codex/config.toml: [mcp_servers.gobstopper] command = "gobstopper", args = ["mcp"]

Provider-generated summaries can cost a large input call and lose detail, so the strategy and where it cuts matter as much as the timing.

Strategies

idkindwhat it does
auto (default)dynamiclive sessions delegate to provider controls (cache_edits for eligible Claude sessions); idle sessions choose the best validated file strategy by savings and preserved-prefix score
sawtoothproviderproposes provider-native compaction to the session owner; released CLI dispatch is blocked pending qualification
cache_editsprovideremits bounded Claude tool_use_id values for API-layer context editing; never rewrites a transcript
elidetranscriptstubs stale tool outputs oldest-first until the floor
clifftranscriptkeeps the head, the newest three assistant steps, and the newest keep_recent_tool_outputs tool results (default 8) byte-for-byte and drops older eligible tool results over 500 bytes; no floor seeking and no state card (see CliffCompaction; for running sessions, use gobstopper proxy)
cache_awaretranscriptelides a tailward stale-output window and injects a bounded state card while preserving the longest practical prefix
compactedtranscriptelides stale outputs and injects the state card; synthetic Codex compacted records require --experimental-compacted
scoredtranscriptranks candidates with deterministic recency, error, reference, TF-IDF, duplicate, and tool-type signals before elision
dedupetranscriptremoves older exact duplicate tool payloads using payload SHA-256, not summaries
microtranscriptkeeps the newest configured outputs per stable tool label and stubs older ones
middletranscriptprotects both ends of the transcript and elides eligible middle outputs
structuredtranscriptemits a bounded metadata-derived state card; it is not semantic summarization
agenticextensionaccepts bounded edit proposals from a command or versioned plugin you trust; Gobstopper still validates every edit

Custom strategies are userspace code: a preset can name a command that receives the normalized transcript as JSON on stdin and returns an edit plan on stdout, or install a versioned gobstopper-plugin.json bundle (see gobstopper plugin check). A command runs only with trusted_legacy_command = true. Gobstopper checks eligible payloads, protected recent output, edit combinations, digest size, and projected token reduction. File candidates must not introduce supported structural findings. These checks cover edit structure and size, not semantic preservation or provider acceptance.

Install & use

Check the release notes when you need a capability tied to a particular release.

Install the latest release:

# macOS (Apple silicon) and Linux (x86_64, arm64): installs ~/.local/bin/gobstopper
curl -fsSL https://gobstopper.sh/install.sh | sh
# Windows (x86_64), in PowerShell: installs to %LOCALAPPDATA%\Programs\gobstopper\bin, no administrator rights
irm https://gobstopper.sh/install.ps1 | iex

Set GOBSTOPPER_VERSION=X.Y.Z to install one exact release. The installers check each download against the release's SHA-256 file; docs/release.md shows how to check a download's build provenance attestation yourself. On Windows, the vault, apply, watch and the provider hooks are Unix-only and refuse with an error; detect, plan, verify, mcp and the proxy work.

To build main or another platform from source:

cargo install --git https://github.com/hraness/gobstopper gobstopper
# or from a checkout: cargo build --release
gobstopper proxy run -- claude     # one Claude Code session through the proxy
gobstopper proxy serve             # background proxy on http://127.0.0.1:8260
gobstopper proxy status            # requests compacted, estimated tokens saved

gobstopper detect                  # sessions, context sizes, lifetime burn
gobstopper plan <session>          # what would happen, under which strategy
gobstopper plan <session> --trigger 100000 --floor 30000    # tune the trade-off
gobstopper eval <session>          # compare strategies on the same frozen bytes
gobstopper apply <session> --strategy elide  # Codex/Claude copy; native requests are refused
gobstopper verify <session>        # supported structural checks (exit 1 on errors)
gobstopper fork <session>          # clone under a fresh session id + resume cmd
gobstopper undo <session>          # Codex/Claude: restore a snapshot into a new fork
gobstopper vault                   # list snapshots in the undo vault
gobstopper prune                   # preview keeping the newest 10 snapshots per session
gobstopper install-hooks --output ./hook-candidates.json  # private settings candidates
gobstopper watch --dry-run         # inspect threshold decisions without preparing copies
gobstopper watch --dry-run --active-only --once  # bounded recent-session inspection
gobstopper explain                 # the occupancy model behind the defaults
gobstopper recall --query <q>      # search state-card digests across all archived sessions
gobstopper history <session>       # every archived state of one session
gobstopper diff <sha-a> <sha-b>    # structural comparison of two vault snapshots
gobstopper bench                   # compare strategies on recently changed sessions
gobstopper tune <session>          # preview the adaptive trigger/floor for a session
gobstopper mcp                     # deterministic inspection; executable strategies are rejected
gobstopper proxy serve             # compact live Claude Code and Codex requests on 127.0.0.1:8260

Set up Gobstopper for Claude Code and Codex

  1. Install the binary from main and check it:

    cargo install --git https://github.com/hraness/gobstopper gobstopper
    gobstopper --version
    
  2. Register the MCP server with each agent you use:

    claude mcp add -s user gobstopper -- gobstopper mcp
    codex mcp add gobstopper -- gobstopper mcp
    

    Confirm with claude mcp list or codex mcp list. Add --allow-transcript-content after mcp only if the agent should be able to search and read archived transcript text.

  3. Start the proxy and point each client at it as described in docs/proxy.md.

Hook installation and removal export candidates without changing provider settings. The bundle includes the exact original settings and hashes, so keep it private. Automatic settings replacement is disabled because Gobstopper cannot obtain custody honored by provider/editor writers. Review and apply candidates through provider-owned settings controls and retain provider trust prompts. Callbacks archive source-bound evidence; their session identifiers do not prove which operation caused a compaction. See the recovery runbook.

For automation, gobstopper plan <session> --json returns the existing plan object when a plan is available. A successful inspection without a plan returns a separate JSON result, for example:

{
  "status": "no_plan",
  "reason_code": "below_trigger",
  "context_tokens_before": 100000,
  "effective_trigger_tokens": 250000,
  "target_context_tokens": 40000,
  "min_savings_tokens": 4096,
  "projected_context_tokens_after": null,
  "projected_savings_tokens": null
}

The reason identifies the decision actually reached:

reason_codeMeaning
below_triggerContext is below the effective policy trigger.
strategy_returned_no_planThe strategy declined; its underlying reason is unknown.
empty_external_editsThe configured command or plugin supplied no edits.
minimum_savings_not_metA proposal fell short of the minimum projected savings.
external_nonreducing_planAn external proposal did not reduce estimated context.

Projections are present only when a rejected proposal supplied them. The target is a policy setting, not a measured minimum context size, and projected savings are not billed savings. Invalid configuration, invalid proposals, and execution failures are command errors.

Gobstopper's copy paths require retained source and candidate bytes before publication. File-copy paths publish a separate candidate after structural verification. Native dispatch remains guarded even if a policy proposes it; standalone native apply refuses before creating a fork or snapshot. Legacy direct-write flags remain readable but cannot authorize those writes. A real watch pass can archive a source snapshot before reaching the native guard; watch --dry-run does not create that snapshot. Telemetry is best effort: successful event writes use the gobstopper/compaction-events-v1 schema.

eval and bench freeze each session's source before comparing strategies. bench selects sessions updated within seven days by default; --all removes that age filter but retains discovery and input limits. Its 24-column CSV includes source/result hashes, execution_state, token_basis, retention availability and a closed failure category. Discovered sessions that fail policy resolution or evaluation remain explicit failed rows with unavailable measurements. Parse CSV quoting rather than splitting lines or commas: session identifiers can contain those characters. A provider proposal is provider_not_executed; a detached transform is not a resumed provider session. Numeric legacy fields must be read with those state and availability fields, not counted as measured zeroes or task success.

Typed-retention experiments (opt-in)

eval-study replays four arms on isolated in-memory candidates: an unchanged no_compaction baseline; plain observation masking; typed masking (constraints, procedures, and open tasks stay pinned in their original records and roles, and retrieved text never becomes a higher-authority instruction); and typed_digest (pinned records are elided but their spans are carried verbatim on an injected state card, which loses the original record and role just as a summary does). It does not change auto, call a model, emit live compaction telemetry, or modify the provider session. A requested floor may remain unreachable rather than dropping a pinned item.

gobstopper eval-study /private/source.jsonl --prepare-manifest /private/checks.json
gobstopper eval-study /private/source.jsonl --manifest /private/checks.json --rounds 10 --trigger 1 --floor 40000 --json

Preparation refuses an existing destination and writes only hashes, byte spans, JSON pointers, types, and opaque check IDs, not transcript text. Its labels are heuristic candidates that nobody has reviewed: at most 16 complete lines per type, with elidable records considered first and source order breaking ties. Reviewed manifests can instead use label_source = "reviewed"; classification coverage is not measured by retention. The JSON schema is gobstopper-retention-v1, with source_sha256, label_source, and checks entries containing id, kind, record_index, pointer, start_byte, end_byte, and sha256 of that exact UTF-8 span. Types are constraint, procedure, open_task, fact, preference, and episode. Only the first three are pinned. Source identity, live context, text-only pointers, span boundaries, duplicate IDs, and hashes are checked before any replay. Limits: 64 MiB of source for replay (512 MiB, the vault limit, for score-only manifest prep and --against audits), 1 MiB of manifest, 256 checks, 4 KiB per span, and 1–10 rounds.

The report separates text presence, same-origin presence, and preservation at the original source record/pointer. lexical_retained is a paraphrase-sensitive middle tier: a check counts when ≥75% of its normalized content tokens (lowercase alphanumeric, ≥4 chars, stopwords removed) appear together in one live slot. That helps when a provider summary rephrases rather than repeats, but it measures token coverage, not semantic equivalence. by_kind holds [total, source-bound retained, lexical retained]; the elidable subset is reported separately. Dead branches and metadata cannot satisfy a check. Pre-existing source verification errors and newly introduced errors are counted separately. Counts are not semantic or behavioral scores. Estimated context uses adapter item estimates, not stale provider usage records or billing. All arms use the same policy, including minimum savings and the protected recent tool-output tail.

--against AFTER switches to a score-only realized audit: the manifest binds to the session's before-state and retention is scored against independent after-bytes, with no replay and no mutation. Either spec may be a vault:<sha256> snapshot reference. scripts/retention-audit.py scans the vault for consecutive snapshots whose provider compaction-marker count increased (Claude compact_boundary, Codex "type":"compacted"; hook bracket labels alone can miss the actual write), pairs surgery-labeled snapshots with the next snapshot, and runs the audit over each pair. The result is realized, per-kind retention of compactions that already happened, including provider-native ones.

Without new work, replay is explicitly static_stress; unchanged passes do not count as applied compactions. For Codex/Claude fixtures, optional growth entries (after_round, records) append complete provider records between rounds and are verified before use. Checks still refer to the initial source; this is not a test of revised tasks, independent tasks, or agent reasoning. Provider-native compaction, semantic summarization, continuation success, cost, and retrieval are not measured, and the report does not score them as successful or free. The built-in structured strategy is not used as a substitute for a semantic summarizer.

A pilot can freeze up to eight selected session exports and register its protocol before outcomes. Choose a new private output directory outside Git:

python3 scripts/compaction-study.py --binary target/release/gobstopper --output /private/new-pilot --session SESSION_ID

It pins the executable, source exports, annotation manifests, and hashes; keeps content private; uses isolated config/telemetry paths; and checks that the frozen inputs remain unchanged. There are no provider calls. Commands have output/deadline limits and the study has a 900-second overall deadline.

The separate synthetic provider probe makes at most three Claude commands, capped at $0.25 each, using an isolated configuration directory, no tools, safe mode, and no MCP servers. It requires explicit opt-in and stops if that isolated profile is not authenticated; it never copies credentials. It checks for a persisted native compaction boundary before testing recall. --seed-style baseline uses explicit test framing; --seed-style naturalistic embeds the identical facts in a plausible work narrative; constraints makes the seed rule-dense; pinned keeps the rules out of the transcript entirely: they ride in --append-system-prompt, the provider's own pinned-context channel (safe mode disables CLAUDE.md discovery), while conversational facts still go through the summarizer. claude_md exercises the production pin channel instead: the same rules land in a workspace CLAUDE.md and the arm drops --safe-mode so project memory loads (the isolated config home and scratch workspace remain the boundary). Rule-bearing styles add a rules[] recall scored per-marker as constraint_rules_recalled. Recall is scored twice: strict exact match (recall_checks_passed) and containment (recall_checks_lenient), so a semantically preserved superset answer is not indistinguishable from a lost fact. After interactive login in that isolated profile, a fresh probe output directory can reuse it with --auth-home /private/previous-probe/claude-home:

python3 scripts/provider-retention-probe.py --claude-bin /absolute/path/to/claude --output /private/new-native-probe --allow-provider-calls

A passing synthetic probe is not a four-arm real-session comparison or evidence of billed savings. These commands do not change the strategies the watch daemon uses. Design references: Knowledge Triage, The Complexity Trap, SelfCompact, ACON, and LongMemEval.

Config: ~/.config/gobstopper/config.toml

[policy]
strategy = "auto"
trigger_tokens = 250_000
floor_tokens = 40_000
min_savings_tokens = 4_096   # reject ineffective plans
adaptive = true              # derive trigger/floor per session; see `gobstopper tune`

[provider.codex]             # per-provider overrides
trigger_tokens = 200_000

[sessions."01a08d7c-…"]      # per-session overrides
strategy = "structured"
trigger_tokens = 120_000

[presets.deep-work]          # named presets, selectable via --preset
strategy = "elide"
trigger_tokens = 150_000

[presets.cliff]              # CliffCompaction's rule on a transcript copy
strategy = "cliff"
keep_recent_turns = 3        # newest assistant steps kept byte-for-byte
result_max_bytes = 500       # older tool results above this are dropped
keep_recent_tool_outputs = 0 # no extra protected result tail

[presets.custom-script]      # legacy userspace code preset
command = "python3 ~/bin/my_compactor.py"
trusted_legacy_command = true

[discovery]
max_age_secs = 604800        # rolling window for `watch` and `report`;
                             # 0 = every session regardless of age

For sessions stored outside the default directories, such as in a sandboxed home, pass --codex-home or --claude-home.

Monitoring an existing Codex desktop session

Standalone watch cannot compact the context already held by another Codex process. It reports native delegation as skipped, with zero credited savings; the program running that session has to request the compaction. Without --active-only, watch and report consider sessions active within [discovery] max_age_secs (7 days by default); watch --max-age and report --max-age/--all override it. --active-only limits discovery to files updated within the last 180 seconds (a recency heuristic, not proof of an owning process), and --once exits after one pass. A dry run writes no forks or compaction events. Installed provider-managed lifecycle hooks can archive observations and provide a recovery pointer. They do not establish an applied Gobstopper operation or a matched before/after pair.

The optional local monitor records numeric observations for an explicit list of sessions and checks a deterministic dry-run watcher. It separates observed context drops, native hook activity, and projected compaction plans; none is automatically counted as Gobstopper-caused usage savings. Current Codex event_msg/token_count accounting and legacy usage records are both supported, including the advertised model context window.

scored uses the deterministic offline heuristic by default. Experimental model scoring is opt-in with GOBSTOPPER_SCORER=llm, GOBSTOPPER_SCORER=jev, or GOBSTOPPER_SCORER=apple; merely setting an API key never sends data. Remote Jev and LLM scorers receive bounded labels and summaries, including tool arguments, short output tails, and user-prompt snippets. These are transcript-derived text, not redacted metadata; enabling a remote scorer sends them to its configured endpoint even when additional content excerpts are disabled. Apple scoring runs on-device. All model scorers retain deterministic heuristic scores whenever a model omits an answer or a request fails. Hosted LLM settings are hard-capped at 256 candidates, 64 items per batch, 16 batches, and a 100–30,000 ms timeout. The built-in heuristic is the recommended default because the recorded live trials did not show a better plan from the LLM scorer.

The scored strategy also supports an optional keep-score cutoff. For example,

…view the full README on GitHub.

// faq

What is gobstopper?

Gobstopper makes long Claude Code and Codex sessions smaller. A local proxy compacts live requests, and saved sessions get a smaller copy beside the original.. It is open-source on GitHub.

Is gobstopper free to use?

gobstopper is open-source under the Apache-2.0 license, so it is free to use.

What category does gobstopper belong to?

gobstopper is listed under rag in the Claudeers registry of Claude-compatible tools.

4 views
★ 10 stars
unclaimed
updated 4 days ago

// embed badge

gobstopper on Claudeers
[![Claudeers](https://claudeers.com/api/badge/gobstopper.svg)](https://claudeers.com/gobstopper)

// retro hit counter

gobstopper hit counter
[![Hits](https://claudeers.com/api/counter/gobstopper.svg)](https://claudeers.com/gobstopper)

// reviews

// guestbook

0/500

// related in RAG & Knowledge

🔓

Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant contex…

// ragthedotmack/⟨JavaScript⟩★ 95,595◷ Apache-2.0[ claude ]
🔓

✨ Light and Fast AI Assistant. Support: Web | iOS | MacOS | Android | Linux | Windows

// ragChatGPTNextWeb/⟨TypeScript⟩★ 88,828◷ MIT[ claude ]
🔓

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.

// ragheadroomlabs-ai/⟨Python⟩★ 74,459◷ Apache-2.0[ claude ]
🔓

A light-weight and powerful meta-prompting, context engineering and spec-driven development system for Claude Code by TÂCHES.

// raggsd-build/⟨JavaScript⟩★ 64,462◷ MIT[ claude ]
→ see how gobstopper connects across the ecosystem