
claude-octopus
๐ Token-frugal session memory & usage tracking for Claude Code โ curated nap/wake handoff instead of full-session resume, plus real 5h/7d rate-limit tracking.
Install with your AI
This project is repository unavailable โ not recommended for automated install, so we don't generate an auto-install prompt for it. Read the repo and decide for yourself.
// compatibility
| Platforms | cli, api, desktop |
|---|---|
| Operating systems | โ |
| AI compatibility | claude |
| License | MIT |
| Pricing | open-source |
| Language | Python |
oct-toolkit
Token-frugal session memory and usage tracking for Claude Code.
When to use it
Short version: it makes Claude Code remember things the way a person would, without getting more expensive the longer you use it.
- Starting work for the day โ
/oct-wake: picks up yesterday's note instead of re-reading the whole old conversation (which costs a lot more). - Mid-task, context getting fat, or stepping away for a bit โ
/oct-nap, then/clear: writes a ~1-2k word note, clears,/oct-wakepicks it back up next time. - "Didn't I run into this before?" โ
/oct-wake <keyword>: searches checkpoints and long-term memory together โ no need to remember which pool something lives in. - Wrapping up for the day โ
/oct-sleep: writes the note and consolidates memory in one shot. In a multi-window day, run this only in the last window you close; the others just/oct-nap. - Wondering what you've spent โ the status bar (turned on by
/oct-pulse) shows it live; for a detailed breakdown, or to check whether you have any wasteful habits, use/oct-checkup. - Not sure the conclusion you just reached is right โ
/oct-consult: a clean sub-agent decides from a neutral one-page brief instead of you switching model and re-reading your whole history to get a second look.
The rest of this doc is for anyone who wants the details, or is picking this up to maintain.
Why
claude --resume reattaches a whole old session โ every message, every file
read, still sitting in context on every future turn. This toolkit's core bet
is that curated, selective recall beats automatic full-context reload:
save a small handoff note, start clean, load only what's actually relevant.
This is a deliberate contrast to auto-capture-everything memory tools. Those are convenient, but real users report they can increase token usage (the compression/injection steps aren't free, and automatic relevance-guessing over-injects). Nothing here loads automatically except a one-line "you have a pending checkpoint" nudge โ everything else is loaded on request.
Does it actually save tokens?
Task type, session length, and model all change the ratio, so don't take anyone else's number as a guarantee for your own setup. Check it yourself, with one command:
/oct-checkup --current
Run it right before you'd normally /oct-nap (context feels fat) to get
Before. Then /oct-nap โ /clear โ /oct-wake, and run it again in
the new session to get After. The gap between the two is your real
saving, for your own workload.
For reference, here's what that comparison actually looked like on my
machine โ 8 real past sessions, resumed for real (claude --resume) vs.
loaded via /oct-wake with one fixed checkpoint, same model pinned for
both sides so the comparison isn't skewed by an unrelated cache miss:
| Old session size | --resume tokens | --resume cost | /oct-wake tokens | /oct-wake cost | Cost saved |
|---|---|---|---|---|---|
| 21K | 21,326 | $0.0880 | 25,569 | $0.0610 | 31% |
| 26K | 25,966 | $0.0616 | 25,848 | $0.0514 | 17% |
| 37K | 36,532 | $0.1030 | 28,310 | $0.0642 | 38% |
| 62K | 61,713 | $0.1971 | 25,620 | $0.0481 | 76% |
| 72K | 71,735 | $0.2738 | 27,887 | $0.0586 | 79% |
| 98K | 98,229 | $0.3364 | 27,410 | $0.0611 | 82% |
| 158K | 158,225 | $0.6025 | 26,910 | $0.0734 | 88% |
| 159K | 158,745 | $0.5476 | 26,959 | $0.0191 | 97% |
Not uniformly dramatic โ the two smallest sessions (under ~30K) save less
because /oct-wake's own fixed overhead (shared cache + the checkpoint
itself) is a bigger fraction of a tiny session. But 6 of these 8 sessions
land at 62K+, which is a normal size for a real working session โ and
that's exactly where savings are consistent, 76โ97%.
If the status bar is on (/oct-pulse), you don't even need to run the
command yourself โ the context size is already sitting there live, before
and after.
Reproduce the table yourself
scripts/oct-benchmark.sh is the actual script that generated the table
above โ it's in the repo, not a one-off. It makes real claude --resume
and claude -p calls against your own session history (same model pinned
on both sides so a version mismatch doesn't blow the shared cache and skew
the result), so running it spends a small amount of real API usage.
# one session vs. one checkpoint
scripts/oct-benchmark.sh <session-id> ~/.claude/summaries/<checkpoint>.md
# auto-pick your largest recent session + latest checkpoint
scripts/oct-benchmark.sh --auto
# sweep N real sessions spread across your own history's size range,
# one fixed checkpoint throughout โ prints a summary table like the one above
scripts/oct-benchmark.sh --sweep 8
Pin the model with OCT_BENCH_MODEL (default sonnet) if you want a
different one on both sides.
Is switching model mid-conversation actually expensive?
It's easy to feel frugal for a bad reason: discuss in a cheap model, only switch to a pricier one when you actually need its judgment. That feels efficient โ you're only "paying the expensive rate" for the moment you need it. What that misses: prompt caching is scoped per model, so the switch itself makes the new model re-read (and pay full cache-write price for) your entire conversation so far, before it answers anything. The longer you'd already been discussing, the bigger that one-time bill.
Real measurements โ resuming three of my own real sessions, switching
model mid-conversation (claude --model fable --resume <session>) and
then switching back (claude --model sonnet --resume <session>, since
that's the realistic case: get an opinion, keep working in the model you
were in), versus asking the same question fresh with only a one-page
neutral brief instead of the full history (what /oct-consult's default
mode automates):
| Conversation size before switching | Switch there | Switch back | Round trip | Cold brief instead | Saved |
|---|---|---|---|---|---|
| ~12K tokens (early in a session) | $0.28 | $0.08 | $0.36 | $0.18โ0.35 | 3โ50% |
| ~48K tokens (a solid working session) | $0.55 | $0.11 | $0.66 | $0.18โ0.35 | 47โ73% |
| ~91K tokens (a full day's session) | $1.13 | $0.23 | $1.36 | $0.18โ0.35 | 74โ87% |
The "cold brief instead" range covers whether the target model's own system prompt happens to already be warm in cache from a recent call (cheap end) or not (pricier end) โ either way it barely moves, because it never has to pay for your conversation history, only its own fixed overhead. That's the whole mechanism: switching model charges you for history, twice if you switch back, and a cold brief doesn't pay it even once.
The naive intuition โ "it's fine, I'll switch back right after, so I'm
only paying the expensive rate briefly" โ is exactly backwards: the return
trip re-reads the same history again, this time at the price of
whichever model you're returning to. Even the smallest real session tested
here, ~12K tokens (already past a skill load and a status check, which is
what a fresh Claude Code session usually looks like before you've said
anything), came out at a wash-to-slight-win for the cold brief once the
round trip is counted โ the "just switch, it's still early" case turned
out not to have a clear real-world example where switching actually wins.
Treat /oct-consult as the default the moment you're mid-task, and save a
plain /model switch for when you're not coming back โ e.g. you're
ending the session in the new model anyway.
Reproduce this yourself: pick one of your own real sessions with
/oct-checkup --sessions, then compare
claude --model fable --resume <session-id> -p "<question>" followed by
claude --model sonnet --resume <session-id> -p "<anything>" against a
fresh claude --model fable -p "<a short neutral brief>\n\n<question>" โ
/oct-checkup <session-id> --last=1 after each shows the real billed
cost.
What's in it
/oct-napโ write a ~1-2k token checkpoint before/clear/oct-wakeโ load the latest checkpoint (not the whole old session) in one shot;/oct-wake <keyword>also searches curated long-term memory/oct-dream//oct-sleepโ periodic memory consolidation/oct-pulseโ per-response token/cost breakdown (Stop hook), plus a persistent status-bar line with Anthropic's official 5h/7d rate-limit % and a reset-vs-ETA comparison so you can tell at a glance whether you'll hit the cap before the window resets on its own/oct-checkupโ per-call token/cost table for a session;/oct-checkup habitsscans the last N days of transcripts for spending patterns (long answer โ immediate short question, chatter on a fat context, repeated corrections, fat session ended without a checkpoint) and proposes one-line memory rules โ you pick which to save- Read-dedup hook โ blocks a byte-identical re-read of a file already read this session (same path, mtime, size, offset/limit), so duplicate content doesn't double up in context
/oct-consultโ get a second opinion without paying for it twice over: instead of switching model yourself (which re-reads your whole history and inherits its framing), the resident model writes a one-page neutral brief and hands it to a clean sub-agent to decide. Fixed cost, no anchoring, no clearing your conversation.--hereis the fallback for when you're about to spend 3+ turns with the bigger model anyway, or the decision leans on context too specific to this conversation to write down โ it re-evaluates in place instead, at the cost of some anchoring. The sub-agent runs onfable; if your plan doesn't include it (e.g. Pro/Team), it falls back toopusonce.
Power-user recipe: reopen every unfinished project at once
If you juggle several projects and each has its own /oct-nap checkpoint,
you can skip re-typing /oct-wake <hash> in each one by hand. Grab the
hashes from /oct-wake --list all, then a few lines of tmux will open one
window per project, each already cd'd in and resuming its checkpoint:
#!/usr/bin/env bash
# wake-all.sh โ one tmux window per project, each resuming its checkpoint.
# Edit this list after checkpoints change (new ones, or /oct-dream archiving old ones).
PROJECTS=(
"project-a|$HOME/code/project-a|a1b2c3d4"
"project-b|$HOME/code/project-b|e5f6a7b8"
)
tmux new-session -d -s wake -n scratch
for entry in "${PROJECTS[@]}"; do
IFS='|' read -r title dir hash <<< "$entry"
tmux new-window -t wake -n "$title" "cd '$dir'; claude '/oct-wake $hash'; exec bash"
done
tmux kill-window -t wake:scratch
tmux attach -t wake
Swap tmux for gnome-terminal --tab / osascript (macOS Terminal) /
your terminal of choice if you'd rather have real windows than tmux panes โ
the pattern is the same either way: one claude "/oct-wake <hash>" per
project, launched from its own directory.
Install
git clone <this-repo> ~/.claude/skills/oct-toolkit
(Not ~/.claude/plugins/ โ a bare clone there is never auto-discovered.
~/.claude/skills/<name>/.claude-plugin/plugin.json is what Claude Code
picks up on its own, no marketplace or install step needed.)
Then open Claude Code. You'll see one line:
๐ oct-toolkit installed. The status bar isn't on yet โ run /oct-pulse to enable it.
Type /oct-pulse, answer yes, done. That's the whole setup.
Why that one step is manual, and what it writes
Commands and hooks (SessionStart, PreToolUse[Read], Stop) load
automatically with the plugin. The bottom status bar is different: it's a
single statusLine setting in ~/.claude/settings.json, and a plugin
shouldn't silently take it over โ you might already be using it for
something else. So /oct-pulse asks first, then writes:
"statusLine": {
"type": "command",
"command": "python3 ~/.claude/skills/oct-toolkit/scripts/usage-analyze.py --statusline-hook"
}
Without it, the per-response cost line still works; the 5h/7d quota
numbers (which come from Anthropic's official rate_limits, delivered
only through the statusLine payload) won't appear.
If your Claude Code prefers marketplace installs, add this repo as a
marketplace source and install oct-toolkit from it โ same result.
Requirements
- Claude Code โฅ2.1.80 for official
rate_limits(older versions fall back automatically) - Python 3
- The bottom status bar (
/oct-pulse) is a terminal-only concept โ it likely won't appear in Claude Desktop, which shares the CLI's underlying engine (so commands and hooks work there) but has no documentedstatusLinesurface. The per-response cost line isn't affected either way.
License
MIT
// faq
What is claude-octopus?
๐ Token-frugal session memory & usage tracking for Claude Code โ curated nap/wake handoff instead of full-session resume, plus real 5h/7d rate-limit tracking.. It is open-source on GitHub.
Is claude-octopus free to use?
claude-octopus is open-source under the MIT license, so it is free to use.
What category does claude-octopus belong to?
claude-octopus is listed under other in the Claudeers registry of Claude-compatible tools.
// embed badge
[](https://claudeers.com/claude-octopus-2)
// retro hit counter
[](https://claudeers.com/claude-octopus-2)
// reviews
// guestbook
// related in Other
็ฌฆๅnature่ฎบๆๅญฆๆฏ่กจ่พพๅ็ง็ ็ปๅพ็Skill
Anti-AI-slop design skill for Claude Code, Cursor, and Codex.
Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.
Huashu Design ยท HTML-native design skill for Claude Code ยท Claude Code ้ HTML ๅ็็่ฎพ่ฎก skill ยท ้ซไฟ็ๅๅ / ๅนป็ฏ็ / ๅจ็ป + 20 ่ฎพ่ฎกๅฒๅญฆ + 5 ็ปด่ฏๅฎก + MP4 ๅฏผๅบ ยท Agent-agnostic