claudeers.
// Claude Skills

skillranker

Rust CLI powered by Jev from TypeSafe.ai that ranks agent skills for the next step using live session context. Includes Claude Code hooks, structured JSON, a…

Actively maintained
100/100
last commit 7 days ago
last release none
releases 0
open issues 2

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up skillranker (git-clone project) into my current project.
Found on https://claudeers.com/skillranker
Repo: https://github.com/Dicklesworthstone/skillranker
Homepage/docs: —
Detected install method: git-clone → git clone https://github.com/Dicklesworthstone/skillranker
Category: skills. Platforms: cli, api.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
active; community-verified: false. Confirm the source before running anything.
// or clone
git clone https://github.com/Dicklesworthstone/skillranker

// compatibility

Platformscli, api
Operating systems—
AI compatibilityclaude
LicenseNOASSERTION
Pricingopen-source
LanguageRust

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

SkillRanker

The right skill for the next step, powered by Jev from TypeSafe.ai.

A standalone Rust CLI that puts TypeSafe.ai's Jev at the center of skill selection: Jev evaluates your agent's live context, compares the available skills, and estimates which ones fit the next step. SkillRanker supplies the session integration, local safeguards, and inspectable feedback around it.

A TypeSafe API key is required to use SkillRanker's ranking system. Sign up at the TypeSafe console to get your own key.

sr demo --case useful     # Inspect an offline fixture before connecting a session
sr rank --allow-network   # Rank skills for the selected session
sr hook claude           # Run the Claude Code prompt-hook integration
sr tui                   # Inspect rankings in an inline terminal display

Contents


TL;DR

The problem. A large skill library gives an agent plenty of procedures to choose from, but choosing is itself a task. Similar descriptions obscure useful distinctions. A skill that helped at the start of a conversation can be irrelevant three turns later. Loading a plausible but unsuitable skill consumes context and can redirect otherwise sensible work.

The solution. SkillRanker (sr) combines the recent conversation, current request, workspace signals, and the selected harness's visible skill inventory. Jev from TypeSafe.ai is the key enabler of the system. It first compares the candidates broadly, then reads richer excerpts from a shortlist and evaluates whether each one fits. Both comparisons include a real “none of these” option. The result is advisory: the agent follows the user's instructions and decides what to consult.

For libraries with more than 254 eligible skills, Quill from FrankenSearch narrows the candidates locally before Jev evaluates them. Smaller rosters reach Jev in full. Explicit skill requests resolve locally before either stage.

SkillRanker does not include a local model or a substitute inference provider. The ranking workflow requires your own TypeSafe account and API key. Local retrieval prepares the candidates; Jev supplies the evaluations that make the recommendations possible.

Why sr?

NeedWhat SkillRanker provides
Evaluate meaning and task fitJev's typed Choice and Noul evaluations from TypeSafe.ai power both ranking passes
Choose for the current stepExact session identity, the newest prompt, recent tool evidence, and project signals
Suggest something the agent can loadHarness-aware visibility, override resolution, stable skill identities, and content revalidation
Respect an explicit requestLocally resolve a requested skill before probabilistic retrieval or ranking
Search a large libraryQuill lexical prefiltering from FrankenSearch, admitting up to 254 skills plus a none option to each Choice
Separate similar skillsDetailed reranking with bounded descriptions and body excerpts
Recognize when no skill fitsRelevance gates, per-candidate fit checks, and sentinel-based abstention
Understand a missing suggestion--why-not traces where a candidate was excluded, with thresholds and concrete recovery hints
Reproduce a surprising resultOpt-in case capture and offline replay compare compatible local policies without another Jev call
Get started without sharing a sessionOffline fixture demos and a readiness report identify the next setup step
Keep the agent movingA failed hook recommendation produces a quiet, non-blocking fallback
Control interruptionsSilent ordinary abstentions and scoped, expiring snoozes preserve explicit skill requests
Bound repeated expenseOptional shared HTTP-attempt allowances and a provider circuit breaker cover concurrent local sessions
Review what happensUsefulness, interruptions, attempts, and cost share a report with explicit label coverage
Evaluate within a budgetOffline replay, explicit live-request caps, and reproducible samples with recorded selection probabilities
Assess recommendation harmControlled comparisons, uncertainty bounds, and optional monitoring across repeated evaluations
Control disclosureNetwork opt-in, field-level disclosure receipts, a minimal context profile, and separate persistence controls

The approach builds on the TypeSafe skill-suggestion recipe. SkillRanker adds session identity, harness visibility, bounded execution, and a local evaluation loop. The comprehensive plan explains the full design and acceptance criteria.

Quick Example

# See a labeled fixture result without a key, network, or private session.
sr demo --case useful

# Inspect local configuration and the available adapters.
sr doctor --json
sr doctor --config
sr capabilities --json

# Inspect visible, shadowed, and excluded skill records.
sr roster --json

# Preview the redacted wide-pass request without network or persistence effects.
sr rank --context scratch/context.json --dry-run

# Evaluate an explicitly selected conversation.
sr rank --context scratch/context.json --allow-network --json

# Find where an expected candidate was excluded; this adds no inference calls.
sr rank --context scratch/context.json --allow-network --why-not SKILL_ID --explain

# Preview the Claude hook settings change, then apply it.
sr install-hook claude
sr install-hook claude --apply

# Review observations without treating adoption as proof of usefulness.
sr stats --since 7d --by-skill

# Replay a labeled evaluation artifact without making network requests.
sr eval --dataset scratch/evaluation.json --explain

# Inspect description quality and suspected coverage gaps locally.
sr doctor --descriptions
sr gaps

Design Philosophy

  1. Choose for the next action. The current request matters, as do the recent failure, the task context, and the instructions already loaded.
  2. Resolve authority locally. The user decides what is required or excluded. The harness determines what can be loaded. A model answer cannot change either.
  3. Separate preference from applicability. Winning a comparison is not enough. A recommendation must survive fit, visibility, loaded-state, and none-option checks.
  4. Keep evidence inspectable. Preserve provider estimates and local arithmetic. --explain exposes computations without inventing model-generated reasons.
  5. Separate adoption from usefulness. Observing a load is useful telemetry. Learning a better policy requires independently judged examples and a holdout.
  6. Bound each evaluation. Input, discovery, subprocesses, networking, retries, persistence, and cleanup consume one ranking deadline. Batch evaluation also has explicit request and total-runtime limits.
  7. Stay standalone. Discovery, parsing, redaction, retrieval, and feedback live in sr. No skill-manager service or private database is required.

How It Compares

These are workflow choices, not benchmark rankings.

ApproachInput to selectionStrengthTradeoff
Manual selectionYour knowledge of the task and libraryDirect control without a ranking serviceRequires remembering each skill's coverage
Keyword searchA query over names and descriptionsCheap local discoverySynonyms and adjacent procedures can be hard to distinguish
Load every skillThe full library's instructionsMakes every procedure available immediatelyConsumes context regardless of relevance
SkillRankerExact session, visible roster, and explicit constraintsEvaluates candidates and can abstainFresh Jev evaluations require authorized network access

SkillRanker recommends procedures. It does not execute skills, grant permissions, or override the agent's governing instructions.

Installation

From source

Build the sr binary with the repository's pinned Rust toolchain and lockfile:

git clone https://github.com/Dicklesworthstone/skillranker.git
cd skillranker
cargo install --locked --path . --bin sr

For a checkout-local binary:

cargo build --locked --release --bin sr
./target/release/sr --help

Runtime setup

Sign up for TypeSafe.ai, then create your own API key in the console. A TypeSafe API key is required to use SkillRanker's ranking system. Jev is the evaluation engine for the entire ranking workflow. Every user supplies their own credential; SkillRanker does not distribute a shared key.

Set TYPESAFE_API_KEY through your shell or secret manager. The environment example lists the service settings. For a local checkout, create .env from the example only if it does not already exist and restrict access with chmod 600 .env before entering your own API key. The .env file is ignored by Git. Export its values into the process environment before running sr or starting an agent whose hooks need the key:

# Run from your checkout, after filling in your own trusted .env file.
set +x
set -a
. ./.env
set +a

Treat .env as a local shell configuration file and source only content you trust. Keep the key out of shell history, logs, and tracked files. Credentials alone do not enable remote transmission.

ComponentRole
TypeSafe API keyAuthenticates fresh Jev evaluations
Network opt-in--allow-network for a run, or network.enabled in trusted user configuration
Visible skill inventoryHarness-resolved skills or an explicit roster file
Session inputClaude hook, normalized context, supported native transcript, or optional cass export
cassOptional archive discovery and access across coding-agent formats

sr does not require ms, a local inference server, or an embedding model. The primary local platform scope is Linux and macOS. Consult sr capabilities --json for the adapters, events, and optional features in a build.

Quick Start

  1. Try an offline fixture. Run sr demo --case useful, then try none, explicit, or unavailable. These labeled examples exercise the local pipeline without reading a private session or contacting Jev. Their output is non-actionable and does not establish live provider health.
  2. Sign up and configure your own TypeSafe API key. Create an account and key in the TypeSafe console, then export TYPESAFE_API_KEY using the runtime setup instructions. SkillRanker relies on Jev for its ranking evaluations.
  3. Check the environment and roster. Run sr doctor --json, sr capabilities --json, and sr roster --json in the agent's workspace. Confirm that candidates are loadable, not merely present somewhere on disk.
  4. Choose the session. Supply --context FILE, --transcript FILE --harness claude_code, or --session PATH for cass. Automatic discovery must resolve one unambiguous session.
  5. Preview and rank. Use --dry-run to inspect the redacted wide payload and disclosure receipt, then --allow-network --json for a fresh evaluation.
  6. Prepare a recorded shadow trial. Initialize optional history with sr ledger init, then explicitly set network.enabled = true in trusted user configuration if you want live hook evaluations. The earlier --allow-network flag authorized only that CLI run. Preview and apply sr install-hook claude; shadow mode evaluates without injecting advice. Without a ready ledger, the hook can rank but cannot promise recorded trial evidence; without network consent, it cannot obtain fresh Jev answers.
  7. Enable advisory output deliberately. Set hook.mode = "advisory" in trusted user configuration after reviewing the integration and its behavior.

Readiness checks

sr doctor reports each prerequisite separately, names the next concrete step for a failed check, and lists local commands that remain useful:

CheckWhat it establishes
Input and rosterA usable source and loadable candidates are available
CredentialA key is present; this alone does not authenticate it
Network authorizationTrusted settings permit a live request
TransportUntested or previously verified by a separately authorized, budgeted live check
LedgerOptional history is ready or degraded; its absence need not block ranking
Hook and snoozesShadow/advisory mode and active scoped interruption controls

Doctor runs locally by default and does not implicitly send a test request, install hooks, migrate storage, or change configuration. Demo uses bundled synthetic contexts and labeled synthetic or recorded responses without touching user configuration or state. Neither command supplies a replacement for Jev.

A previous transport check includes its time, scope, and runtime/endpoint/configuration identity. Incompatible changes invalidate it; a stored check does not establish current provider availability.

sr capabilities --json separates implemented adapters from tested harness versions/features and unverified versions. Adapter conformance covers prompt timing, branch identity, visibility, restrictions, compaction, load evidence, and hook output, deadlines, and delivery. Native advice requires passing real-harness evidence for every dimension on the installed version; protocol fixtures and official schema documentation alone cannot authorize it. Cass remains an explicit archive source. Native Codex, omp/pi, and Grok integrations each need their own support record. Incompatible identity or visibility semantics disable advice. The adapter contract defines these checks.

Command Reference

Bare sr is equivalent to sr rank: a table on a TTY, JSON otherwise. It does not start a TUI. Source flags are mutually exclusive, and piped stdin is consumed only by an explicit input mode. sr capabilities --json lists the commands, schemas, and optional features available in the installed build.

Ranking and inspection

CommandPurposeExample
sr demo --case CASEInspect a labeled offline fixturesr demo --case unavailable
sr rankRank the next stepsr rank --allow-network --json
sr rank --context FILERead normalized context; - means stdinsr rank --context scratch/context.json --dry-run
sr rank --transcript FILE --harness NAMERead a supported native transcriptsr rank --transcript scratch/session.jsonl --harness claude_code --offline
sr rank --session PATHExport an exact session through casssr rank --session scratch/session.jsonl --allow-network
sr rank --why-not ID --explainTrace an expected candidate's exclusionsr rank --allow-network --why-not SKILL_ID --explain
sr rank --save-case FILEExplicitly capture a bounded redacted replay casesr rank --allow-network --save-case scratch/case.json
sr replay FILERecompute a historical case offlinesr replay scratch/case.json --compare-policy scratch/candidate.toml
sr hook claudeHandle the Claude prompt-hook protocolsr hook claude --shadow
sr roster --jsonInspect visibility, overrides, records, and exclusionssr roster --json
sr roster --snapshot FILEExplicitly export a bounded private roster manifestsr roster --snapshot scratch/roster-snapshot.json
sr roster --diff FILECompare fresh discovery with a saved roster snapshotsr roster --diff scratch/roster-snapshot.json
sr doctor --jsonInspect local configuration and readinesssr doctor --json
sr doctor --configExplain effective non-secret values and their sourcessr doctor --config
sr capabilities --jsonDescribe commands, schemas, features, limits, and exitssr capabilities --json
sr tuiOpen the inline viewersr tui

Hooks, feedback, and analysis

CommandPurposeExample
sr install-hook claudePreview a managed hook settings changesr install-hook claude --apply
sr uninstall-hook claudePreview removal of the managed entrysr uninstall-hook claude --apply
sr statsReport observation and operational metricssr stats --since 7d --by-skill
sr observeReconcile structured load eventssr observe --transcript scratch/session.jsonl --harness claude_code
sr feedbackRecord an explicit usefulness judgmentsr feedback EVENT_ID --skill SKILL_ID --verdict useful
sr feedback --instead IDRecord a better alternative for an eventsr feedback EVENT_ID --skill SKILL_ID --instead ALTERNATIVE_ID
sr snoozePreview a scoped temporary advisory mutesr snooze EVENT_ID --skill SKILL_ID --for 30m
sr budgetInspect or preview a shared HTTP-attempt allowancesr budget --max-attempts 100 --window 1h
sr evalReplay a labeled evaluation artifact offline by defaultsr eval --dataset scratch/evaluation.json --explain
sr calibrateReport a candidate threshold configurationsr calibrate --evaluation scratch/report.json
sr calibrate --rollback REVISIONPreview restoration of managed policy fieldssr calibrate --rollback POLICY_REVISION
sr doctor --descriptionsCheck description quality locallysr doctor --descriptions
sr gapsReport suspected coverage gapssr gaps
sr ledger initInitialize local history explicitlysr ledger init
sr ledger migratePreview a supported schema upgradesr ledger migrate --apply
sr ledger prunePreview retention cleanupsr ledger prune --before 2026-09-01
sr ledger clearPreview clearing local historysr ledger clear

--apply performs a previewed hook, snooze, budget, calibration, or ledger mutation. Calibration consumes a labeled evaluation artifact; it does not silently change project settings after a number of observed loads. Description audits use the network only with an explicit online request and network authorization.

Evaluation controls

sr eval defaults to replay with zero network requests. A live run needs --online, trusted network authorization, and an explicit --max-requests cap. That cap counts HTTP attempts across the entire batch, including retries. Each case also has its own ranking deadline; the batch stops scheduling work when a request or runtime limit is reached and reports unfinished cases.

FlagDefaultMeaning
--dataset FILERequiredVersioned, consented evaluation data and compatible recorded responses for replay
--onlineOffPermit fresh Jev evaluations when network access is separately authorized
--max-requests NRequired for live runsMaximum HTTP attempts across the batch, including retries
--max-runtime-ms N600000Overall batch deadline, in addition to per-case deadlines
--sample-size NFull supplied frameSelect a bounded sample of task-family representatives
--seed SFresh recorded random seed when samplingDeterministic diagnostic selection or reproduction of a recorded sample; a fixed seed alone is not probability-sampling evidence
--explainOffInclude equations, substituted values, assumptions, and interpretation in the report
# Reproduce a diagnostic selection; seed 42 alone supports no sampling guarantee.
sr eval --dataset scratch/evaluation.json --sample-size 100 --seed 42 --explain

# Draw and record a random sample, then authorize a bounded live evaluation.
sr eval --dataset scratch/evaluation.json --sample-size 100 \
  --online --allow-network --max-requests 400 --max-runtime-ms 600000

Sampling does not grant network access or enlarge the request budget. Missing stage responses remain unevaluated in replay; they are never replaced with invented scores. See evaluation and sampling for the report's denominators and uncertainty rules.

Ranking controls

FlagDefaultMeaning
--messages N12Recent logical messages
--budget-chars N12000Rendered context budget, including the latest request
--top K5Maximum eligible suggestions returned
--shortlist M8Real candidates admitted to the rerank
--gate F0.30Overall need threshold
--fits F0.30Minimum candidate fit
--timeout-ms N3000Whole one-shot ranking deadline; applied separately to each TUI/watch refresh
--roster FILEHarness discoveryReplace discovery with an explicit inventory
--require-skill IDNoneResolve an explicit required skill; repeatable
--latestOffExplicitly choose the newest discovered session
--context-profile PROFILEstandardChoose standard context or the bounded minimal disclosure profile
--no-toolsOffRemove tool arguments and results from outgoing context
--no-cacheOffDisable response-cache reads and writes
--no-ledgerOffDisable all ledger and ingestion-cursor access; use transient evidence
--no-persistOffDisable all persistent state, including cache, cursors, and locks
--offlineOffGuarantee zero network calls
--allow-networkOffAuthorize network evaluation for this invocation
--explainOffInclude distributions, exclusions, truncation, and score contributions
--why-not IDNoneTrace a candidate from the current snapshot with --explain, without adding requests
--save-case FILEOffCLI-only, explicit capture for offline replay; cannot overwrite an existing file
--dry-runOffPreview the redacted request without network or persistence effects

Sizes satisfy 1 ≤ K ≤ M ≤ 32; fewer available candidates is normal. Parsing is strict, with documented aliases only. Invalid or conflicting privacy flags produce an error rather than being silently corrected.

An explicit chunk-overflow experiment uses bounded groups and reduction rounds. It has separate request limits and availability in the capabilities contract; the normal overflow policy uses local prefiltering.

JSON output

The versioned output contract specifies decision, error, quality, trace, and non-actionable report schemas.

This illustrative result has two eligible candidates. The score arithmetic uses w_fit = 1, with priors and phase weighting disabled; timing and usage are examples.

{
  "schema_version": 1,
  "event_id": "example-event-001",
  "decision": "ranked",
  "reason": "eligible-candidates",
  "harness": "claude_code",
  "context_quality": "complete",
  "quality": {
    "prompt_complete": true,
    "task_anchor_known": true,
    "history_windowed": true,
    "attachments_omitted": false,
    "source_gaps": false
  },
  "roster": {
    "total": 2,
    "eligible": 2,
    "wide_candidates": 2,
    "shortlist": 2,
    "partial": false,
    "retrieval": "full",
    "provenance": {
      "snapshot_id": "000000000000000000000000000000000000000000000000000000000000000a",
      "policy_version": "ranking-v1",
      "wide_set_id": "000000000000000000000000000000000000000000000000000000000000000b",
      "rerank_set_id": "000000000000000000000000000000000000000000000000000000000000000c"
    }
  },
  "needs_skill": 0.74,
  "choice_confidence": 0.81,
  "none_probability": 0.10,
  "phase": "debugging",
  "skills": [
    {
      "rank": 1,
      "skill_id": "s_01",
      "name": "rust-test-triage",
      "invocation_name": "rust-test-triage",
      "rank_score": 0.888889,
      "rerank_probability": 0.60,
      "wide_probability": 0.55,
      "fits": 0.80,
      "path": ".claude/skills/rust-test-triage/SKILL.md",
      "content_hash": "0000000000000000000000000000000000000000000000000000000000000001"
    },
    {
      "rank": 2,
      "skill_id": "s_02",
      "name": "rust-code-review",
      "invocation_name": "rust-code-review",
      "rank_score": 0.111111,
      "rerank_probability": 0.30,
      "wide_probability": 0.35,
      "fits": 0.50,
      "path": ".claude/skills/rust-code-review/SKILL.md",
      "content_hash": "0000000000000000000000000000000000000000000000000000000000000002"
    }
  ],
  "omitted_rank_mass": 0.0,
  "cache": {
    "hit": false,
    "wide_hit": false,
    "rerank_hit": false,
    "age_ms": null,
    "stale": false
  },
  "model": {
    "requested": "jev-latest",
    "wide_returned": "jev-latest",
    "rerank_returned": "jev-latest",
    "immutable_revision": null
  },
  "usage": {
    "requests": 2,
    "http_attempts": 2,
    "input_tokens": 6400,
    "output_tokens": 480,
    "unknown_usage_attempts": 0
  },
  "persistence": "recorded",
  "warnings": [],
  "warnings_omitted": 0,
  "elapsed_ms": 720
}
DecisionMeaning
rankedUp to K eligible suggestions from a successful evaluation
explicitLocally resolved user requests, with no invented model certainty
abstainValid input and policy produced no advisory recommendation
unavailableAn operational, input, privacy, or coverage problem prevented a decision

Demo and replay use separately versioned, non-actionable envelopes around their synthetic or historical decisions. A demo's context is synthetic; any recorded provider response retains its own provenance. It remains demonstration evidence, not a live evaluation or a quality-gate result. Neither artifact uses the live hook output channel.

choice_confidence describes the rerank distribution. fits is a model estimate of suitability. rank_score is a relative local score over eligible candidates. They are different quantities. Top-K truncation preserves the original eligible normalization and reports omitted mass. Fields from an unexecuted stage are null, not fabricated zeros.

Cache hits and returned model identities are recorded separately for each stage. When reranking is required, a wide-stage hit alone is not a complete offline result. An alias such as jev-latest does not identify an immutable model revision. persistence: "recorded" means the ranking metadata was committed before output; it does not mean the harness acknowledged or consumed the recommendation.

Quality metadata includes prompt_complete, task_anchor_known, history_windowed, attachments_omitted, and source_gaps. These describe the admitted input; context_quality: "complete" does not claim that the entire conversation history was read. Warnings are bounded to 32 details plus an omitted count. Live decision JSON and individual demo/replay/report summary envelopes are capped at 2 MiB. The separate 256 MiB evaluation-artifact limit bounds a streamed dataset or report with at most 10,000 case records; it does not enlarge an individual output envelope. Roster listings and full-wide explanations paginate against a fixed snapshot; a changed snapshot requires restarting. Trace pages contain at most 128 entries and bind both snapshot and query identity. Unevaluated entries keep operands and reasons null instead of inventing evidence.

Exit codes

CodeMeaning
0Ranked, explicit, valid abstention, or another successfully completed command
2Invalid usage or configuration
3Missing or ambiguous session
4Provider, authentication, or network failure; request-admission refusal
5Empty/unusable roster, unresolved explicit request, or Quill retrieval failure
6Overall deadline exhausted
7Malformed, oversized, or unsupported input
8Network transmission disallowed
9Required storage or administrative mutation failed
10Invalid structured provider response
11No complete valid result under offline/cache-only constraints

JSON errors include schema_version, decision: "unavailable", and an error object with code, kebab-case kind, message, hint, and retryable. Ordinary ranking can succeed with a storage warning; an explicit feedback write cannot claim success when its required write failed. Incomplete essential context and output-limit failures use exit 7, unresolved explicit requests and Quill retrieval failures use 5, and superseded input uses 3. Overall deadline exhaustion uses 6. retryable means a fresh invocation with the same intended inputs may succeed; it does not grant network access or relax a deadline.

Evaluation and replay reports distinguish execution from quality: run_status is complete or partial; gate_status is passed, failed, not-established, or not-applicable. Exit zero means the report was produced, not that every comparison was replayable or a policy passed. Promotion requires the intended complete cohort, compatible evidence, and explicit passed gates. Fatal errors retain their nonzero exit code and available partial-work/usage data. Each summary accounts for requested/completed cases and required/completed stages. Partial, incompatible, empty, or synthetic evidence cannot carry a passed gate. A completed report can faithfully describe a failed historical request without treating that historical error as a failure to generate the report.

The dedicated hook maps recommendation failures to quiet exit-zero behavior so it never blocks the agent. CLI failures retain their meaningful exit codes.

Configuration

Ordinary settings resolve from lowest to highest priority:

built-in defaults
  -> trusted user configuration
  -> allowlisted workspace configuration
  -> recognized SR_* environment variables
  -> command-line flags

On Linux, user configuration falls back to ~/.config/sr/config.toml; project configuration is .sr/config.toml at the workspace root. Other platforms use native configuration directories.

Trusted user settings for an advisory hook include:

[network]
enabled = true

[hook]
mode = "advisory"

Without those choices, remote transmission is disabled and the hook runs in shadow mode. --allow-network can authorize a single CLI evaluation.

VariablePurpose
TYPESAFE_API_KEYTypeSafe bearer credential; never serialized or stored in project config
TYPESAFE_ENDPOINTTrusted HTTPS base origin; sr appends /v1/systemone once
SR_MODELRequested model; default jev-latest
SR_MESSAGES, SR_BUDGET_CHARSContext-volume limits: 1–12 messages and 1–12,000 Unicode scalar values
SR_TOP, SR_SHORTLISTOutput and rerank sizes, satisfying 1 ≤ top ≤ shortlist ≤ 32
SR_GATE, SR_FITSFinite thresholds in [0, 1]
SR_TIMEOUT_MSWhole-invocation deadline, 201–60,000 ms

Workspace configuration may tune bounded ranking values and exclusions. It cannot authorize networking, change endpoints/proxies, supply credentials, expand transcript access, disable redaction, or enable raw retention. Unknown or duplicate keys, forbidden project settings, and invalid values are errors after bounded configuration reads and before discovery, networking, or mutation. Unrecognized SR_* environment names are errors too. Project context settings may reduce message/character budgets, select minimal, or remove tool content; they cannot widen a trusted user's disclosure settings. Exclusions and skill roots merge as unions. Project roots must remain relative to the workspace; only trusted user configuration can authorize absolute roots. Readers also check symlink containment when opening files.

Credentials are environment-only and kept outside serializable configuration. The v1 endpoint override is environment-only; model overrides use trusted user configuration or SR_MODEL. Proxy, redaction-disable, and raw-retention settings are reserved and rejected. The configuration contract lists every key, bound, source layer, and restriction.

The endpoint accepts an empty or root path and rejects URL credentials, query strings, fragments, and non-root paths. Credentials, cache identity, and request allowance scope use the canonical origin. A development-only loopback HTTP exception cannot carry production credentials.

sr doctor --config shows non-secret effective values and their winning sources. Disallowed overrides are configuration errors; no valid policy or fingerprint is reported for them. Credential presence is reported without its value, and endpoint overrides expose only their presence. --offline and --allow-network are accepted here but conflict; this inspection command never sends a request.

The library's cli::ConfigFiles shares bounded initial configuration reads with file refreshes. Refresh keeps the invocation's validated environment and CLI layers, returning the current configuration and a boundary-specific comparison against its typed policy receipt. Malformed, unreadable, or late reads fail rather than authorize output. Rank orchestration and its HTTP/publication call sites are not implemented yet; these library checks do not establish live ranking support. Shared request allowances and snoozes are explicit trusted-user controls. Project configuration cannot enable, raise, or disable the allowance.

--offline and --dry-run conflict with --allow-network. Case capture conflicts with --dry-run and --no-persist. --no-cache and --no-ledger disable their respective stores independently; --no-persist also disables persistent runtime coordination. Ordinary configuration reads remain allowed.

Each provider attempt rechecks the effective disclosure and admission policy. Publication separately rechecks the fields governing eligibility and hook mode. A revoked permission or changed governing value withholds the affected action; it never authorizes a replacement request. An edit hidden by an unchanged CLI or environment override does not change the effective policy.

How Ranking Works

1. Establish the exact context

The Claude hook uses the incoming prompt as the current request, even when the transcript has not yet recorded it. Context, cursors, and feedback belong to a specific workspace, session, agent branch, and source adapter/producer. An explicit source that fails does not silently fall through to another conversation.

Native JSONL reads process complete records within a bounded tail. Replacement, truncation, compaction, and incomplete final lines are handled explicitly. An empty first transcript can still yield prompt-only context; malformed existing history is a different condition.

Windowing drops reasoning blocks, binary/media payloads, and prior sr advice. Tool summaries preserve invocation/result association and useful failure lines. The latest request gets budget priority, with explicit head/tail truncation for oversized input. Redaction runs on complete bounded fields before truncation, then on the assembled provider payload.

Explicit directives are resolved from the full bounded local request before redaction or truncation. The normalized input envelope carries local identity and events; it is validated and reduced to a separate provider schema, never sent wholesale to Jev.

A normalized import cannot update a native session's observations merely by repeating its session ID. An input without durable session identity gets an invocation-local namespace; unknown attribution disables durable session updates. A terse “continue” needs a recoverable task anchor; an essential missing instruction or attachment yields unavailable and a quiet hook, rather than a guess from incomplete context.

Project signals use language/framework filenames, allowlisted tools on a trusted PATH, and bounded repository-relative dirty paths. Absolute workspace paths and branch names remain local by default. Git inspection disables filesystem-monitor hooks, optional index locks, and submodule traversal; an unsupported safe invocation omits the optional signal instead of executing project helpers.

The library boundary context::signals::collect implements these optional local signals with an explicit authorized workspace and trusted executable roots. Its 250 ms stage intersects the invocation budget; status output is capped at 64 KiB and 100 paths. Non-UTF-8/unsafe paths and partial inventories are reported rather than treated as complete absence. Results are not serializable provider payloads: callers must still redact them. CLI/ranking integration remains separate. Synchronous filesystem/spawn calls retain the subprocess boundary's documented uninterruptible-kernel-I/O limitation.

2. Resolve what is loadable

An explicit --roster FILE replaces discovery. Otherwise, a harness inventory or its visibility adapter determines roots, overrides, plugins, and load targets. The presence of a directory does not mean the selected harness loads its skills. Generic file mode uses explicitly configured roots and exposes uncertain visibility.

Each skill has an opaque stable ID, its actual invocation name, display name, source, content hash, load target, metadata, and visibility. Same-name skills remain distinct where the harness permits; shadowed or ambiguously invocable entries are excluded from hook suggestions.

Invocation restrictions are part of eligibility. Claude skills with disable-model-invocation: true or an effective user-only restriction are excluded from automatic advice; user-invocable: false alone does not exclude agent use. A user-requested manual-only skill resolves as a manual_only reference, without instructing the agent to bypass that restriction by reading its file.

ResourceDefault bound
Hook stdin1 MiB
Normalized context1 MiB / nesting depth 64
Each user/project/policy configuration file256 KiB / nesting depth 32
Explicit roster32 MiB / 10,000 records / nesting depth 64
Transcript tail2 MiB / 2,000 records
Observation ingestion8 MiB per invocation; a separate committed cursor
One transcript record256 KiB
cass stdout8 MiB
Skill file / frontmatter256 KiB / 16 KiB
Discovery10,000 files / 32 MiB parsed bytes
Quill query128 distinct terms / 4,096 Unicode scalar values after escaping
Replay case capture/import16 MiB / nesting depth 64, with per-field limits
Local replay policy64 KiB / nesting depth 32
Streamed evaluation dataset/report256 MiB / 10,000 case records / nesting depth 64
Live decision or artifact summary envelope2 MiB / nesting depth 64
Wide description160 characters
Rerank description / body excerpt1,000 / 700 characters
Serialized provider request / decoded response96 KiB / 2 MiB

Reject duplicate keys and duplicate record definitions within each schema's collection/namespace in normalized inputs, rosters, configuration, replay/evaluation artifacts, frontmatter, and provider responses. The same skill can still be referenced across wide and rerank stages. Ambiguous skill metadata excludes that record and marks coverage partial. Configuration cannot execute interpolation or recursive includes. Evaluation imports use bounded streaming and enforce per-case limits as well as the total cap.

Source snapshots supply both hashes and excerpts. Before publishing a live advisory decision or no-match claim, sr checks membership, precedence, and indexed/wide content that conditioned the decision, plus the entire shortlist's content and restrictions. A changed candidate outside the shortlist or a new overflow match can invalidate the result too. Explicit resolution checks every target and its name's precedence. Trusted adapter generations can avoid a rescan only when they cover all required dependencies. Otherwise bounded re-enumeration/content checks use the same deadline; missing required validation withholds output. A runner-up cannot replace an answer conditioned on stale alternatives. These are last-validation observations, not a freeze of the filesystem; the harness still checks its later load. A supplied roster replaces discovery but grants no new path access or invocation permissions.

sr roster --snapshot FILE explicitly exports an owner-only manifest, bounded to 32 MiB and 10,000 records, without implicitly overwriting an existing file. Manifests can contain private skill names. sr roster --diff FILE compares the snapshot with fresh authorized discovery in the same workspace, adapter, and source namespace. It reports additions/removals, content and restriction changes, shadowing, and invocation-name changes. Incompatible manifests are identified; incomplete source coverage remains unknown rather than becoming a confirmed deletion. Saved paths grant no new read access and cannot restore a removed skill.

3. Retrieve, then compare

Explicit requirements are resolved from the complete visible roster first. They bypass probabilistic retrieval and cannot be vetoed by a low gate. Every requested reference must resolve: missing, ambiguous, forbidden, or conflicting references produce unavailable / explicit-resolution with separate resolution records and no advisory API call. Successful explicit lists are not truncated to top-K; the input limit is 32 explicit references.

For advisory ranking, Quill, the native lexical engine in FrankenSearch, supplies BM25 retrieval when the eligible roster exceeds 254 real skills. Every Choice also includes __none__, for at most 255 total options. Retrieval uses the latest request plus bounded task and error context; a terse “continue” retains useful prior evidence.

Eligible rosterQuill matchesReal skills admitted to Jev
1–254 skillsPrefilter skippedThe full eligible roster
More than 254 skillsAt least 254The first 254 matches under the deterministic ordering
More than 254 skills1–253Only those matches; the set is not padded with nonmatching skills
More than 254 skillsNoneNo provider call; unavailable / retrieval-empty, exit 5

Configured sizes first satisfy 1 ≤ K ≤ M ≤ 32. The effective rerank size is min(M, admitted_wide_count), and the output cap is min(K, effective_M). For example, three Quill matches produce a wide Choice with three skills plus none, a rerank of at most three skills, and at most three returned suggestions. A single match is valid and still competes against none. An initially empty roster is a roster failure; a valid roster reduced to zero by explicit exclusions or proven available references yields a local abstention.

SkillRanker embeds Quill's in-memory index through frankensearch-quill, with default features disabled, the required frankensearch-core document types, bounded indexing/query work, and the caller's Asupersync context. Names and aliases are searchable titles; descriptions and tags are searchable content, rather than stored-only metadata. Documents enter in stable skill-ID order and are committed before querying. Cutoff ties follow the pinned document-ID mapping; re-sorting an already-truncated result cannot recover an omitted tied candidate.

Queries are a deduplicated OR of escaped literal terms. Only the adapter adds the OR separators; conversation text cannot introduce Boolean operators, wildcards, ranges, or field syntax. The limit of 128 distinct terms and 4,096 Unicode scalar values applies after escaping and separators; analysis input and work are bounded too. No analyzed terms means retrieval-empty. Parser diagnostics and truncation are reported. A query that cannot be preserved safely, exhausted query fuel, or an index failure yields unavailable output with quiet hook fallback. Retrieval failures use exit 5; exhaustion of the overall invocation deadline uses 6. Partial work never becomes a complete candidate set, and there is no silent fallback to another engine.

Overflow results identify retrieval: "quill-bm25", the admitted count, and engine/schema provenance. Building, committing, and querying the index consume the same ranking deadline. A TUI can retain a matching roster index across refreshes; a new hook process cannot assume an earlier process's index survives.

Quill is the only lexical search engine used by this project. Tantivy is not used for runtime search, fallbacks, tests, benchmarks, or reference code. The hybrid search facade, legacy lexical engine, Quill gauntlet, and optional oracle/compatibility features are excluded. Dependency checks cover normal, build, and development feature graphs. Quill verification uses native tests and independent expected-result fixtures. SkillRanker does not need the fsfs command, an embedding model, a search service, or an imported foreign index.

The wide pass combines a Choice, phase distribution, and three oriented gates:

needs_skill = mean(
    specialized_method,
    material_help,
    1 - context_suffices
)

needs_skill is a heuristic score. Below the default 0.30 threshold, sr abstains without a rerank. The wording includes planning, analysis, writing, and explanation skills; acting on files is not a prerequisite for needing a method.

When the gate passes, up to eight real candidates proceed to a detailed Choice with another none option and one fit Noul per candidate. If none wins the wide comparison, the detailed comparison still runs when the need gate passes: richer skill excerpts can resolve ambiguity left by short descriptions. The client uses the TypeSafe HTTP API, preserving typed answers and validating every requested option before scoring.

4. Apply eligibility and rank survivors

A candidate is removed if it is excluded, below the fit threshold, or a reusable reference whose relevant content is proven present in the current context epoch. Workflows and unknown usage kinds remain eligible for repeat invocation. A changed shortlist invalidates the in-flight result.

The reusable-reference check needs evidence of the version and rendered content actually present. A source file's current hash cannot establish what a past read consumed. Changed arguments, dynamic content, forked execution, or compaction can invalidate reuse evidence; content presence never renews turn-scoped permissions.

Every remaining candidate must individually beat the none option's raw rerank probability. Ties are excluded. If no candidates survive, sr abstains; local priors and fit blending cannot re-admit a candidate that failed this check.

For each eligible candidate:

eps = 1e-6
clip(x) = min(1 - eps, max(eps, x))
log_odds(x) = ln(clip(x) / (1 - clip(x)))

utility_i = ln(clip(p_rerank_i))
          + w_fit   * log_odds(fits_i)
          + w_prior * prior_delta_i
          + w_phase * phase_match_i

rank_score_i = softmax(utility)_i

Defaults are w_fit = 1.0, w_prior = 0.0, and w_phase = 0.0. Priors and phase weighting are optional evaluated policy choices. A skill is not penalized just because a previous suggestion went unobserved. After compaction, uncertain loaded-state evidence cannot suppress a skill indefinitely.

Explain And Replay A Result

Find where a candidate was lost

--why-not SKILL_ID --explain follows a candidate through discovery, visibility and restrictions, local policy, Quill admission, the wide shortlist, fit/none eligibility, final ordering, and publication. It reports the first decisive exclusion and any later stages actually evaluated.

sr rank --context scratch/context.json --allow-network \
  --why-not SKILL_ID --explain --json

An unevaluated stage reports not-evaluated; an unknown ID reports not-in-snapshot. Neither receives a fabricated zero fit. Explanations include threshold operands, tie handling, content/policy versions, and bounded recovery hints. They do not expand discovery, insert the target into a shortlist, change the ranked result or provider request bytes, or add a provider call. Hints identify actions and arguments for review; they do not execute commands or relax policy.

Save a case and compare local policies offline

Explicit capture turns a surprising result into a reproducible case:

# Opt in to retaining this run's bounded redacted inputs and recorded answers.
# Use a new output path in an existing private directory.
sr rank --context scratch/context.json --allow-network \

_…[view the full README on GitHub](https://github.com/Dicklesworthstone/skillranker)._

// faq

What is skillranker?

Rust CLI powered by Jev from TypeSafe.ai that ranks agent skills for the next step using live session context. Includes Claude Code hooks, structured JSON, abstention, and local feedback. Requires a TypeSafe API key.. It is open-source on GitHub.

Is skillranker free to use?

skillranker is open-source under the NOASSERTION license, so it is free to use.

What category does skillranker belong to?

skillranker is listed under skills in the Claudeers registry of Claude-compatible tools.

7 views
★ 125 stars
unclaimed
updated 19 days ago

// embed badge

skillranker on Claudeers
[![Claudeers](https://claudeers.com/api/badge/skillranker.svg)](https://claudeers.com/skillranker)

// retro hit counter

skillranker hit counter
[![Hits](https://claudeers.com/api/counter/skillranker.svg)](https://claudeers.com/skillranker)

// reviews

// guestbook

0/500

// related in Claude Skills

🔓

An agentic skills framework & software development methodology that works.

// skillsobra/⟨Shell⟩★ 292,190◷ MIT[ claude ]
🔓

Public repository for Agent Skills

// skillsanthropics/⟨Python⟩★ 178,324[ claude ]
🔓

💫 Toolkit to help you get started with Spec-Driven Development

// skillsgithub/⟨Python⟩★ 138,919◷ MIT[ claude ]
🔓

AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs,…

// skillsGraphify-Labs/⟨Python⟩★ 123,800◷ MIT[ claude ]
→ see how skillranker connects across the ecosystem