claudeers.

Projects tagged #benchmark

// benchmark (4)

🔓

Quickly find bottlenecks in Rust - one profiler for CPU, time, memory, SQL and async code.

// uncategorizedpawurb/Rust1,681MIT[ claude ]
🔓

Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments…

// skillsUiPath/Python113Apache-2.0[ claude ]
🔓

An agent skill that plays ARC-AGI-3. One rule: say what an action will do before you spend it. Claude Code on Opus 5 finished all 25 public games at 100.00 R…

// skillspbshgthm/Python50[ claude ]
🔓

Turn every Claude Code and Codex session into reusable training, evaluation, and verification assets.

// automationjinzijian/Python26Apache-2.0[ claude ]