claudeers.

Projects tagged #ai-evaluation

// ai-evaluation (3)

🔓

Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.

// uncategorizedopenlit/⟨TypeScript⟩★ 2,792◷ Apache-2.0[ claude ]
🔓

Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments…

// skillsUiPath/⟨Python⟩★ 141◷ Apache-2.0[ claude ]
🔓

AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-…

// skillsagentvitals/⟨Shell⟩★ 95◷ MIT[ claude ]