claudeers.

Projects tagged #evals

// evals (5)

🔓

AI Observability & Evaluation

// automationArize-ai/⟨Python⟩★ 11,618◷ NOASSERTION[ claude ]
🔓

Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform…

// uncategorizedFrancyJGLisboa/⟨Python⟩★ 2,400◷ MIT[ claude ]
🔓

Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery, updates, rollback, quarantine, and 17-platform…

// uncategorizedFrancyJGLisboa/⟨Python⟩★ 2,401◷ MIT[ claude ]
🔓

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

// mcp-serversMCPJam/⟨TypeScript⟩★ 2,227◷ NOASSERTION[ claude ]
🔓

Make Claude Opus 4.8 behave like Claude Fable 5 — doctrine output style, drift-catching hooks, and an eval loop against golden Fable transcripts. Claude Code…

// pluginsrennf93/⟨Shell⟩★ 34◷ MIT[ claude ]