Agent Evaluation · seb1n/awesome-ai-agent-skills

Design reproducible evaluations for AI agent quality

Helps design evaluation datasets, rubrics, graders, and regression gates for AI agents, so teams can compare prompts or models, measure tool-use reliability, and decide whether an agent is ready for production.

Good for

  • Build a test set for an AI agent
  • Compare two prompt versions objectively
  • Set a release gate before shipping an agent
Category
Coding

Open-source skills are maintained by their authors and listed as published, with attribution. Results depend on how well the skill fits your task and material.

A good place to start

Help me design an evaluation suite with rubrics and a regression gate before releasing a new version of my support agent.

Make your next great thing.

Bring a question, a file, or an idea that’s not quite there yet.