Agent Evaluation · seb1n/awesome-ai-agent-skills
Design reproducible evaluations for AI agent quality
Helps design evaluation datasets, rubrics, graders, and regression gates for AI agents, so teams can compare prompts or models, measure tool-use reliability, and decide whether an agent is ready for production.
Good for
- Build a test set for an AI agent
- Compare two prompt versions objectively
- Set a release gate before shipping an agent
- Source repository
- seb1n/awesome-ai-agent-skills
- Category
- Coding
Open-source skills are maintained by their authors and listed as published, with attribution. Results depend on how well the skill fits your task and material.
A good place to start
Help me design an evaluation suite with rubrics and a regression gate before releasing a new version of my support agent.

Make your next great thing.
Bring a question, a file, or an idea that’s not quite there yet.