Design Ai Benchmarking · Aperivue/medsci-skills
Design rigorous AI-vs-human-expert benchmark studies
Reviews the design of a study that compares AI system outputs to a human-expert panel before any ratings are collected, covering rubric design, calibration probes, reviewer panel setup, and inter-rater reliability targets. Useful for researchers planning an AI evaluation or reader study.
Good for
- Design a rubric to score AI vs expert outputs
- Plan calibration probes for a reader study
- Set inter-rater reliability targets before rating begins
- Source repository
- Aperivue/medsci-skills
- Category
- Medicine
Open-source skills are maintained by their authors and listed as published, with attribution. Results depend on how well the skill fits your task and material.
A good place to start
Help me design a rigorous benchmark comparing our AI system's outputs against expert reviewers.

Make your next great thing.
Bring a question, a file, or an idea that’s not quite there yet.