Design Ai Benchmarking · Aperivue/medsci-skills

Design rigorous AI-vs-human-expert benchmark studies

Reviews the design of a study that compares AI system outputs to a human-expert panel before any ratings are collected, covering rubric design, calibration probes, reviewer panel setup, and inter-rater reliability targets. Useful for researchers planning an AI evaluation or reader study.

Good for

  • Design a rubric to score AI vs expert outputs
  • Plan calibration probes for a reader study
  • Set inter-rater reliability targets before rating begins
Source repository
Aperivue/medsci-skills
Category
Medicine

Open-source skills are maintained by their authors and listed as published, with attribution. Results depend on how well the skill fits your task and material.

A good place to start

Help me design a rigorous benchmark comparing our AI system's outputs against expert reviewers.

Make your next great thing.

Bring a question, a file, or an idea that’s not quite there yet.