Model Evaluation · Aperivue/medsci-skills
Compute held-out metrics for a medical-imaging model
Runs task-appropriate evaluation code on held-out predictions for medical-imaging models — Dice, AUROC, FROC, or similarity metrics with confidence intervals, calibration, and subgroup breakdowns — producing a per-case results table for reporting.
Good for
- Compute Dice and boundary metrics for segmentation results
- Report AUROC and AUPRC with bootstrap confidence intervals
- Break down model performance by patient subgroup
- Source repository
- Aperivue/medsci-skills
- Category
- Medicine
Open-source skills are maintained by their authors and listed as published, with attribution. Results depend on how well the skill fits your task and material.
A good place to start
Evaluate my segmentation model's held-out predictions and report Dice, HD95, calibration, and performance broken down by subgroup.

Make your next great thing.
Bring a question, a file, or an idea that’s not quite there yet.