Model Evaluation · Aperivue/medsci-skills

Compute held-out metrics for a medical-imaging model

Runs task-appropriate evaluation code on held-out predictions for medical-imaging models — Dice, AUROC, FROC, or similarity metrics with confidence intervals, calibration, and subgroup breakdowns — producing a per-case results table for reporting.

Good for

  • Compute Dice and boundary metrics for segmentation results
  • Report AUROC and AUPRC with bootstrap confidence intervals
  • Break down model performance by patient subgroup
Source repository
Aperivue/medsci-skills
Category
Medicine

Open-source skills are maintained by their authors and listed as published, with attribution. Results depend on how well the skill fits your task and material.

A good place to start

Evaluate my segmentation model's held-out predictions and report Dice, HD95, calibration, and performance broken down by subgroup.

Make your next great thing.

Bring a question, a file, or an idea that’s not quite there yet.