MIDL 2024 (oral)  ·  Radiotherapy and Oncology 2025

Which contour edits actually matter?

Seven clinicians and a dose-prediction model, judging the same 54 glioblastoma target edits

MIDL 2024 Amith Kamath1,*, Zahira Mercado1,*, Robert Poel2, Jonas Willmann3, Ekin Ermiş2, Elena Riggenbach2, Nicolaus Andratschke3, Mauricio Reyes1,2

RadOnc 2025 Jonas Willmann3,4,*, Amith Kamath1,*, Robert Poel1,2, Elena Riggenbach2, Lucas Mose2, Jenny Bertholet2, Silvan Müller2, Daniel Schmidhalter2, Nicolaus Andratschke3, Ekin Ermiş2,*, Mauricio Reyes1,2,*

1ARTORG Center, University of Bern    2Inselspital, Bern University Hospital    3University Hospital Zurich    4Memorial Sloan Kettering Cancer Center    *equal contribution

In short. Fourteen glioblastoma patients, 54 plausible edits to the target volume, every one re-planned so its dosimetric consequence is known. Seven experienced clinicians, shown only the contours, agreed with one another on barely half of the 21 pairings, called 42% of harmless edits harmful, and never once said an edit helped, though four of them did. A dose-prediction model, given the same contours, beat all three oncologists it was benchmarked against in the conference study.

Four panels: an axial CT slice with the reference target contour in white and an edited one in cyan, the edited region filled yellow; the re-optimised dose plan as a heat map; the change in dose as a red and blue diverging map; and a strip of coloured verdicts in which every judge says the edit was harmful while the ground truth says it was not.
One edit, and everything anyone said about it. The target volume was enlarged by 1.7 cc (yellow, left). The re-optimised plan (middle) redistributes dose, and the change it makes (right) stays inside every clinical constraint. All seven clinicians, and the model, called this edit harmful. It was not.

The problem

Automatic segmentation now draws most radiotherapy contours, which turns the clinician into a reviewer. A reviewer has to decide, for each difference they see, whether it is worth correcting: correcting everything throws away the time auto-contouring was meant to save, correcting nothing risks the patient. That decision is dosimetric, but it is made by looking at outlines, long before any dose is computed.

So we measured how well it is made. Fourteen glioblastoma patients, up to four plausible edits each to the clinical target volume, and every one of the 54 results re-planned in Eclipse, which establishes what each edit does to the 13 organs at risk. Then we asked clinicians to predict that from the contours alone, and asked a dose-prediction model the same question.

What we found

How "impact" came to be defined, twice

The conference study needed a rule it could compute for all 54 re-planned edits, so it used a relative one: an edit is impactful if the maximum dose to at least one organ moves by more than 10%. Presenting that work, the clinicians pushed back on the rule itself. A 10% change in maximum dose is not on its own clinically meaningful. Ten percent added to an organ far below its tolerance changes nothing for the patient, while a much smaller rise on an organ already near its limit can change the plan.

The journal study therefore rebuilt the ground truth around the constraints clinicians plan against: +1 for each organ an edit pushes past its guideline limit, −1 for each it brings back under, summed over all 13.

Ground truthTestVerdicts
MIDL 2024 the maximum dose to at least one organ at risk moves by more than 10% 33 worse, 17 no change, 4 better
RadOnc 2025 a clinical dose constraint is crossed, scored +1 or −1 per organ and summed over all 13 5 worse, 45 no change, 4 better

The two labellings agree on 18 of the 54 edits. Of the 33 the first rule called harmful, only 3 cross a clinical limit; 27 are neutral and 3 are improvements. It also explains the dose predictor's two scores: its own rule is a relative one, so it does well against the first definition and its recall falls to 0.33 against the second.

Who judged, and when. The conference paper reported three radiation oncologists; the journal reports seven, adding a fourth oncologist whose answers came from the same round and three medical physicists surveyed afterwards. None could have been influenced by the earlier result: all four oncologists were surveyed in January and February 2024, before the conference paper appeared, and their columns in the journal data are identical to the answers recorded then. The physicists judged under the same protocol, with no access to the dose distribution, the ground truth, the earlier results or each other's answers.

Watching it happen

Each clip sweeps through one patient. Left is what the evaluators were given: the reference target in white, the edited one in cyan, with tissue the edit added and tissue it removed filled in. Middle is the re-optimised plan. Right is what the edit did to the dose, which nobody was allowed to see. The strip carries both ground truths, all seven evaluators and the model, and the two lines beneath give the reason behind each verdict: a crossed constraint for the plan, a relative dose shift for the model. Both are generated from the data, and both rules reproduce all 54 published labels exactly.

All 54 edits, one after anotherThe whole study in one pass. Watch the verdict strip rather than the anatomy: the orange "worse" column runs down it almost unbroken while the ground truth beside it stays blue, and green appears only in the truth row.
Everyone agreed, and everyone was rightAll seven, and the model, called this edit harmless, and the re-optimised plan bears them out.
Everyone agreed, and everyone was wrongA 1.7 cc enlargement all seven called harmful. The dose moves, but never crosses a constraint. Consensus is not evidence.
Harm that nobody caughtThis edit takes the brainstem from 52.6 to 56.2 Gy against a 54 Gy limit. Not one of the seven flagged it; the model did.
Harm that was caughtSix of seven flagged this one, correctly. Where an edit pushes the target hard towards an organ, clinicians read it well; the failures are in the subtle cases.
An improvement nobody sawThis edit brought the left eye back under its limit. Five of seven called it harmful and the other two neutral.
The edit they disagreed about mostThree of seven called it harmful, four did not. Same image, same instructions, same profession.
The model right, most evaluators wrongFour of seven called this harmful; the plan disagrees, and so did the predictor.
The model wrong, evaluators rightIts characteristic failure: it predicts a smooth dose rather than the plan's beam structure, so it over-reads small changes near an organ.
Unanimous, and unanimously wrongEvery evaluator called this harmless, and so did the model. The re-optimised plan crosses a constraint.

Results

How much the evaluators agreed with each other

Cohen's Kappa for all 21 pairs. Above 0.6 is conventionally read as strong agreement; ten pairs fall below it, and each of the five weakest involves R‑1, the evaluator who used "worse" least often and who also scored closest to the ground truth.

two oncologists two physicists one of each
Agreement between evaluators (Cohen’s Kappa) 0 0.2 0.4 0.6 0.8 R-1 / R-3 R-1 / R-3: 0.24 (oncologist pair) 0.24 R-1 / M-2 R-1 / M-2: 0.28 (mixed pair) 0.28 R-1 / M-1 R-1 / M-1: 0.34 (mixed pair) 0.34 R-1 / R-4 R-1 / R-4: 0.37 (oncologist pair) 0.37 R-1 / R-2 R-1 / R-2: 0.41 (oncologist pair) 0.41 R-3 / M-3 R-3 / M-3: 0.54 (mixed pair) 0.54 R-1 / M-3 R-1 / M-3: 0.56 (mixed pair) 0.56 M-2 / M-3 M-2 / M-3: 0.57 (physicist pair) 0.57 R-4 / M-2 R-4 / M-2: 0.59 (mixed pair) 0.59 R-2 / R-4 R-2 / R-4: 0.59 (oncologist pair) 0.59 R-2 / M-2 R-2 / M-2: 0.62 (mixed pair) 0.62 R-4 / M-1 R-4 / M-1: 0.63 (mixed pair) 0.63 R-4 / M-3 R-4 / M-3: 0.63 (mixed pair) 0.63 R-2 / R-3 R-2 / R-3: 0.64 (oncologist pair) 0.64 R-2 / M-3 R-2 / M-3: 0.64 (mixed pair) 0.64 M-1 / M-2 M-1 / M-2: 0.66 (physicist pair) 0.66 R-3 / R-4 R-3 / R-4: 0.67 (oncologist pair) 0.67 R-3 / M-1 R-3 / M-1: 0.67 (mixed pair) 0.67 M-1 / M-3 M-1 / M-3: 0.69 (physicist pair) 0.69 R-3 / M-2 R-3 / M-2: 0.71 (mixed pair) 0.71 R-2 / M-1 R-2 / M-1: 0.74 (mixed pair) 0.74 0.6 = strong agreement; 10 of 21 pairs fall short

How each judge spent their verdicts

Every evaluator had 54 edits and three labels. The ground truth uses all three; no evaluator did. R‑1 called 44 edits harmless and R‑3 called 30 of them harmful, on identical images.

better no change worse
How each judge spent their 54 verdicts 0 18 36 54 R-1 R-1: 44 × No change 44 R-1: 10 × Worse 10 R-2 R-2: 32 × No change 32 R-2: 22 × Worse 22 R-3 R-3: 24 × No change 24 R-3: 30 × Worse 30 R-4 R-4: 27 × No change 27 R-4: 27 × Worse 27 M-1 M-1: 29 × No change 29 M-1: 25 × Worse 25 M-2 M-2: 30 × No change 30 M-2: 24 × Worse 24 M-3 M-3: 37 × No change 37 M-3: 17 × Worse 17 Majority Vote Majority Vote: 33 × No change 33 Majority Vote: 21 × Worse 21 Model Model: 3 × Better Model: 17 × No change 17 Model: 34 × Worse 34 ground truth: 4 better, 45 no change, 5 worse

The model against three oncologists

The conference comparison, on the conference ground truth: weighted precision and recall over the three categories, and the mean time one edit took to judge. The model's advantage is not that it reads anatomy better, but that it answers a dosimetric question by predicting a dose rather than inferring one from the shape of an outline.

precision recall
MIDL 2024: the model against three oncologists 0 0.2 0.4 0.6 0.8 Oncologist 1 Oncologist 1: precision 0.41 0.41 Oncologist 1: recall 0.35 0.35 44 s Oncologist 2 Oncologist 2: precision 0.48 0.48 Oncologist 2: recall 0.46 0.46 50 s Oncologist 3 Oncologist 3: precision 0.55 0.55 Oncologist 3: recall 0.57 0.57 71 s Dose predictor Dose predictor: precision 0.57 0.57 Dose predictor: recall 0.57 0.57 30 s per edit

What to do with this

Citation

@article{willmann2025predicting,
  title     = {Predicting the impact of target volume contouring variations on the organ at risk dose: results of a qualitative survey},
  author    = {Willmann, Jonas and Kamath, Amith and Poel, Robert and Riggenbach, Elena and Mose, Lucas and Bertholet, Jenny and M{\"u}ller, Silvan and Schmidhalter, Daniel and Andratschke, Nicolaus and Ermi{\c{s}}, Ekin and Reyes, Mauricio},
  journal   = {Radiotherapy and Oncology},
  pages     = {110999},
  year      = {2025},
  publisher = {Elsevier}
}
@inproceedings{kamath2024comparing,
  title        = {Comparing the Performance of Radiation Oncologists versus a Deep Learning Dose Predictor to Estimate Dosimetric Impact of Segmentation Variations for Radiotherapy},
  author       = {Kamath, Amith and Mercado, Zahira and Poel, Robert and Willmann, Jonas and Ermi{\c{s}}, Ekin and Riggenbach, Elena and Andratschke, Nicolaus and Reyes, Mauricio},
  booktitle    = {Medical Imaging with Deep Learning},
  volume       = {250},
  year         = {2024},
  organization = {PMLR}
}

Errata

Values in the published papers that do not follow from the released data. None affects a reported conclusion. Full audit in results/radonc_verification.csv and results/midl_verification.csv.

Radiotherapy and Oncology 2025

MIDL 2024