IEEE ISBI 2023  ·  Cancers 2023

How sensitive is a dose predictor to the contours you feed it?

Deep-learning dose prediction for glioblastoma, and what it can tell you about contour quality

ISBI 2023 Amith Kamath1, Robert Poel1,2, Jonas Willmann3,4, Nicolaus Andratschke3, Mauricio Reyes1

Cancers 2023 Robert Poel2,1, Amith J. Kamath1, Jonas Willmann3, Nicolaus Andratschke3, Ekin Ermiş2, Daniel M. Aebersold2, Peter Manser2,5, Mauricio Reyes1,2

1ARTORG Center, University of Bern    2Inselspital, Bern University Hospital    3University Hospital Zurich    4Center for Proton Therapy, PSI    5Division of Medical Radiation Physics, Bern

In short. A cascaded 3D U-Net reproduces a clinical glioblastoma dose plan to about 0.94 Gy mean error in seconds rather than hours. It is sensitive to realistic redrawings of an organ contour, tracking the dose change a re-optimised plan delivers, and it can be made robust to target shapes the training set never contained by adding a handful of them. That combination is what a dose-aware check on contouring needs.

One axial slice of a glioblastoma case: the planned dose on the left and the model's predicted dose on the right, both as a heat map over the planning CT, with the target volume and organs at risk outlined.
The task. Given a planning CT and the contours — target volume plus 13 organs at risk — predict the 3D dose distribution a treatment planning system would produce. Left, the clinical plan; right, the prediction. This is the largest target of the 20 test cases and the one with the most organs in play.

Abstract

Radiotherapy planning for glioblastoma starts with contours: the target volume and the organs at risk around it. Those contours take hours of a clinician's time, vary between experts, and are only checked dosimetrically much later, once a plan has been computed. A model that predicts the dose distribution in seconds could move that check forward — if the prediction reacts to a contour edit the way a re-optimised plan does.

We trained a two-level cascaded 3D U-Net on 60 curated glioblastoma VMAT cases and evaluated it on 20 held-out ones with the openKBP dose and DVH scores. The model reaches a mean dose error under 1 Gy against plans prescribing 60 Gy. We then tested it two ways: against ten plausible redrawings of one left optic nerve, each with its own re-optimised reference plan; and against target shapes the training set never contained — concave targets and targets split into several lesions — before and after adding six such cases to training. The model tracks the dose changes the re-planned contours produce, and the shape gaps close with a small amount of targeted data, which together support using such a model for dose-aware quality assurance of contours.

What we found

The model

A two-level cascaded 3D U-Net, the architecture that won the openKBP challenge: the second U-Net receives the first one's output concatenated with its own input. 15 input channels at 128³ — the normalised planning CT, the target volume and 13 organ-at-risk masks — and one continuous output channel scaled to 0–70 Gy.

SettingValue
Cohort125 glioblastoma cases planned to 60 Gy in 30 fractions (double full co-planar VMAT arc, 6 MV, Eclipse AAA)
Split60 train, 15 validation, 20 test; 5 excluded for missing contours
Loss0.5 · L1(reference, coarse) + L1(reference, refined)
Schedule80 000 iterations, He initialisation, five independent runs
Inferencefour-flip test-time augmentation; ~15 s on an A5000, ~45 s on Apple silicon

Both papers use the same architecture and the same 20-case test set. The journal version adds the out-of-distribution target shapes and three retrained models.

Watching it work

Every clip puts the planning CT in grey under a dose heat map, 20–65 Gy, with the target volume in white and one hue per organ at risk. Below 20 Gy the overlay is transparent so the anatomy stays readable, and opacity rises with dose so the high-dose region is where the eye goes.

Test cases are picked by measured properties — target size, how many organs the plan pushes above 20 Gy, and whether the prediction reproduces the plan's fine streak structure — not by eye.

Plan versus prediction, sweeping up through the headThe largest target in the test set, 278 cc, with seven organs at risk above 20 Gy and the worst dose score of the twenty. The prediction tracks the shape of the high-dose region and smooths the gradient at its edge.
A target reaching down to the skull baseHere the target touches the chiasm, both optic nerves, the pituitary and the brainstem, so the plan has real competing constraints. This is also the case whose fine radial streak structure the model reproduces best.
Ten plausible left optic nervesEach frame is one redrawing of the nerve, drawn solid with the original dashed beside it, and each has its own re-optimised plan. Left, the plan; right, the prediction. Under each panel is the mean dose that contour receives, and the subtitle gives the shift from the original contour for both — the two quantities the correlation above is computed on.
A second large target235 cc, and a case where the plan pushes only one organ above 20 Gy — so the high-dose region is shaped by the target alone. The prediction's dose gradient is visibly smoother than the plan's.
Two separate lesions, before and after retrainingThe harder of the two shape gaps: with several targets close together the plan carves dose between them, and this is where the retrained model gains least.
A concave target, before and after retrainingThree panels: the plan, the initial model, and the model retrained with six concave cases. Only the model changes across the row. Both this clip and the next use cases the retrained model never saw.

Results

Two openKBP metrics, both in Gray and both lower-is-better. The dose score is the mean absolute error between predicted and planned dose inside a region; the DVH score is the mean absolute difference of dose-volume criteria — mean dose and D(0.1 cc) for an organ, D1/D95/D99 for the target.

Test setInitial+ concave+ multi-lesion+ both
Dose score
Standard, 20 cases0.940.940.920.98
Concave targets0.870.810.810.87
Multiple lesions1.300.841.241.02
DVH score, organs at risk
Standard, 20 cases2.011.731.851.89
Concave targets2.111.671.992.08
Multiple lesions3.051.863.052.67

Reading across the first row: retraining does not cost anything on the standard test set. Reading down the concave and multiple-lesion rows: the concave update helps most, and helps on both kinds of hard shape.

Which organs are hardest

Dose score per organ at risk, averaged over the 20 test cases. Small organs sitting in the steep part of the dose gradient are hardest; the whole-volume score is much lower than any of them because most of the volume is far from the target.

Dose score by organ at risk (Gy, lower is better) 0 1 2 3 Chiasm 2.99 Hippocampus R 2.60 Cochlea R 2.43 OpticNerve R 2.27 Eye R 2.21 OpticNerve L 2.12 Hippocampus L 2.10 LacrimalGland R 1.94 Pituitary 1.89 Cochlea L 1.86 Eye L 1.49 LacrimalGland L 1.45 BrainStem 1.40 whole-volume score 0.94

Sensitivity to a redrawn contour

Ten redrawings of one left optic nerve, chosen because the target sits close to it, each validated as plausible by radiation oncologists and each re-planned in the treatment planning system. For every alternative, how far the mean dose to that contour moves from the original — in the plan, and in the prediction.

re-optimised plan model prediction
Change in mean dose to the contour (Gy) 0 2 4 6 8 contour 1 Contour 1: plan 0.15 Gy Contour 1: prediction 0.42 Gy contour 2 Contour 2: plan 0.28 Gy Contour 2: prediction 0.22 Gy contour 3 Contour 3: plan 0.36 Gy Contour 3: prediction 1.03 Gy contour 4 Contour 4: plan 0.44 Gy Contour 4: prediction 0.52 Gy contour 5 Contour 5: plan 2.09 Gy Contour 5: prediction 3.17 Gy contour 6 Contour 6: plan 2.40 Gy Contour 6: prediction 2.49 Gy contour 7 Contour 7: plan 3.03 Gy Contour 7: prediction 1.48 Gy contour 8 Contour 8: plan 4.82 Gy Contour 8: prediction 5.44 Gy contour 9 Contour 9: plan 7.59 Gy Contour 9: prediction 5.63 Gy

The pairs move together: a contour that costs the plan 7.6 Gy costs the prediction 5.6 Gy, and contours the plan barely notices the model barely notices either. Over the nine alternatives the correlation is +0.926, against −0.471 for the Dice coefficient over the same contours. Mean absolute dose change is 2.12 Gy for the plan and 2.04 Gy for the prediction, and no contour variation moved any organ dose by more than 5% of the prescription.

What to do with this

Citation

@inproceedings{kamath2023doseprediction,
  title        = {How sensitive are deep learning based radiotherapy dose prediction models to variability in Organs At Risk segmentation?},
  author       = {Kamath, Amith and Poel, Robert and Willmann, Jonas and Andratschke, Nicolaus and Reyes, Mauricio},
  booktitle    = {2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI)},
  pages        = {1--4},
  year         = {2023},
  organization = {IEEE}
}
@article{poel2023deep,
  title     = {Deep-Learning-Based Dose Predictor for Glioblastoma--Assessing the Sensitivity and Robustness for Dose Awareness in Contouring},
  author    = {Poel, Robert and Kamath, Amith J and Willmann, Jonas and Andratschke, Nicolaus and Ermi{\c{s}}, Ekin and Aebersold, Daniel M and Manser, Peter and Reyes, Mauricio},
  journal   = {Cancers},
  volume    = {15},
  number    = {17},
  pages     = {4226},
  year      = {2023},
  publisher = {MDPI}
}