A Lesion-adaptive Segmentation Approach for Tumor Delineation on FDG PET/CT in Diffuse Large B-cell Lymphoma Patients

This study was designed to develop a strategy for selecting the preferred segmentation method for individual DLBCL lesions on FDG PET-based lesion-specific characteristics at different treatment timepoints. Our findings support previous research that no semi-automated delineation method is superior across all lesions and treatment timepoints [14]. Therefore, there is a need for an adaptive approach based on a model with lesion-specific input to identify the preferred segmentation method for each lesion. In this study, we addressed the latter by deriving a simple lesion-adaptive Decision Rule as the primary outcome and, only secondarily, exploring classical ML classifiers for the same task.

When considering a single segmentation method, MV2 was most consistently preferred, achieving a rate of 3 in 74.2% of segmentations overall and 81.4% at baseline assessments. SUV4.0 was previously recommended as preferred method for delineating DLBCL lesions at baseline, with performance comparable to MV2 but greater ease of use [4, 6, 9, 13]. However, in our study, SUV4.0 was considered preferred in only 64.1% of lesions. Interestingly, MV3 demonstrated better performance on I-PET, particularly for lesions with low contrast or heterogeneous metabolic activity, consistent with our previous findings on segmentation method preference for delineating DS4-5 DLBCL lesions at I-PET [14].

Lesion characteristics, such as SUVpeak, TBRpeak, and SUVbg, substantially affected the delineation performance score of each segmentation method, whereas imaging timepoint and lesion location type had a much smaller impact. Our findings align with Weisman, Berthon, and others, emphasizing that integrating lesional uptake metrics such as SUV and TBRpeak is crucial for improving segmentation accuracy and significantly improving the performance of machine-learning-based automated tumor delineation methods [17, 19, 30].

Using these features, a straightforward lesion-adaptive decision rule was identified that uses only two variables (SUVpeak and SUVbg) and three standard methods (SUV4.0, MV2 and MV3). This decision rule, refined from our earlier proposed strategy, replaces lesional SUVmax (> 10) with SUVpeak (> 8) as the threshold for selecting SUV4.0 and introduces a new branch for lower-uptake lesions based on SUVbg: MV3 is selected when SUVbg is above 0.8; otherwise, MV2 is preferred [14]. This approach is particularly effective because MV3, which requires agreement among three methods to include a voxel, tends to be more accurate in cases of low lesional uptake, thereby minimizing the risk of over-segmentation. However, it can lead to under-segmentation when both lesional and background uptakes are low, as both SUV4.0 and SUV2.5 may fail (or in cases of a high lesional or tumor-to-background ratio due to incomplete segmentations by A50peak and 41%max), making MV2 a better option in such scenarios due to its less strict criteria.

Despite optimization over our earlier I-PET-based strategy, the current Decision Rule still required user interaction for a small subset of lesions (18%) to ensure visually accurate delineation across all lesions. These lesions were typically small and/or low-uptake or immediately adjacent to intense physiological activity; the most extreme positive outliers in Fig. 3 arose from such cases, where the chosen threshold occasionally flooded into kidney, pericardial/myocardial or pleural regions contiguous with liver. In the rare case where repeated thresholding attempts fail to delineate a small lesion near high-physiological uptakes, a pragmatic safeguard is to use a small fixed-size VOI centred on the visually identified lesion for lesion-level or radiomics/dissemination analyses, while applying the Decision Rule to the remainder of the disease burden. Although these flooded VOIs were intentionally left uncorrected in our analysis, the decision-rule approach still yielded lesion volumes more frequently within ± 10% (and, to a lesser extent, within ± 3 mL) of the reference MTV than those obtained with the benchmark SUV4.0 method. While both methods correlated strongly with reference MTVs, SUV4.0 consistently undersegments or misses smaller lesions, more prevalent at I-PET and E-PET (Fig. 3). This tendency toward undersegmentation may necessitate more frequent manual adjustments, particularly in studies focusing on quantitative features like lesion count, DmaxBulk, and other dissemination features [31, 32].

At the patient level, TMTVs derived with SUV4.0, the Decision Rule and the best-performing ML approach all showed firm overall agreement with the reference, indicating that lesion-level biases partly, but not completely, cancel when volumes are summed. However, at interim and EoT PET the Decision Rule provided slightly better TMTV agreement and fewer negative deviations (underestimation) than SUV4.0, which is notable given that the limited number of substantial outliers of the Decision Rule were not manually corrected, in contrast to typical clinical workflows where such cases would be adjusted.

In this multi-class setting, using classical ML classifiers for automated method selection is challenging, because several segmentation methods can be acceptable (rate 3) for the same lesion and there is no single uniquely correct class for model training. Even so, all classical ML classifiers we evaluated achieved similar accuracies (0.79–0.82) with a highest-probability strategy for preferred method selection, and more complex, elaborate schemes (hierarchical rules, restricted method sets, stacked ensembles) did not materially improve performance. Overall, these findings indicate that, in our current setting, increasing algorithmic complexity does not necessarily translate into better method selection than the pragmatic lesion-adaptive Decision Rule.

At the same time, the limited gain from more complex models might also reflect the restricted set of input variables we used. To maintain simplicity and implementability, we confined our models to practical, routinely available inputs with the location variable contributing little beyond the uptake metrics, as expected for a non-organ-aware metric. Likewise, SUVbg is derived from a generic local background shell and, although it captures local background activity to some extent, cannot fully represent organ-specific uptake patterns. More advanced lesion-level features (e.g., shape, texture, intralesional heterogeneity), combined with explicit anatomical context, could reasonably be expected to improve preferred-method prediction by better capturing both lesion characteristics and organ-dependent background activity patterns, thereby more closely approximating the anatomical/contextual reasoning implicit in expert visual ratings [17]. Therefore, future work could include such extensions on top of the current lesion-adaptive presegmentation framework, once larger, well-annotated lymphoma datasets across treatment phases are available.

It is important to acknowledge that various other advanced, machine- or deep learning-based semi- and full-automated lesion segmentation algorithms have been proposed in DLBCL and other FDG-avid lymphomas [19, 30]. While several AI-based models for TMTV predictions show considerable promise, most have been trained and validated on baseline scans from single-centre cohorts, with single-observer “ground truth” and presegmentation-based VOIs derived from a single method, with limited assessment across treatment phases or observers. Their generalisability to DS4-5 DLBCL at interim and EoT PET is therefore limited. In HOVON-84, only 33 of 574 patients had DS4-5-positive scans at both these timepoints. Still, these yielded 598 lesions and 3,588 rated segmentations, which we consider adequate for the technical question addressed, while recognising that external validation of the Decision Rule on technical generalisability and reproducibility in larger, independent cohorts will be essential [27].

At baseline, SUV4.0 and the Decision Rule produced highly concordant TMTVs. Substituting the Decision Rule for SUV4.0 is thus unlikely to change established baseline models such as the IMPI materially. At present, however, there are no validated prognostic or treatment‑predictive models in DLBCL that incorporate interim/EoT TMTV or longitudinal TMTV change, so we cannot yet determine whether the Decision Rule improves prognostic performance over SUV4.0 at these timepoints and whether this will be affected by missing or underestimating smaller lesions when using SUV4.0.

Developing prognostic and predictive PET models for (early) treatment-adaptive strategies that explicitly incorporate interim and EoT PET-based MTV or other radiomics requires large, high-quality datasets, particularly because increased FDG uptake at these timepoints may partly reflect inflammation/repair, making MTV/TMTV and derived radiomics biologically more complex and their prognostic value potentially harder to establish even with expert-guided segmentations. However, this subgroup of incomplete DLBCL responders, characterized by significantly poorer outcomes, is rare. Consequently, adequately powered prognostic analyses will require pooling of multiple small patient cohorts across centres and trials, which is only scientifically meaningful if lesion delineation methodology is harmonised so that MTV/TMTV and dissemination PET metrics are comparable. Early delta‑radiomics studies in DS4-5-positive subsets reported promising signals, especially for dissemination features, but all identified non-standardised lesion delineation beyond baseline as a major limitation [33,34,35,36]. Notably, two such studies applied a 41% SUVmax threshold at I-PET, an approach that performed poorly at these timepoints in our data (Supplementary Table 1). Moreover, Baseline MIP-CNN work on HOVON-84, externally validated on PETAL and extended across additional trials, showed that a segmentation-free deep learning model was outperformed by segmentation-dependent PET and clinical PET models incorporating MTV, DmaxBulk, and SUVpeak, in predicting 2-year time-to-progression [37]. This explicitly supports the relevance of DLBCL segmentations for improving prognostically valuable radiomics-ML and DL models at baseline PET and further underscores the need for rigorous segmentation methodology when extending such models to interim and EoT PET.

The present lesion-adaptive Decision Rule can be viewed as a complementary, transparent tool that improves and standardises the semi-automatic selection of the preferred segmentation method for individual DLBCL lesions, regardless of treatment phase, and may serve as a predelineation tool for radiomics-ML or DL-based PET model development in future prognostic or response-adapted DLBCL studies. Reducing method-related variability in lesion segmentation, particularly at interim and EoT PET, helps to bridge the current methodological gap between baseline-focused segmentation practice and the needs of longitudinal, treatment-adaptive PET studies in DLBCL.

Comments (0)

No login
gif