Neural network-driven direct CBCT-based dose calculation for head-and-neck proton treatment planning

Proton therapy has established itself as a highly effective treatment modality for cancer patients. However, the unique physical properties of proton beams, characterized by the Bragg peak phenomenon, offer superior dose conformity compared to conventional photon therapy, whilst introducing heightened sensitivity to anatomical variations (Newhauser and Zhang 2015). Daily anatomical changes, including tumor shrinkage, weight loss, organ motion, and setup variations, can result in substantial alterations to dose distributions that may compromise target coverage or exceed normal tissue tolerance limits (Zhang et al 2007, Paganetti et al 2021). To mitigate these effects, many centers employ daily image guidance and adaptive workflows (Albertini et al 2020, Paganetti et al 2021).

Most modern proton therapy facilities are equipped with cone beam computed tomography (CBCT) systems for daily setup verification (Hua et al 2017). CBCT provides three-dimensional anatomical information at the treatment position, making it an attractive modality for dose calculation in adaptive therapy workflows. However, direct dose computation on CBCT is challenged by scatter, beam hardening, motion artifacts, and limited soft tissue contrast relative to planning CT (Giacometti et al 2020, Liu et al 2023).

A core difficulty is mapping CT numbers (Hounsfield units, HU) to stopping-power ratios for accurate proton range calculations (Landry et al 2019). CBCT often exhibits systematic HU biases of 50–100 HU vs fan-beam CT used for treatment planning, with spatial non-uniformities that can translate to $ \gt $5 mm range uncertainty in heterogeneous regions (Peters et al 2023). These effects degrade dose accuracy and motivate correction strategies.

Traditional approaches to CBCT-based dose calculation employ multi-step correction workflows that attempt to address image quality limitations through various methodologies (Giacometti et al 2020, Liu et al 2023). Scatter correction algorithms utilize physical models or measurement-based approaches to estimate and subtract scatter contributions from CBCT projections (Trapp et al 2022). In contrast, synthetic CT generation methods employ image registration techniques to deform planning CT images to match daily CBCT anatomy, creating hybrid images with improved CT number accuracy (Landry et al 2015). Alternative approaches include histogram matching (Arai et al 2017), intensity correction (Kurz et al 2015), and density override (Dunlop et al 2015) methods that aim to standardize CBCT images for dose calculation applications.

Despite technological advances, these correction methods introduce additional computational complexity, potential error propagation, and workflow inefficiencies that limit their practical implementation in time-constrained adaptive therapy protocols (Shen et al 2020). The multi-step nature of traditional approaches increases the overall treatment planning time and introduces multiple potential failure points that require quality assurance verification (Liu et al 2023). Furthermore, the accuracy of correction methods often depends on the quality of the original CBCT images and the magnitude of anatomical changes, leading to variable performances across different clinical scenarios (Giacometti et al 2020).

The emergence of deep learning methodologies in medical physics has opened new possibilities for addressing the CBCT dose calculation challenge through end-to-end learning approaches (Cui et al 2020, Wang et al 2021). Recent studies have explored various neural network architectures, including convolutional networks, long short-term memory (LSTM) (Hochreiter and Schmidhuber 1997), and Transformer (Vaswani et al 2017) approaches for dose prediction applications (Neishabouri et al 2021, Pastor-Serrano and Perkó 2022).

Sequence modeling approaches, particularly extended LSTM (xLSTM) architectures (Beck et al 2024), have shown exceptional promise for medical imaging applications. The xLSTM architecture offers enhanced memory mechanisms, improved computational efficiency, and superior performance compared to traditional LSTM implementations. Recent applications in MRI-based dose calculation have demonstrated close to Monte Carlo (MC) accuracy with substantial speed improvements, establishing the feasibility of sequence-based approaches for radiotherapy planning (Li et al 2025).

Building upon these advances, this study investigates the adaptation of xLSTM-based neural network architectures for direct CBCT-based proton dose calculations. The approach leverages the sequence modeling capabilities of xLSTM to capture complex spatial relationships in beam-wise dose deposition patterns while eliminating the need for traditional CBCT correction workflows. The methodology enables direct transformation from CBCT images to accurate dose distributions, with potential advantages in adaptive proton therapy workflows through improved computational efficiency and enhanced clinical practicality.

The primary objectives of this study were to validate a deep learning approach against MC reference calculations using clinical patient data with realistic anatomical variations, to develop a deep learning-based dose calculation engine capable of accurate proton dose prediction directly from CBCT images, and to characterize model performance across different beam and anatomical configurations. Secondary objectives included assessment of dose-volume histogram preservation for critical structures and demonstration of computational efficiency suitable for integration into clinical treatment planning workflows.

2.1. Dataset and patient cohort

This retrospective study utilized paired CT–CBCT images from 45 head-and-neck cancer patients treated with intensity-modulated proton therapy at our institution between 2021 and 2023. The dataset comprised 77 paired CT–CBCT images, with each patient contributing 1–4 image pairs. All patient data were anonymized for research purposes following institutional review board guidelines.

Patient imaging was performed using standardized clinical protocols to ensure consistency across the dataset. Planning CT images were acquired using a Siemens Sensation Open scanner with 120 kV and reconstructed at 1 mm × 1 mm × 2 mm voxel size with a soft tissue kernel. Daily CBCT images were obtained using ProBeam® Proton Therapy System with a standard head-and-neck protocol (Varian Medical Systems, Palo Alto, CA, USA), reconstructed at 0.56 mm × 0.56 mm × 2 mm resolution and subsequently resampled to match planning CT resolution.

Image registration was performed using a two-step approach to ensure accurate anatomical alignment between planning CT and corresponding CBCT images. Initial rigid registration was performed using the ANTsPy (Tustison et al 2021) library with mutual information-based algorithms to establish global alignment. Subsequently, deformable registration was applied using an internally developed deformable registration tool (Li et al 2024) to account for soft tissue deformations. Registration accuracy was verified through comprehensive visual inspection to ensure adequate anatomical alignment quality.

For model development, the dataset was systematically divided into training, validation, and independent test cohorts following standard machine learning practices. The training dataset included 33 patients (55 image pairs) for model parameter optimization. The validation dataset comprised seven patients (12 image pairs) for hyperparameter tuning, early stopping criteria, and model selection during the development process. Critically, an independent test dataset comprised five additional patients who were completely held out from all aspects of model development, including training, validation, and hyperparameter selection. These five patients represent truly unseen data and were used exclusively for final clinical performance evaluation. The independent test patients underwent patient-specific fine-tuning (using only their planning CT data) followed by comprehensive validation on their treatment CBCT images to assess clinical performance under realistic adaptive therapy scenarios.

2.2. Beam configuration sampling and data augmentation

Comprehensive beam sampling was performed based on the PSI Gantry2 proton therapy setup (Pedroni et al 2011) to generate a diverse training dataset. For each CT–CBCT image pair, 1500 proton beams were systematically sampled using randomized parameters centered on the image isocenter, with stochastic variations to reflect clinical scenarios.

Beam sampling employed a probabilistic approach with the following parameters: gantry angles were randomly sampled from −30∘ to 180$\circ$, and couch angles from −180$\circ$ to 180$\circ$. Nozzle extraction distances were sampled from 1 to 27 cm, representing the full range of air gaps available with the PSI Gantry2 configuration and enabling simulation of various beam delivery scenarios from close-proximity to extended-distance treatments. Field centers were positioned at the image volume center with random spatial offsets to account for setup variations and target positioning uncertainties.

Each beam configuration included a single pencil beam spot with random 2D lateral offsets ranging from −10 to +10 cm to sample different beam positions within the field. Beam weights were set to 1000 monitor units (MUs), corresponding to approximately 104 protons per MU, ensuring consistent total particle numbers for dose calculations across all configurations. Energy selection utilized a probabilistic distribution derived from the beam energies used in the 40 clinical head-and-neck treatment plans from the training and validation datasets. All beam energy values from these plans were collected and fitted to create a sampling distribution that accurately reflects institutional clinical practice. During training data generation, beam energies for randomly sampled configurations were drawn from this fitted distribution, ensuring the training dataset energy spectrum matched actual clinical energy utilization patterns.

Beam’s-eye-view (BEV) patch extraction represents a critical methodological component that transforms the 3D dose calculation problem into a sequence modeling task suitable for xLSTM processing. For each beam configuration, the CT/CBCT volume was geometrically transformed and resampled to create a BEV perspective aligned with the proton beam direction. The extraction process involved: (1) coordinate system transformation to align the image volume with the beam axis, (2) definition of the extraction volume centered on the beam path, and (3) systematic resampling to create standardized patch dimensions. Each BEV volume encompassed 24 × 24 × 255 voxels at 2 mm isotropic resolution, where the 24×24 transverse dimensions correspond to a 4.8×4.8 cm field-of-view and the 255-voxel depth dimension (51 cm total depth) provided adequate coverage for full-range proton penetration in head-and-neck anatomy. The depth dimension was designed to encompass the complete proton range for clinical energies while maintaining computational tractability for neural network training and inference operations.

2.3. MC simulation and ground truth generation

Ground truth dose distributions were generated using the FRED MC simulation platform (Schiavi et al 2017), which has been validated for proton therapy applications (Gajewski et al 2021). Simulation parameters were optimized to ensure statistical accuracy while maintaining computational feasibility for large-scale dataset requirements.

Each beam simulation utilized 1 million primary proton histories to achieve statistical uncertainties below 1% in high-dose regions ($ \gt $50% of maximum dose) and below 2% in intermediate-dose regions (10%–50% of maximum dose). Dose grid resolution was maintained at 2 mm isotropic spacing to match BEV patch dimensions and provide adequate spatial resolution for capturing typical lateral dose gradients in proton therapy.

Output dose matrices were normalized to absolute dose per beam and subsequently scaled to clinical prescription levels for training purposes. Parallel computing infrastructure utilizing NVIDIA RTX 4090 GPUs enabled efficient MC calculations, with average simulation times of 2 seconds per beam configuration.

2.4. Neural network architecture and implementation

The CBCT-NN architecture implemented an encoder-decoder framework enhanced with xLSTM-based sequence modeling specifically designed for beam-wise proton dose prediction (figure 1). The architecture follows a modular design that processes BEV patches through hierarchical feature extraction, sequence modeling, and dose reconstruction stages. It is important to note that this model is designed as a dose calculation (recalculation) engine, not a treatment plan optimization tool. Consequently, at inference (test time), the model’s inputs consist only of the CBCT image volume and the relevant beam metadata (e.g. energy). Contours (such as CTV or OARs) are not required or used as inputs. This contour-free design was an intentional choice to prevent biasing the physics prediction with planning ‘intent’ and to ensure the model provides a faithful, unbiased dose recalculation for daily plan quality assurance.

Figure 1. CBCT-NN method overview and processing pipeline.

Standard image High-resolution image

The encoder component (ConvEncoder) utilized a hierarchical 2D convolutional neural network structure with three encoding blocks to process BEV patches slice by slice. The first convolutional layer employed 64 feature channels with 5×5 kernels, followed by group normalization (16 groups) and SiLU activation functions. The second convolutional layer maintained 64 channels with identical kernel configuration, while the third layer projected features to the desired embedding dimension. Max pooling operations with 2×2 kernels and stride 2 were applied after the first two convolutional layers to progressively reduce spatial dimensions while preserving essential spatial information.

Spatial positional encoding was implemented using learnable embeddings that preserved sequential relationships across the depth dimension of BEV patches. These embeddings were crucial for maintaining spatial coherence during sequence modeling operations and enabling accurate dose gradient reconstruction in subsequent processing stages.

An innovative energy token encoding system was implemented to condition the network on beam energy. We utilize a learnable energy dictionary (i.e. an embedding layer) that maps each discrete energy level available on the machine to a d-dimensional vector. Concretely, for a set of K discrete energies, we learn an embedding matrix $W_E \in ^$. For a given proton pencil beam with energy index k, the corresponding energy token $e_k = W_E[k]$ is obtained via lookup. This ek vector is then prepended as a conditioning token to the input sequence before the xLSTM blocks. The entire embedding matrix WE is trained end-to-end with the rest of the model, ensuring that energy-dependent physics effects, including range modulation and lateral scattering variations, were appropriately captured throughout the dose prediction process.

The xLSTM sequence modeling component formed the core innovation of architecture, leveraging enhanced memory mechanisms to capture complex spatial dependencies in dose deposition patterns (Beck et al 2024). The xLSTM encoder utilized a block stack configuration incorporating both matrix LSTM (mLSTM) and scalar LSTM (sLSTM) variants. The mLSTM blocks employed four-head attention mechanisms with 1D convolutional kernels (size 4) and QKV projection block sizes of 4, while sLSTM blocks featured CUDA-optimized backends with four-head attention and power-law block-dependent bias initialization. The configuration included 2 blocks total with sLSTM positioned at the second block, dropout rate of 0.2, and feedforward networks with 1.3 × projection factors and GELU activation functions.

The decoder architecture (ConvDecoder) implemented symmetric expansion paths that progressively reconstructed full-resolution dose distributions from encoded sequence features. The decoder featured three main processing stages with skip connections from the encoder’s convolutional layers. Each stage employed 5 × 5 convolutional kernels with group normalization and SiLU activation functions, followed by nearest-neighbor upsampling (factor 2) to restore spatial resolution.

Skip connections between corresponding encoder and decoder levels preserved fine-grained spatial information by concatenating encoder feature maps with decoder features at matching spatial resolutions. The decoder processed the sequence output by first removing the energy token through slicing operations and reshaping the remaining tokens to match the spatial dimensions, then applying the hierarchical reconstruction process to produce the complete 3D dose distribution.

The final output layer consisted of a single 5 × 5 convolutional filter that generated the predicted dose distribution for each BEV slice, with the complete output reshaped to match the original beam geometry for clinical dose calculation applications.

2.5. Training protocol and optimization strategy

Model training employed a multi-stage approach designed to maximize performance while ensuring stable convergence and generalization to unseen data. The training protocol incorporated population-based pre-training followed by patient-specific fine-tuning to balance between broad applicability and individualized accuracy.

The initial pre-training phase utilized the dataset of 82 500 paired BEV patches (55 image pairs × 1500 beams). The LAMB optimizer was employed with an initial learning rate of $1\times 10^$, which was reduced by a factor of 0.5 when validation loss plateaued for more than 10 and 20 epochs. The loss function utilized mean absolute error between predicted and reference dose distributions to ensure dose calculation accuracy. This simplified loss formulation focused model optimization on fundamental dose prediction accuracy while maintaining training stability.

2.6. Validation and performance metrics

Comprehensive validation was performed using an independent test dataset of 5 patients (separate from both the 33-patient training and 7-patient validation cohorts), each with treatment plans optimized on the planning CT. For each patient, two distinct imaging pairs were utilized: (1) a planning-phase CBCT–CT pair for fine-tuning, where the CBCT was acquired during the first treatment fraction (temporally closest to planning CT acquisition), and (2) a late-treatment CBCT–CT pair for testing. To clarify terminology, a ‘control CT’ (also referred to as ‘repeated CT’ or ‘re-CT’) is a diagnostic-quality fan-beam CT acquired intermittently during the treatment course, based on clinical indication, and on the same day as a daily CBCT. This same-day ‘control CT’-not the original planning CT-was used to generate the MC ground truth dose for validation. The test data pair consisted of the CBCT and control CT acquired on the same day during the treatment fraction closest to the end of the treatment course. This time point was chosen to evaluate the model on the latest available anatomy, which generally represented the largest observed anatomical deviation from the planning CT within our dataset.

Patient-specific fine-tuning was implemented to simulate realistic adaptive workflow scenarios and represents a key methodological innovation of this approach. For each test patient, the pre-trained population model was fine-tuned using their planning CT data to adapt the model to patient-specific anatomical features, tissue compositions, and geometric characteristics. This fine-tuning process involved generating approximately 1500 beam configurations from the planning CT using the same sampling methodology described above, followed by training for 100 epochs using a reduced learning rate of $2.5\times 10^$ to prevent overfitting while optimizing patient-specific accuracy. The approach mimics a realistic clinical implementation scenario where the initial planning CT scan would be used to calibrate the neural network model for each individual patient prior to treatment delivery, enabling subsequent rapid dose calculations on daily CBCT images throughout the treatment course.

Image pairing and registration procedures were carefully implemented to ensure accurate validation conditions. Planning CT and treatment CBCT images were paired based on temporal proximity (CBCT acquired within 1–2 weeks of planning CT) and anatomical consistency verified through visual inspection by experienced medical physicists. A two-step registration process was performed between each CT–CBCT pair: initial rigid registration using mutual information-based algorithms, followed by deformable registration using an internally developed deformable registration tool to account for soft tissue deformations and anatomical changes between imaging sessions. Registration quality was assessed by examining anatomical landmark correspondence for both bony structures and soft tissue boundaries after the complete registration process. To clarify the validation workflow, DIR was used to register the daily CBCT to the same-day control CT frame, generating a ‘warped-CBCT’. The CBCT-NN dose was then calculated on this ‘warped-CBCT’. To assess the quality of this image registration step, we performed a quantitative and qualitative analysis. Figure 2 presents the deformable registration results for each test case, showing the fused overlay of the ‘warped-CBCT’ and the control CT image. The normalized mutual information score for each case is reported on the figure to quantify the alignment. Potential residual DIR error in the ‘warped-CBCT’ input image nonetheless remains a key limitation of this validation methodology.

Figure 2. Visual and quantitative validation of the ‘warped-CBCT’ inputs for the five test patients. The 3D Mattes mutual information (Mattes MI) score, calculated within the body mask, is reported for each patient.

Standard image High-resolution image

Contour propagation for dose-volume histogram analysis was performed using the established registration transformations. Target volumes and organ-at-risk contours originally defined on the planning CT were propagated to the corresponding CBCT images using the registration parameters. Contour accuracy was verified through visual inspection and manual adjustment where necessary to ensure anatomical correspondence. This approach enabled consistent (DVH) comparisons between MC calculations on repeated CT images and CBCT-NN predictions on corresponding CBCT images while maintaining identical geometric conditions for evaluation.

The validation protocol involved recalculating initial treatment plans on both the same-day control (repeated) CT (using FRED MC simulation) and corresponding CBCT (using CBCT-NN) to enable direct comparison under identical geometric conditions. This recalculation-only protocol, using the identical initial beam parameters for both modalities, was intentionally designed to isolate the accuracy of the dose engine and imaging modality, rather than confounding the comparison with effects from plan re-optimization.

Gamma analysis served as the primary validation metric, implemented using standard clinical criteria of 2 mm/2% and 2 mm/3% with global normalization and a 10% low-dose threshold. Mean percentage dose error (MPDE) analysis was performed at various dose threshold levels to assess prediction accuracy across different dose regions. MPDE was calculated as the mean absolute difference between NN predicted and reference MC doses, normalized by the prescription dose, for voxels receiving doses above specified thresholds (5%, 10%, 50%, and 90% of maximum dose). This analysis provided complementary and more sensitive information to the gamma evaluation by quantifying dose accuracy in clinically relevant dose regions.

Dose-volume histogram analysis provided clinically relevant validation through comparison of key parameters for target volumes and organs at risk. Primary endpoints included clinical target volume (CTV) coverage (V95%), parotid gland mean doses for xerostomia assessment, and spinal cord maximum doses.

Computational performance evaluation measured inference times for both single-beam calculations and complete treatment plans to assess clinical implementation feasibility. Timing measurements were conducted using clinical-grade hardware configurations (NVIDIA RTX 4090 GPUs with 24 GB memory) representative of modern treatment planning environments, with multiple repeated measurements to ensure reliability and reproducibility of performance assessments.

3.1. Overall model performance and dose calculation accuracy

The CBCT-NN model demonstrated good performance in direct dose calculation from CBCT images, achieving dose calculation accuracy suitable for clinical implementation across all validation metrics (table 1). Gamma analysis results using 2 mm/2% criteria revealed a mean pass rate of 95.1 ± 2.7% across all test cases. These results are comparable to previous neural network-based proton dose prediction studies that employed similar gamma analysis metrics on planning CT images (Neishabouri et al 2021, Pastor-Serrano and Perkó 2022), demonstrating that direct CBCT-based prediction can achieve similar accuracy to established CT-based neural network approaches. Individual patient results showed consistent performance, with pass rates ranging from 91.0% to 97.5%. To substantiate the importance of the patient-specific fine-tuning (FT) step, we performed an ablation study comparing the performance of the general population model (without FT) to the fine-tuned model on five independent test cases. The results of this comparison, including Gamma pass rates (2%/2 mm) and MPDE in both the high-dose region ($ \gt $90% $D_}$) and the overall body, are summarized in table 2. The FT model resulted in a marked accuracy gain, as demonstrated by the clear improvements shown in table 2. These results confirm that FT is a crucial step for achieving high-fidelity, patient-specific dose prediction.

Table 1. Comprehensive model performance metrics and clinical validation results for CBCT-NN across 5 test patients. Dose percentages are reported relative to the prescribed dose.

PatientGamma (2 mm/2%)Gamma (2 mm/3%)CTV V95 (%)Parotid $D_\mathrm $ (%)Spinal Cord Dmax (%)CT-MC vs CBCT-NNCT-MC vs CBCT-NNCT-MCCBCT-NNCT-MCCBCT-NNCT-MCCBCT-NN191.092.295.292.662.861.344.244.3297.597.899.599.6————397.198.097.296.857.558.667.065.2496.097.299.298.839.740.148.246.2594.295.098.598.747.345.311.815.4

Table 2. Per-case ablation of patient-specific fine-tuning (FT): w/ FT vs w/o FT on five test cases.

PatientGamma (2 mm/2%) $\uparrow$MPDE High (%) $\downarrow$MPDE Body (%) $\downarrow$w/ FTw/o FTw/ FTw/o FTw/ FTw/o FT190.9882.692.462.667.9811.05297.4596.992.133.004.394.45397.1082.901.121.214.608.20495.9785.262.333.174.638.54594.1793.424.978.127.887.973.2. MPDE analysis

MPDE analysis (calculated as mean absolute differences) provided detailed assessment of prediction accuracy across clinically relevant dose regions. For high-dose regions ($ \gt $90% of prescribed dose), the mean MPDE was 2.6 ± 1.4%. Intermediate-dose regions (50%–90% prescribed dose) showed mean MPDE of 3.3 ± 0.9%, while lower-dose regions (10%–50% prescribed dose) achieved 6.2 ± 2.0%. The overall body MPDE averaged 5.9 ± 1.9%. Detailed per-patient values are summarized in table 3. Higher accuracy was observed in high-dose regions compared to lower-dose regions. The spatial distribution of dose errors is illustrated through detailed dose maps and error visualizations (figure 3).

Figure 3. Spatial dose distribution analysis and error visualization comparing CBCT-NN predictions with Monte Carlo reference calculations across five patient cases with varying anatomical complexity.

Standard image High-resolution image

Table 3. Mean percentage dose error (MPDE, %) by dose region. MPDE is computed as the mean absolute percentage difference.

Patient$ 90\%\,D_$50%–$90\%\,D_$10%–$50\%\,D_$Overall (body)12.464.318.397.9822.133.874.144.3931.122.025.424.6042.332.724.674.6354.973.638.357.883.3. Dose-volume histogram analysis and clinical relevance

DVH analysis demonstrated good preservation of clinical parameters (figure 4). Target volume analysis revealed minimal deviations between CBCT-NN predictions and MC reference calculations. CTV V95% showed mean differences of −0.6 ± 1.1%, indicating a slight conservative bias (table 1).

Figure 4. Dose-volume histogram comparison for target volumes and critical organs comprehensive DVH analysis comparing CBCT-NN predictions (dashed lines) with Monte Carlo reference calculations (solid lines) across multiple patient cases.

Standard image High-resolution image

OAR analysis showed good performance for critical structures. Parotid gland mean dose predictions had mean differences of −0.5 ± 1.5% compared to MC calculations. Spinal cord maximum dose assessment revealed mean differences of 0.1 ± 2.5%. Close agreement between predicted and reference DVH curves for all critical structures is demonstrated across multiple patient cases (figure 4).

3.4. Computational performance and treatment planning efficiency

Computational performance analysis demonstrated timing characteristics suitable for clinical implementation using clinical-grade hardware (NVIDIA RTX 4090 GPUs with 24 GB memory).

Individual beam dose calculations required an average of 3 milliseconds per pencil beam. Complete treatment plan calculations for typical head-and-neck cases (40 000–50 000 pencil beams) were completed in 1–3 min. The neural network approach provides substantial computational advantages over traditional MC methods, though specific comparisons depend on MC implementation and required statistical accuracy.

This study establishes a proof-of-concept for direct CBCT-based proton dose calculation using xLSTM neural networks, achieving accurate dose prediction without traditional correction workflows. The gamma pass rates of 95.1 ± 2.7% using 2 mm/2% criteria are consistent with previous neural network-based proton dose prediction studies on planning CT images (Neishabouri et al 2021, Pastor-Serrano and Perkó 2022), demonstrating that our direct CBCT-based approach achieves comparable accuracy to established CT-based deep learning methods while eliminating the need for image correction workflows. The computational efficiency of under 3 minutes for complete treatment plans significantly enhances the practical utility of CBCT-based dose calculations.

The xLSTM-based architecture represents a significant methodological advancement over traditional sequence modeling for radiotherapy applications. Enhanced memory mechanisms enable capture of complex spatial dependencies while maintaining practical processing times (Li et al 2020). The energy token encoding system provides a novel approach to incorporating physical parameters, enabling dynamic adaptation to energy-dependent physics effects. The BEV sequence modeling successfully transforms the complex 3D dose calculation problem into a tractable sequence prediction task while preserving essential spatial relationships.

Comparison with conventional approaches highlights several advantages. Traditional CBCT correction workflows require multiple processing steps including scatter correction, image registration, or synthetic CT generation, each introducing potential error sources and computational overhead. The direct prediction approach eliminates these intermediate steps while achieving comparable accuracy. Deep learning approaches using traditional convolutional architectures struggle with long-range spatial dependencies essential for accurate proton dose calculation, while Transformer-based approaches have been shown to yield inferior performance for this specific dose prediction task (Li et al 2025). The xLSTM approach, by contrast, proves more effective, demonstrating a superior capability to model the complex spatial dependencies required for accurate dose calculation.

The demonstrated accuracy and computational efficiency position this approach as particularly suitable for adaptive radiotherapy applications. The ability to rapidly calculate accurate doses directly from CBCT images eliminates time-consuming correction workflows, enabling practical implementation of online adaptive protocols (Albertini et al 2020) where rapid dose assessment is critical for treatment adaptation decisions.

The patient-specific fine-tuning methodology presents both advantages and limitations. The primary advantage lies in adapting the model to individual patient characteristics using readily available planning CT data. However, this introduces an additional processing step requiring computational resources and time (approximately 30 min per patient). Future research could explore alternative personalization strategies, including few-shot learning approaches, to reduce computational overhead while maintaining patient-specific optimization benefits.

The comprehensive validation methodology provides robust evidence for clinical applicability. The focus on independent test data with realistic anatomical changes ensures reported performance reflects true clinical scenarios rather than optimistic laboratory conditions. The validation using repeated imaging sessions during actual treatment courses demonstrates model performance under conditions representative of adaptive therapy implementation.

Several considerations are important for broader clinical implementation. The current validation was performed on head-and-neck cancer patients, providing focused validation in anatomically consistent regions. Extension to other treatment sites including thoracic and abdominal regions represents a natural progression that will benefit from the established methodology while accounting for site-specific challenges such as respiratory motion and larger anatomical variations (Alraddadi 2021). The demonstrated robustness in head-and-neck applications provides confidence for systematic expansion to additional anatomical sites. Several important

Comments (0)

No login
gif