Comparative analysis of three thyroglobulin immunoassays: analytical performance and clinical implications

The decision to change an analytical method in clinical laboratories is particularly critical when dealing with tumor markers like thyroglobulin (Tg) working as indicators of surgical radicality [12]. Immunoassays face technical challenges, including variability in antibody quality and interference from endogenous antibodies. For instance, monitoring elevated thyroglobulin (Tg) levels via immunometric assays is essential in the follow-up of patients with differentiated thyroid cancer (DTC) to detect residual or recurrent disease [13, 14] .

In recent years, the development of second-generation immunometric assays (Tg-IMAs) and mass spectrometry-based methods has significantly improved the sensitivity and specificity of Tg detection [15, 16].

While LC-MS/MS offers enhanced specificity [17, 18], its application is still limited by cost, technical complexity, and lack of standardization across laboratories [19].

Notably, newer Tg-IMAs demonstrate functional sensitivity below 0.1 ng/mL, enabling the detection of low levels of Tg in patients under TSH suppression therapy [20].

While exploring alternative methods is essential, a thorough evaluation of its impact on patient management must be faced. Inter-assay variability among commercial immunoassays persists despite standardization efforts, such as the use of CRM-457 as a reference material [10]. Studies have demonstrated significant discrepancies in Tg values across different platforms, which may lead to misinterpretation of longitudinal trends in patient monitoring [21, 22].

Our comparative analysis of the Access (Beckman Coulter), Liaison (Diasorin), and Atellica (Siemens) immunoassays demonstrated strong overall concordance across methods, with some variability at different thyroglobulin (Tg) concentration ranges.

Tg-L demonstrated a very strong overall correlation with Tg-B, with moderate correlation at low Tg levels (0–2 ng/mL) and very strong correlations for both 2–50 ng/mL and > 50 ng/mL.

Tg-A showed a very strong overall correlation with Tg-B with moderate correlation at low Tg levels (0–2 ng/mL), very strong correlation for 2–50 ng/mL and strong correlation for > 50 ng/mL. Notably, Tg-B showed a high concordance in detecting negative Tg values (< 0.2 ng/mL) with Tg-L (96%) and Tg-A (98%), with all discrepant cases occurring in TgAb-negative patients.

Despite a slight analytical bias at low Tg concentrations (0–2 ng/mL), the three methods showed strong concordance.

Among the samples falling within the critical range of 0–2 ng/mL, a total of six samples exhibited discordant results across the three Tg assays, with values falling below the functional sensitivity threshold of 0.2 ng/mL in one assay and detectable in another. All these samples were from TgAb-negative patients, and none showed signs of heterophile antibody interference, based on standard laboratory flags and consistent analyte behavior. However, it is important to note that no heterophile-blocking agents were specifically used, which represents a limitation of the study.

Clinical follow-up data were available for four of the discordant cases. All patients had undergone total thyroidectomy for differentiated thyroid carcinoma and were under regular surveillance with no evidence of structural disease at the time of sampling. Importantly, none of these minor discrepancies in Tg levels led to changes in therapeutic management, as Tg values were still interpreted within the “excellent response” category according to ATA guidelines.

Additionally, we assessed whether changing the analytical cutoff from 0.2 ng/mL to 0.1 ng/mL would affect agreement rates. While this adjustment resulted in a slight increase in discordant classifications, the overall concordance between Tg-A and the reference method Tg-B remained high, and the impact on classification of patient response categories was minimal.

Moreover, the Bland-Altman analysis revealed a significant negative bias for Tg-L relative to Tg-B, indicating that Tg-L consistently reports lower values across the measured range. This systematic difference may be attributed to several analytical factors. First, although all three assays are standardized against CRM-457, differences in calibrator formulation, matrix composition, and traceability procedures can lead to divergence in quantification. Additionally, variability in capture and detection antibodies, particularly in their affinity for different Tg isoforms or glycoforms, may contribute to inter-assay discrepancies. Tg-L may be more selective for certain molecular forms of thyroglobulin or less reactive to circulating fragments, resulting in lower measured concentrations. Differences in signal detection technologies and epitope recognition could further exacerbate this underestimation. Understanding these assay-specific biases is critical when interpreting longitudinal Tg trends, especially if a change in assay platform occurs during patient follow-up.

The results highlight the need for re-baselining when switching between the three immunoassays, rather than supporting a general feasibility of transitioning across assays.

It is important to note that the serum samples analyzed in this study were not exclusively obtained from patients with confirmed thyroid cancer, but rather from a heterogeneous population with and without thyroid pathology. This constitutes a limitation of the present work and may partially influence the generalizability of the results in a strictly oncological setting. Furthermore, due to the inherent molecular heterogeneity of thyroglobulin and the variability in antibody design across different immunoassay platforms, we strongly recommend re-baselining Tg values when switching assay methods in longitudinal patient monitoring. This step is essential to ensure accurate trend interpretation and consistent clinical decision-making over time. Although the assays are highly correlated, differences between methods must be taken into account when transitioning between platforms, particularly for consistent longitudinal patient follow-up

Moreover, we suggest that Tg-A results between 0.06 and 0.2 µg/L should be accounted for ‘grey zone’ values and these subjects should undergo a stimulation test until long-term outcomes can be clarified through extended prospective studies to minimize clinical and therapeutic uncertainty and reduce potential anxiety among patients with DTC.

Despite the use of certified reference materials, discrepancies among different immunoassays are well documented. These differences are most likely attributable to variations in the source and composition of kit calibrators, the intrinsic heterogeneity of the analyte, and the use of distinct antibody sets in each assay, which differ in their specificity toward various Tg isoforms [13, 23].

There are currently no universally accepted Tg cut-off values for assessing disease status in DTC. This is due to inter-assay variability, differences in assay sensitivity, and individual patient factors. As a result, monitoring longitudinal trends in Tg levels using the same assay, ideally under stable TSH conditions, provides more reliable clinical information than relying on single absolute values. Consistency in the assay platform is essential to accurately detect subtle changes that may indicate disease recurrence or progression [23, 24].

In our study anti-thyroglobulin antibody (TgAb) status was assessed using a single immunoassay method. We acknowledge this as a limitation, since confirming TgAb negativity with two independent methods would have provided stronger evidence for the absence of interference. Nonetheless, all patients included in the study tested negative for TgAb, and TgAb negativity was consistently observed across all samples, including those showing discrepant Tg values. Additional studies involving larger patient cohorts and subjects with positive anti-thyroglobulin antibodies are warranted to confirm these findings and to better define the utility of Tg-A in clinical decision-making, especially when used to monitor patients after thyroidectomy for biochemical recurrence.

In addition, the analytical measurement range (AMR) represents a crucial parameter for assay suitability in routine follow-up [25]. The Atellica assay (Tg-A) demonstrated the narrowest AMR among the three methods tested, spanning from 0.050 to 150 ng/mL, compared to the broader range of 0.1 to 500 ng/mL reported for both Tg-B and Tg-L. The relatively limited upper range of Tg-A may necessitate frequent manual or automated dilutions in patients with residual thyroid tissue or biochemical recurrence. This limitation could affect laboratory workflow and turnaround time, especially in settings managing large volumes of thyroid cancer patients. Therefore, both the risk of hook effect at high concentrations and the constraints of AMR must be taken into account when selecting an assay platform and interpreting longitudinal trends in Tg monitoring.

A major limitation of this study lies in the composition of the sample cohort, which includes residual serum samples from a heterogeneous population of individuals, both with and without thyroid pathology. While this reflects the variety of Tg levels encountered in routine clinical practice and supports analytical comparisons across a broad dynamic range, it limits the clinical specificity of the findings for DTC surveillance. Ideally, assay comparisons should be conducted in a cohort of post-thyroidectomy DTC patients under TSH-suppressed or stimulated conditions, where precise Tg quantification has direct implications for disease monitoring and therapeutic decision-making. Therefore, while the present results provide important analytical insights, further studies in targeted DTC populations are needed to confirm their applicability in a clinical follow-up context. Nevertheless, this comparative analysis highlights the critical need for standardization in Tg measurement and for awareness of assay-specific biases to support accurate risk stratification and effective long-term management of patients with DTC [26].

In conclusion, although limited follow-up data suggest a minimal clinical impact in a subset of cases, further studies are required to validate these findings in a larger DTC cohort.

Meanwhile, a personalized and closely monitored follow-up strategy, as proposed in this study, is recommended.

Comments (0)

No login
gif