Plasma proteomic signatures improve risk stratification and personalized screening for gastric cancer

Study population

In the discovery stage, a case-control study was initially designed with 194 participants, comprising individuals with gastric cancer as the case group, and individuals without cancerous lesions as the non-gastric cancer control group. All participants were recruited from Zhejiang Cancer Hospital between 2022 and February 2023. The case group included newly diagnosed gastric cancer patients, who had not undergone prior treatments such as chemotherapy, radiotherapy, targeted therapy, or biological therapy. Patients with two or more malignancies were excluded from the study. Pathological staging followed the 8th edition of the UICC & AJCC TNM classification. The control group comprised healthy individuals who underwent endoscopy as part of routine physical examinations. Participants in this study ranged in age from 32 to 81 years, with 135 males and 59 females. The case group consisted of 100 gastric cancer patients, classified as follows: 10 in stage I, 22 in stage II, 50 in stage III, and 18 in stage IV. No participants with completely normal gastric mucosa were identified. The control group included 16 individuals with superficial gastritis, 31 with chronic atrophic gastritis, and 47 with intestinal metaplasia. Detailed clinical information for the enrolled participants, including age, sex, body mass index, smoking status, alcohol drinking, family history of cancer, Lauren classification, and tumor-node-metastasis staging, is provided in Fig. 2A and Table S1.

Fig. 2Fig. 2The alternative text for this image may have been generated using AI.

Liquid chromatography-tandem mass spectrometry-based plasma proteomic profiling of gastric cancer and controls in the discovery stage. A Overview of plasma proteomics using data-independent acquisition mass spectrometry. B Comparison of the number of proteins identified in gastric cancer and non-gastric cancer groups. C Proteins identified in gastric cancer and non-gastric cancer groups ranked by median intensity. D Principal component analysis of plasma proteomic profiles in gastric cancer and non-gastric cancer groups. GC, gastric cancer; NGC, non-gastric cancer

The validation cohort was derived from the UK Biobank, with participants recruited between 2006 and 2010. Gastric cancer diagnoses in the UK Biobank cohort were obtained through linkage to hospital data, cancer registries, and death registries, using the International Classification of Diseases (ICD). The time to gastric cancer occurrence was calculated from the cohort entry date until the occurrence of gastric cancer, death, or administrative censoring (October 31, 2022), whichever occurred first. Of the UK Biobank cohort, 92 gastric cancer and 52,460 gastric cancer-free participants with proteomic data were involved in the analysis, with a median (interquartile range) follow-up period of 13.63 (12.90–14.38) years. Baseline characteristics of participants in the validation cohort were shown in Table S2.

Ethics

The discovery stage study was approved by the Institutional Review Board of Zhejiang Cancer Hospital (No. IRB-2025-368(IIT)). Ethics approval for the UK Biobank study was obtained from the North West Centre for Research Ethics Committee (11/NW/0382). The proteomic profiling of the UK Biobank was approved by the Access Subcommittee of UK Biobank under Access Management System Application No. 65,851. All participants provided written informed consent, and the study was conducted in accordance with the Declaration of Helsinki.

Sample collection and processing

Fasting blood samples (5 ml) were collected from each participant prior to any gastroscopy or surgery in the discovery stage. The samples were transported to the laboratory within 1 h at 4 °C for processing. Plasma was separated by centrifugation at 1800 g at 4 °C for 10 min, using EDTA as an anticoagulant. After centrifugation, the plasma layer was collected and stored at − 80 °C for further analysis.

Protein extraction, proteomic profiling, data processing, and protein quantification

Plasma samples were processed using a plasma protein enrichment kit (PH035, Beijing Precise Health Biotechnology Co., Ltd). Briefly, 20 µL of plasma was mixed with 80 µL of enrichment buffer. Subsequently, 20 µL of enrichment magnetic beads were added, and the mixture was incubated at room temperature for 30 min. The enriched sample was washed twice using washing solution. Next, digestion buffer containing 2 µg of trypsin was added, and the solution was vortexed thoroughly. The mixture was incubated at 37 °C overnight for digestion. After digestion, 100 µL of termination buffer was added, and the sample was centrifuged at 16,000 g for 20 min. The supernatant was collected for desalting.

For chromatography, mobile phase A consisted of a 0.1% formic acid (FA) solution, while phase B was 80% acetonitrile (ACN) containing 0.1% FA. Lyophilized peptides were resuspended in 12µL mobile phase A and centrifuged at 16,000 g for 15 min. The supernatant was transferred to a sample vial. 3 µL sample was loaded onto an UHPLC system (UHPLC 3000, Thermo Fisher Scientific, USA) at a flow rate of 3 µL/min. Peptides were separated using a C18 analytical column (1.9 μm, ID150, 30 cm length) at a flow rate of 600 nL/min. Mobile phase B was increased from 8 to 50% in 60 min. The sample was analyzed using a Thermo Scientific Orbitrap Eclipse mass spectrometer (Thermo Fisher Scientific, USA) with a Nanospray Flex ion source. The ion spray voltage was set to 2.2 kV, and the ion transfer tube temperature was set to 320 °C.

The mass spectrometer was operated in Data Independent Acquisition (DIA) mode. The first-stage scan resolution was set to 60,000, with a scan range of 350–1150 m/z and an automatic maximum injection time. The second-stage scan resolution was set to 30,000, with 30 scan windows and a collision energy of 32%. The DIA raw files were processed using Spectronaut 17.4 (Biognosys) software for database searching. The search parameters were set to include trypsin digestion, allowing up to two missed cleavages permitted. Fixed modifications were set to Carbamidomethyl (C), and variable modifications included Oxidation (M) and Acetylation (N-terminal). A false discovery rate (FDR) threshold of less than 1.0% was employed to protein level. The search was conducted against the UniProt human protein database.

The detailed methodology for antibody-based protein profiling using the Olink Explore 3072 platform, including study-wide protein measurements, processing, and quality control procedures, is described in a previous study [10].

Statistical analysesIdentification of plasma proteomic signatures associated with gastric cancer

In the discovery stage, principal component analysis (PCA) was used to visualize the proteomic profiles of subjects with gastric cancer and those without gastric cancer. Differentially expressed proteins were defined as upregulated if the Log2 fold change (Log2FC) was greater than 0, or downregulated if Log2FC was less than 0. Multiple testing correction was performed using the FDR < 0.05 as the significance threshold. In the validation stage, the association between each protein and gastric cancer risk was assessed using Cox proportional hazards regression. Nominal statistical significance (P < 0.05) with consistent direction of effect between discovery and validation stages was used as the primary criterion for replication. To further address multiple comparisons, false discovery rate correction was applied across all tested proteins, and proteins with FDR < 0.05 were considered high-confidence findings.

Polygenic risk score (PRS)

The genotyping, imputation, and quality control procedures employed in the UK Biobank have been described elsewhere [11]. In brief, genotyping was performed by using two similar chips, the UK Biobank Axiom Array and the UK BiLEVE Axiom Array (Affymetrix). Imputation was constructed using the merged UK10K and 1000 Genomes Phase 3. We selected 112 single nucleotide polymorphisms (SNPs) associated with gastric cancer at a genome-wide significant level (P < 5 × 10−8) from the Finngen R12 (Table S3) [12]. The genotype data for each SNP associated with gastric cancer were obtained from the UK Biobank, and a PRS for gastric cancer was calculated using the following equation: \(\:\text\text\text\text\text\text\text\text\:\text\text\text=\sum\:_^__\). Where SNPi represents the dosage of the effective allele for SNPi, βi is the effect estimate for SNP i for gastric cancer derived from previous genome-wide association studies (GWAS), and n is the number of instrumental variables obtained from the GWAS [12] .

Construction of risk prediction model and risk score

Clinical characteristics of participants in the validation cohort were derived from the UK Biobank dataset. Initially, a univariable Cox proportional hazards regression model was used to explore potential clinical factors associated with gastric cancer risk. A multivariable Cox regression model with backward selection was then applied to identify variables for inclusion in the prediction model. Hazard ratios (HRs) and 95% confidence intervals (CIs) were calculated. Proteins identified as being associated with gastric cancer risk in the two-stage study were further selected using LASSO-penalized Cox regression. To reduce the risk of overfitting given the limited number of gastric cancer events, stability variable selection among validated proteins was performed using LASSO-penalized Cox regression with 200 bootstrap iterations; each iteration randomly drew 80% of the sample without replacement. In each subsample, LASSO-Cox regression with 10-fold cross-validation was conducted, and variables selected at lambda.1se were recorded. The selection frequency for each variable was calculated across all iterations. Proteins with a selection frequency ≥ 0.7 were considered stable predictors and were included in the primary protein model. Proteins with a selection frequency ≥ 0.4 were selected to construct an exploratory extended protein model. Four prediction models were constructed in the validation cohort using the Cox proportional hazards model, including a clinical model (Model 1), a clinical model plus polygenic risk score (Model 2), a primary protein model incorporating clinical factors and stable proteomic biomarkers (selection frequency ≥ 0.7) (Model 3), and an exploratory extended protein model incorporating additional proteomic biomarkers (selection frequency ≥ 0.4) (Model 4).

Internal validation was performed using 1000 bootstrap resamples to estimate optimism-corrected performance, including the C-index and calibration slopes. Shrinkage factors were calculated to quantify the degree of potential overfitting. Model discrimination was assessed using optimism-corrected C-index and time-dependent ROC curves. C-index was compared using the ‘compareC’ function in R. Calibration plots were generated to assess the agreement between predicted and observed gastric cancer-free probabilities. A weighted risk score was calculated for all participants in the validation cohort based on the combination of clinical and proteomic signatures using the ‘predict’ function in R. The formula for the risk score is as follows: Risk Score = exp (β1⋅X1 + β2⋅X2+…+βn⋅Xn) where Xn represents the level of each variable, and βn​ is the corresponding coefficient derived from the Cox regression model. Cut-off values were determined using Youden index. In addition to hazard ratios, absolute cumulative incidence estimates at 5, 10, and 15 years were calculated for each risk group using the Kaplan-Meier method. The distribution of gastric cancer events across risk groups was also reported to complement the relative risk estimates.

Decision curve analysis was conducted to compare the net benefits of the different models.

All analyses were performed using R (version 4.5.3), unless otherwise specified.

Comments (0)

No login
gif