Introduction:
Lipidomics and metabolomics have become key approaches for understanding and diagnosing human diseases, including type 2 diabetes, Alzheimer’s disease, cancer, and kidney dysfunction. This study provides a comprehensive overview of the evolution of these disciplines through a bibliometric and text-mining analysis of scientific production from 2004 to 2024, based on data from Scopus and validated through a multi-database comparative approach.
Methods:
A total of 9,628 articles were harmonized and analyzed using Bibliometrix, Scimago Graphica, OpenRefine, and custom R scripts to identify the most productive journals, authors, countries, and institutions, and to map thematic structures and keyword dynamics. Beyond traditional bibliometric indicators, our integrative approach combined quantitative trends with semantic and conceptual mapping to trace the methodological and translational evolution of the field. To ensure robustness and generalizability, equivalent searches were conducted in the Web of Science Core Collection and PubMed, and cross-database validation assessed concordance in journal and country rankings, as well as temporal and thematic trends.
Results:
The field shows rapid expansion, with an annual growth rate of 32.6%. The United States and China lead global output, followed by major European contributors. Core topics include Alzheimer’s disease, obesity, and breast cancer, while emerging areas focus on artificial intelligence, multi-omics integration, and Mendelian randomization. Analytical methodologies such as liquid chromatography-mass spectrometry, gas chromatography-mass spectrometry, and nuclear magnetic resonance, together with metabolic diseases, remain central to the field. In contrast, niche themes such as microbiota-COVID-19 interactions and oxidative stress-cancer associations represent emerging interdisciplinary bridges.
Discussion:
Overall, lipidomics and metabolomics are evolving toward integrative and computational frameworks with strong diagnostic potential, underscoring the need for validated biomarkers, standardized data pipelines, and open repositories to enable clinical translation.
1 IntroductionMetabolomics systematically studies metabolites, the chemical compounds produced, used, or transformed during cellular metabolism. Metabolites, such as amino acids, fatty acids, and carbohydrates, are the downstream end products of biological processes and can function as biomarkers for clinical diagnosis (Patti et al., 2012). In addition, metabolomics can elucidate metabolite dynamics, composition, interactions, and changes in response to external interventions such as drugs, diet, or environmental conditions (Pang and Hu, 2023). These shifts can directly impact within cells, tissues and biofluids, underscoring the role of metabolomics in linking human physiology, lipid metabolism, and nutrition (Johnson et al., 2016). On the other hand, lipids are a diverse group of metabolites that serve as structural components of cell membranes, sources of energy storage, and intermediates in signaling pathways (Han and Gross, 2022).
Like genes and proteins, metabolites can also serve as signatures of biochemical activity, facilitating their correlation with disease phenotypes. Remarkably, a single mutation in a gene can affect several metabolic pathways, thereby exerting a functional influence on distinct cellular processes (Gieger et al., 2008). Nowadays, multi-omics studies scrutinize the human lipidome and/or metabolome to facilitate the development of personalized treatments, for example, by monitoring changes in metabolites and lipid levels relative to genetics or epigenetics (Astarita et al., 2023). Beyond discerning the coordination of each gene with its molecular function to promote human health, contemporary research aims to elucidate the relationship between nutrients, host metabolism, and gut microbiota, to achieve a more comprehensive understanding of biological processes and their influence on phenotypes and diseases (Orešič, 2009).
Although the terms lipidomics and metabolomics gained prominence in the early 2000s, the proposition that the quantitative analysis of metabolites and chemical constituents could be beneficial to monitor a patient’s health can be traced back several decades (Gross, 2017; Hasin et al., 2017). The advent of mass spectrometry in the early 20th century, coupled with the development of separation techniques such as gas chromatography (GC) and liquid chromatography (LC), enabled researchers such as Nobel laureate Linus Pauling and his colleague Arthur B. Robinson in 1971 to profile breath and urine vapor with the following premise: “The thorough quantitative analysis of body fluids might permit differential diagnosis of many diseases in a more effective way than is possible at the present time” (Pauling et al., 1971).
Lipidomics and metabolomics have become increasingly relevant for disease diagnosis. Specific metabolites, including isoleucine, leucine, valine, tyrosine, and phenylalanine, are associated with insulin resistance and β-cell function before the onset of type 2 diabetes (Wang et al., 2011). The role of lipids in disease ranges from dysregulation of the plasma lipidome in Alzheimer’s disease (Liu et al., 2021) to new research showing promising results with mass spectrometry (MS) for faster diagnosis of prostate cancer (Buszewska-Forajta et al., 2021) or improving the diagnosis and prevention of kidney function (Baek et al., 2022). Moreover, researchers have investigated lipidomic and metabolomic signatures in COVID-19 patients, showing that, despite clinical heterogeneity, patients consistently display distinct profiles when compared with healthy controls. In addition, analyses before and after tocilizumab treatment revealed a partial reversion of these alterations, suggesting that lipidomic and metabolomic profiling may help monitor therapeutic effects (Meoni et al., 2021).
Additionally, bibliometric studies are essential for monitoring the growth of research in a field. These studies enable the analysis and prediction of trends, including the estimation of the extent, timing, and participants in specific research areas (Kokol et al., 2021). These studies offer a comprehensive overview and facilitate collaboration among current researchers by highlighting analogous projects. There are two types of bibliometric analysis: performance analysis and science mapping. The former examines the contributions of authors, countries, institutions, and journals. The science mapping examines the relationships between these constituents to determine the social, intellectual, and conceptual structures of the topic of interest (Donthu et al., 2021).
The utilization of lipidomics and metabolomics in the diagnostic process has garnered significant attention, as evidenced by multiple bibliometric studies addressing specific disease contexts (see below). However, lipidomics and metabolomics encompass a broad spectrum of medical conditions, and a preliminary analysis of existing literature revealed that no study had comprehensively examined their conceptual and translational evolution across diseases.
Recently, Li et al. (2025) published a data mining-based bibliometric study that mapped the research landscape and emerging frontiers of lipidomics from 2003 to 2024, highlighting publication growth, international collaborations, and emerging themes such as ferroptosis, gut microbiota interactions, and multi-omics integration (Li et al., 2025). Their analysis, while providing a valuable goal overview of lipidomics, did not address the diagnostic dimension or the cross-omic convergence with metabolomics. In contrast, the present study jointly examines lipidomics and metabolomics to capture their conceptual convergence, research interconnections, and diagnostic translation.
The following bibliometric studies, identified through a keyword ((Metabolomics OR Lipidomics) AND Bibliometric*) search, are notable examples in this field: metabolomics in coronary artery disease (Yu et al., 2022), research on metabolic dysfunction in polycystic ovary syndrome (Xu et al., 2023), cancer metabolic reprogramming (Yang et al., 2024), the metabolomics of osteosarcoma (Tu et al., 2024), and the relationship between exercise and metabolomics (Lv et al., 2022).
The objective of this study is to explore and characterize the global scientific production in lipidomics and metabolomics applied to disease diagnosis from the early 2000s to 2024 – a period that has seen remarkable expansion in omics-based research. We seek to map the most relevant studies, diseases addressed, and analytical strategies, as well as the collaborative and thematic evolution of these fields. Through bibliometric performance and science mapping analyses, we identify key contributors, conceptual structure, and emerging research trends shaping diagnostic innovation in lipidomics and metabolomics.
What are the current publication and collaboration patterns, including main sources, contributors, and author networks?
How have research topics evolved over time according to publication growth, thematic evolution, and keyword dynamics?
Which authors, institutions, and countries have been most productive or influential?
Which diseases, analytical platforms, and concepts dominate current research?
Which emerging themes and conceptual connections signal a transition toward clinical translation and precision medicine?
Ultimately, this study aims to provide a comprehensive overview of how lipidomics and metabolomics have evolved as diagnostic and translational sciences, highlighting disease contexts, analytical methods, and research collaborations that are driving the integration of these omics into precision medicine.
2 Materials and methods2.1 Data source and search strategyScopus was used as the primary data source due to its broad coverage of high-quality scientific literature in health and biomedical sciences. A comprehensive literature search was conducted on February 8, 2025, to identify studies addressing the role of lipidomics and metabolomics in human disease-related contexts.
The search strategy was designed to retrieve original research articles focused on human biomedical studies while excluding publications based on animal or plant models and those belonging to non-biomedical disciplines. Only articles written in English and published between 2004 and 2024 were considered. As data collection occurred in early 2025, publications from later in that year may not have been fully indexed at the time of retrieval.
The search yielded a total of 9,728 records, which constituted the primary dataset for subsequent bibliometric analyses. A detailed description of the search query, inclusion and exclusion criteria, data handling procedures, and data analysis considerations are provided in the Supplementary File 1.
Figure 1 illustrates the systematic identification, screening, and inclusion of records retrieved from Scopus, followed by data cleaning, harmonization, and visualization workflows, following PRISMA principles adapted from bibliometric studies (Ahmed et al., 2023; Murthy et al., 2025). The raw and processed datasets derived from Scopus are provided in Supplementary Files 2, 3.

PRISMA and bibliometric analysis workflow steps.
2.2 Data filteringData filtering was performed using the Biblioshiny interface of the Bibliometrix package to retain only records with complete bibliographic metadata (Aria and Cuccurullo, 2017; Derviş, 2020). From the initial set of 9,728 retrieved articles, a total of 9,628 records met this criterion and were retained for subsequent analyses.
The final curated dataset used in the analyses is provided in Supplementary File 4, together with a metadata completeness report generated by Biblioshiny (Supplementary File 5).
2.3 Data harmonizationBibliometric analyses require a preliminary data harmonization step to address inconsistencies and variations in bibliographic metadata, which can affect the reliability of thematic analyses. In this study, harmonization focused on author’s keywords, as these terms best capture the conceptual content of each publication.
To unify semantically equivalent terms expressed in different lexical forms, we applied a combined automated and manual harmonization strategy using the open-source tool OpenRefine (Ham, 2013; CS&S, 2024). Automated clustering methods were first used to group closely related keyword variants, followed by manual inspection to prevent overgeneralization and to ensure conceptual accuracy. It is noteworthy to mention that manual inspection was limited to validating automated clustering outcomes and correcting a small number of clearly overgeneralized cases, accounting for less than 1% of the total keyword terms.
This process allowed for the consolidation of synonymous disease names, analytical platforms, and methodological terms into standardized representations, resulting in a harmonized keyword dataset suitable for robust downstream analyses. Detailed descriptions of the harmonization procedures and outputs are provided in Supplementary Files 6-8.
2.4 Data analysisData analysis was conducted using a combination of Bibliometrix/Biblioshiny, OpenRefine, and custom R scripts to perform performance analysis and scientific mapping. Custom scripts were developed to address known limitations of automated bibliometric tools, particularly for author disambiguation and country level productivity metrics.
To avoid inflation of national research output caused by multi-author publications, scientific production was quantified at the document level rather than by author counts. Scimago Graphica was used for graphical representation of selected indicators (Hassan-Montero et al., 2022). Document-level analyses were performed to address research questions related to disease representation and thematic focus. Author keywords were analyzed to identify the most frequently studied diseases, which were grouped into broader conceptual categories using a curated disease dictionary. Temporal analyses were conducted to examine shifts in research emphasis over time, including the emergence of pandemic-related research after 2020.
Journal performance was assessed using Scopus-derived indicators, including CiteScore (CS) (Teixeira da Silva and Memon, 2017), SCImago Journal Rank (SJR) (Guz and Rushchitsky, 2009), and Source Normalized Impact per Paper (SNIP) (Moed, 2010), to provide a multidimensional view of source influence across study areas.
The PICOS framework was not applied, as this study focuses on bibliometric and text-mining analyses rather than on structured clinical outcomes typical of systematic reviews.
2.5 Multi-database validationTo assess the robustness and comprehensiveness of the results obtained from the Scopus dataset, additional searches were conducted in the Web of Science Core Collection (WoSCC) and PubMed. Equivalent search strategies, based on the same keywords, Boolean operators, and a time frame restricted to 2004-2024, were adapted to the syntax of each database. Inclusion and exclusion criteria were applied consistently across all sources to ensure comparability.
For each supplementary database, the same general analytical framework was applied, including verification of metadata completeness. Keyword harmonization was not performed, as differences in indexing terminology across databases were expected. Analyses focused on total record counts, annual publication trends, and country-level productivity based on publication and corresponding author information.
Cross-database comparisons emphasized conceptual consistency rather than exact numerical agreement. Concordance in temporal trends was evaluated using linear regression analyses of annual publication counts. This approach allowed us to determine whether thematic focus, geographical distribution, and temporal patterns identified in the Scopus-based analysis were reproducible across other major bibliographic sources.
This validation framework follows recent bibliometric studies (Bai et al., 2025) and aligns with methodological recommendations for cross-database comparison and data cleaning (Lim et al., 2024).
3 Results3.1 Overview3.1.1 Main informationOur preliminary analysis of the 9,628 documents reveals a wealth of insightful information regarding the evolution of lipidomics and metabolomics in disease diagnosis. Table 1 summarizes key descriptive statistics of the dataset and highlights fundamental trends in the literature. We observed a notable annual growth rate of 32.58%, which aligns with the accelerated expansion of research in lipidomics and metabolomics for disease diagnosis and personalized healthcare. This metric demonstrates the growing recognition of the clinical and research relevance of lipidomics and metabolomics.
FeatureExplanationCountMain information about dataDocumentsTotal number of scientific publications9628SourcesThe frequency distribution of sources such as journals and books1891Annual growth rate %The average increase in the number of documents over a year32.58Document average ageAverage age of a document given in years5.47Average citations per docThe average number of quotes in each article29.55Document contentsKeywords plus (ID)Total number of words or phrases that frequently appear in the title of an article’s references37177Author’s keywords (DE)Total number of keywords14302AuthorsAuthorsTotal number of authors65979Authors of single-authored docsThe number of single authors per article83Authors collaborationSingle-authored docsTotal number of single-authored documents85Co-authors per docThe average number of co-authors in each document9.91International Co-authorships %The average number of articles with international collaboration31.38Summary of descriptive information on the dataset collection found from 2004 to 2024.
We also observed that the documents have an average age of 5.47 years over the period (2004-2024), indicating a recent surge in publications. This trend underscores sustained research activity and highlights the increasing momentum of the field. The identification of 14,302 author keywords underlines the comprehensive thematic scope, encompassing the application of lipidomics and metabolomics in diagnostics. These keywords provide the foundation for further analysis and encapsulate the core concepts and thematic directions emphasized in current research.
Our analysis of collaboration trends offers valuable insights. The data show that three out of every ten articles (31.38%) resulted from international collaborations with an average of 9.91 co-authors per publication. This high degree of co-authorship reflects the multidisciplinary nature of lipidomics and metabolomics. These fields draw on expertise from medical and biological sciences, technological developments (e.g., LC and GC -MS), systems biology, and bioinformatics. By integrating these disciplines, researchers have fostered global cooperation and enabled significant advancements in high-throughput analytical techniques and computational modeling, both of which are essential for deciphering complex biological systems.
3.1.2 Annual scientific productionFigure 2 illustrates the progression of annual scientific output in the domains of lipidomics, and metabolomics as applied to disease diagnosis. The data reveals a continuous growth since 2004, with a more pronounced upward trajectory commencing in 2010 that signaled the consolidation of this field of study. A notable milestone was attained in 2011 when the number of publications surpassed 100 for the first time. A decade later, in 2021, publications exceeded 1,000, further solidifying the field’s expanding relevance. In 2024, output reached its highest level to date, indicating sustained interest within the scientific community.

Annual distribution of scientific publications and polynomial forecast for 2025 in lipidomics and metabolomics for disease diagnosis.
To analyze the evolution of scientific production and project future trends, we fitted a second-degree polynomial regression model (Figure 2). This model captures the variability in acceleration and deceleration of growth over time, providing a more accurate representation of the annual growth rate compared to an exponential model. Our fitted model achieved an adjusted R2 of 0.99, indicating excellent agreement with the observed data. Because this feature measures the relative predictive power of the polynomial regression model, a value close to 1 indicates an excellent fit to the observed data. The polynomial model also indicated a growth trend but shows a slowdown in recent years (2021-2023).
Projections based on the model estimated that researchers working on lipidomics and metabolomics would publish approximately 1,535 scientific papers in 2025. However, updated Scopus data reviewed up to mid-April 2026 identified 1,788 publications for the year 2025. This discrepancy suggests that the recent growth of the field may be accelerating beyond the trend captured by the model, reinforcing the overall conclusion of sustained and expanding research activity in lipidomics and metabolomics for diagnosis.
We provide the annual production data in Supplementary File 9.
3.2 SourcesThe document’s source refers to the title of a journal, book, or conference proceeding. In this bibliometric analysis, we evaluated the impact and productivity of sources in lipidomics, and metabolomics applied to disease diagnosis following Bradford’s Law (Brookes, 1969). This identifies the most relevant sources within a research domain. We found that the core zone of the most productive and influential periodicals includes 28 sources (1.48%), which collectively published 3,205 documents (33.29%). Table 2 presents 20 of these 28 core sources, ranked by the number of publications (TP), along with their h-index (h), g-index (g), and m-index (m). These metrics provide a multi-dimensional view of each journal’s influence:
SourceTPTChgmYFPCSSJRSNIPHPQScientific Reports4411179554804.15420137.50.9001.18292ndQ1Metabolites329416231452.38520135.70.9030.84360thQ2Metabolomics283829345782.14320056.60.7400.76370thQ2Journal of Proteome Research26613197641033.20020069.01.2990.89684thQ1PLoS ONE25313132641063.36820076.20.8391.08489thQ1International Journal of Molecular Sciences209214525341.92320138.11.1791.12090thQ1Analytical Chemistry150933953942.524200512.11.6211.27491stQ1Frontiers in Immunology98129619312.37520189.81.8681.19377thQ1Journal of Pharmaceutical and Biomedical Analysis88191025391.47120096.70.5840.92776thQ1Clinica Chimica Acta86183726401.529200910.11.0161.08686thQ1Journal of Chromatography B83295029521.52620075.60.5390.83565thQ2Analytica Chimica Acta83235928461.556200810.40.9981.06791stQ1Analytical and Bioanalytical Chemistry81267732491.68420078.00.6860.85983rdQ1Journal of Lipid Research59327229571.813201011.12.0901.39489thQ1Frontiers in Endocrinology5849512191.33320175.71.2401.12260thQ2eBiomedicine54155924382.400201617.73.1931.72897thQ1Journal of Clinical Endocrinology and Metabolism53172523401.533201111.41.8991.57490thQ1Journal of Translational Medicine53100620301.667201410.01.6111.23494thQ1Cancers5176516252.28620198.01.3911.03079thQ1Journal of Chromatography A50182425421.47120097.90.7170.92385thQ1List of journals with the highest number of publications on the subject.
TP, Total Publications; TC, Total Citations; h, h-index; g, g-index; m, m-index; YFP, Year of First Publication; CS, CiteScore 2023; SJR, SJR 2023; SNIP, SNIP 2023; HP, Highest Percentile; Q, Quartile.
H-index (h): counts the number of papers (h) that have received at least h citations, capturing both productivity and impact (Hirsch, 2005).
G-index (g): refines the h-index by requiring that the top g articles receive at least g2 citations, which always makes g greater than or equal to h (Egghe, 2006).
M-index (m): divides h by the number of years since the first publication, helping evaluate emerging researchers or sources (von Bohlen und Halbach, 2011).
Our analysis highlights that the most productive journals in this area include S
Comments (0)