Biochip for Determination of 92 Human Y-Chromosome Haplogroups by SNP Markers and Frequency Distribution of These Haplogroups in Slavic Population of European Russia

Selection of the haplogroups that make the panel most informative for the purposes of investigation was a primary task in designing a tool for genetic personal identification. A test system for genotyping at several thousands of markers would certainly provide an ideal variant. However, current methods fail to combine a performance of the kind with a sensitivity and implementation potential required for forensic routine. Known next-generation sequencing techniques and high-density biochips are demanding in terms of the amount and quality of test DNA or are too complicated to use in an ordinary forensic laboratory. Moreover, instruments for high-throughput screening are not manufactured in Russia, and their purchases abroad are associated with certain sanction risks. The Russian biological microchip technology does not have the above drawbacks, but the number of test markers must be reduced to a reasonable optimum because of the limitations characteristic of multiplex PCR.

Haplogroup Selection

Almost all core haplogroups were included in the panel: B, C, D, E, G, H, I, J, L, N, O, Q, R, and T. The haplogroups M and S were exceptions because their occurrence is virtually zero outside the islands of Southeastern Asia and Oceania [13, 14]. The presence of the haplogroups M and S is possible to establish indirectly. In that very unlikely case, hybridization with only ancestral-type DNA probes will be observed for all of the 14 core haplogroups.

Then publicly available data were analyzed for subclades of each core haplogroup to consider whether a particular subclade is reasonable to include in the panel. Because only fragmentary occurrence data are available for many haplogroups and their subclades in the literature, public databases were used as an additional information source. These primarily included FamilyTreeDNA (https://discover.familytreedna. com/y-dna) and Haplogroup Research Analytical Suite (HRAS, https://hras.yseq.net). Haplogroup frequencies given below were from these databases, unless otherwise indicated.

The core haplogroups B, D, H, and T were not detailed to the subclade level because their occurrences are low in the target region. The haplogroup B is found predominantly in countries of Central and Southwestern Africa: the Central African Republic, Congo, Angola, Namibia, Botswana, etc. The haplogroup D (CTS3946) shows the highest occurrence in Japan (28%), China (5%), Mongolia (3%), Vietnam (3%), and Kazakhstan (2%), while its frequency in Russia is lower than 1%. The haplogroup H (M2920) is specific to the Romani (44%) [15] and is widespread in India (19%), Pakistan (8%), Uzbekistan (5%), and Azerbaijan (2%) and very rare in Russia (<<1%). The haplogroup is presumably common in the Romani population of Russia, but reliable data are unavailable. The haplogroup T (M272) is most characteristic of populations of the Middle East (its maximum frequency, 19%, is observed in the United Arab Emirates), North Africa (3‒16%), Armenia (5%), Azerbaijan (5%), Ukraine (2%), and Belarus (2%); its frequency in Russia is approximately 1%. As is seen from the above data, it is unreasonable to divide the four core haplogroups into subclades.

A list of subclades selected in each haplogroup is shown below. The logic of selection is illustrated with the example of C-M130.

The haplogroup C-M130 is divided into C1-F3393 and C2-M217. Because there are only two subclades, C2a-F1906 and C2b-Z4063, in C2-M217, the cladees were included in the panel in place of C2. The subclades C1-F3393 is extremely rare in the target region and can be established indirectly, by the presence of C-M130 with simultaneous lack of C2a-F1906 and C2b-Z4063. Variants of the kind are designated with an asterisk placed after the upstream haplogroup: C-M130* in this case. As for a further cladeing of C2a-F1906, C2a1‒C2a1a is its major lineage. Diverging clades are extremely rare in Latin America (C2a2) and Japan (C2a1b) and were consequently not included in the panel. The subclade C2a1a is divided into three large subclades (C2a1a1-Z18160, C2a1a2-F6370, and C2a1a3-M504), which are important in the Central Asian region, and minor C2a1a4-F9992, which is rare and low significant. C2a1a1-Z18160 and C2a1a2-F6370 were included in our panel. In place of C2a1a3-M504, we used its downstream major subclade C2a1a3a1-Y4630. Of the downstream subclades of the haplogroup C2b-Z4063, the subclade C2b1a1a-M407, which is more specific to Kazakhs, Yakuts, and Buryats, was included in the panel.

The haplogroup E1b-P177 and downstream E1b1b-M5017 are found in three regions: East European (8‒18%) and Baltic (2‒11%) countries, the Caucasus (8% in Azerbaijan, 7% in Armenia, and 4% in Georgia), and Central Asia (7% in Turkmenistan and Uzbekistan and 4% in Tajikistan). The frequency of the haplogroup in Russia is 4%. Ethnic specificity was generally not observed for downstream subclades of the haplogroup. However, several cladees that are major in the target region and improve the discriminating power of the method were included in the panel: E1b1b1a-L539, E1b1b1a1b1-L618, E1b1b1a1b2-L677, E1b1b1b-Z827, E1b1b1b2a1-M123, E1b1b1b2a1a1-L29, and E1b1b1b2a1a4-L791.

The haplogroup G-P257 shows the highest frequencies in Kuban and the Kaucasus, including Adygea (47%), North Ossetia (43%), Abkhazia (35%), Georgia (33%), Armenia (12%), and Azerbaijan (9%). Moderate frequencies are observed in Central Asia, including Kazakhstan (9%), Tajikistan (5%), and Uzbekistan (5%), and Eastern Europe, including Lithuania (7%), Moldova (8%), Belarus (6%), Ukraine (5%), and Latvia (4%). In Russia, the haplogroup is detected in 6% of the population. The side subgroup G1-F858 shows the highest frequency (5%) in Kazakhstan and low frequencies (≤ 1%) in the Caucasus, Russia, and East European countries. The major clade G2-L156 has two subclades, relatively rare G2b and major G2a-P15. The latter is divided into larger G2a2-L1259 and smaller G2a1-Z6552. Both of them were included in the panel.

The haplogroup I-U179 is a West European haplogroup and has two distribution foci, North European countries (27‒48%) with carriers of the subclade I1-M253 and Balkan countries (30‒47%) with carriers of the subclades I2-M438. I1-M253 is sufficiently common in Estonia (15%). Its two downstream subclades, I1a1-Z2336 and I1a2-Z58, were included in the panel. The haplogroup I2 is widespread in Ukraine (13%), Belarus (7%), and Russia (5%). Three down-stream subclades of the haplogroup were selected for the panel: I2a1a-S2648, I2a1b-M436, and I2a2-L596.

The haplogroup J-M304 shows the highest frequency in the Greater Middle East with a maximum on the Arabian Peninsula (50‒74%). The haplogroup is divided into two approximately equal subclades, J1-PF4641 with a focus shifted to the region of North Africa and the Arabian Peninsula and J2-CTS886 with a distribution shifted to the north. Of its downstream subclades, J1a2a1a˜-Z2363 is more characteristic of East Europe, Baltic countries, Belarus, and Ukraine (4‒9%) and J1a2b˜-CTS3569 is more common in Russia (5%). To represent the clade J2-CTS886, we selected J2a-M410; its subclades J2a1a1a-PF5125, J2a1a1a2b2-PF5130, J2a1a1a2b2a3˜-Z7671, and J2a1a1b-Z2221; and J2b-Z2417.

The haplogroup L-M185 has a relatively low occurrence. Its major distribution foci are in India (11%), Central Asia (2‒16%), Chechnya (13%), Turkey (5%), and Caucasian countries (2‒4%). L1 is its major downstream subclade, while L2 is an extremely rare haplogroup. L1a-M2481 is divided into L1a1 (not included because it is characteristic of Pakistan and India) and L1a2. In turn, the latter is divided into L1a2a-Z5919 (more common in Central Asia) and L1a2b˜-Y6288, which is highly specific to Chechnya (13%) and Ingushetia (10%). L1b-PH3615 is found in the Caucasian region and Central Asia at a low frequency.

The haplogroup N-M231 is widespread in North Eurasia, mostly in Finland (61%), Yakutia (83%), Karelia (63%), Baltic countries (18‒42%), and Mongolia (10%). Its frequency in Russia is approximately 15%. N-N1-N1a is its major line and is divided into N1a1-L395 and N1a2-CTS10075. Of the subclades downstream of N1a1-L395, N1a1a1a1a1-CTS10760 and N1a1a1a1a2-CTS10082 were selected for the panel.

The haplogroup O-M175 is a main haplogroup in East Asian populations and accounts for ~75% of the male population in China [16] and ~87% in Southeastern Asia [17]. O-M175 is divided into two cladees, O1-M1422 and O2-F36. O1-M1422 shows the highest frequencies in Southeastern Asia (from 23% in the Philippines to 60% in Indonesia), China (15%), South Korea (34%), Japan (24%), and Mongolia (3%). Its frequency in Russia and republics of the former Soviet Union is lower than 1%. The subclade O2-F36 is maximally represented in Myanmar (57%) and China (38%) and is found in the island part of Southeastern Asia (9–30%), Mongolia (13%), Kazakhstan (5%), Uzbekistan (3%), Kyrgyzstan (2%), and Tajikistan (1%). Its frequency in Russia is <1%. O-M134 is among the most common subclades of O-M175. The haplogroup O2-F36 has a number of downstream cladees of low significance, while its subclade O2a2b1-M134 is important for the target region and is the most abundant. The haplogroup occurs in East Asia and is often found in Kazakhs [18] with a higher prevalence in the Middle Zhuz and, in particular, Naimans (42%). The presence of the haplogroup O-O2 and the absence of O2a2b1 indicates that the specimen belongs to one of the subclades intermediate between O2 and O2a2b1 (O2b*, O2a1*, O2a2a*, and O2a2b2*).

The haplogroup Q-M1105 is most characteristic of native American populations (up to 70%) and is uncommon in Eurasia, being found mostly in Central Asia (4‒9%) and the Caucasus (up to 5%). Its frequency in Russia is 2%. The clades Q1-L472 and Q2-L275 and the Q1 subclades Q1b1a-L74 and Q1b1b-YP344 were selected. Q1a1-M7361 is a rare subclade and is characteristic of Turkmenistan (13%) and Chechnya (4%). The subclade was included in the panel because its marker is near the marker of Q1b1b-YP344 and does not require separate amplification; only a discriminating DNA probe should be included in the biochip.

The haplogroup R-P224 is the most widespread in Russia, and its 27 subclades were consequently included in the panel. First, three major cladees were isolated: the most abundant R1a-M420; West European R1b-S224; and Indian-Pakistani R2-M479, which is relatively common in Central Asia (2‒6%). Further cladeing is shown in Fig. 1. Most clades are distributed over vast areas and have a low investigational potential in the Slavic population, which constitutes an absolute majority of the total Russian population. However, regional specificity is observed for certain clades. For example, the subclade R1a1a1b2a2a3˜-S23592 is found in 34% of the Kyrgyz population; R1a1a1b2a2a1-Z2123 is specific to Karachay-Cherkessia (66%), Bashkiria (56%), and Kabardino-Balkaria (44%); and downstream R1a1a1b2a2a1c˜-Y20746 is characteristic of Bashkiria (33%), but not the other regions of Russia.

The haplogroup panel formed on the basis of the above data (Table 1) was included in a test system, which was named Y-expert. It should be noted that the haplogroup frequencies were obtained from public databases. The databases are based on commercial works, which are performed mostly for urban clients. This introduces an error into the estimates of actual population frequency distributions. We had to use these approximate data because there were no relevant publications concerning many of the populations of our interest. Using our test system among other tools, particular regions and ethnic groups are possible to study with more stringent criteria of population research in the future.

To facilitate the utility of data obtained with forensic specimens, we created a reference guide with information on all of the haplogroups included in the panel. The guide describes the general phylogeny of Y haplogroups and provides detailed information about each clade, including ethnogeographic distribution maps, haplogroup frequencies in countries that were parts of the former Soviet Union, a list of countries ranked by the frequencies of particular haplogroups, and links to public database pages with current information on haplogroups of interest (see S1, Supplementary).

Characteristics of Markers

Once the list of the target subclades was formed, the next task was to select the SNP markers associated with the subclades. Although a single marker was often known for a subclade, there were several synonymous markers in the Y chromosome for the majority of the subclades. In that case, we selected the marker that was the most promising in terms of genotyping, lacked nonspecific sequences in the flanking regions, had a GC content of approximately 50%, and allowed an amplicon of a minimal possible size. The markers selected by these criteria are summarized in Table 1. It should be noted that a selected marker was not always the synonym most commonly used for the respective subclade according to the ISOGG Y-DNA Haplogroup Tree 2019-2020, Version: 15.73 (https://isogg.org/tree).

Given that degraded DNA testing (forensics and paleogenetics) was one of the intended applications of our test system, we tried to minimize the amplicon size. The following values were achieved: amplicon size range, 50–120 bp; mean amplicon size, 72 bp; and median amplicon size, 68 bp. The values are possible to improve in further studies.

Testing of Russian Population

Various regional samples from the Russian populations have many times been examined to detect various, mostly major, subclades of Y haplogroups. We were the first to study the Y haplogroups in Slavs from central Russia at the given level of subclade particularization by SNP markers. Genotyping results obtained in the population under study are summarized in Table 2.

Table 2. Haplogroup frequencies in the sample of Slavs from central Russia

As is seen from Table 2, derivatives of the R lineage (58.15%) and especially its subclade R1a (51.12%) were the most common haplogroups. The haplogroups I (17.98%) and N (9.55%) were the second and third most common, respectively, and were followed, in order of decreasing frequency, by the haplogroups E (4.78%), J (3.93%), G (2.53%), Q (1.69%), O and T (0.56% each), and C (0.28%). The haplogroups B, D, H, and L were not detected in the sample. The frequency distribution obtained for the main haplogroups generally agreed with data reported by Balanovsky et al. [19] in 2008, but the spectrum tended to shift towards central and southern regions of Russia. The trend was probably associated with northward migration processes.

Although R is the most common haplogroup in Russian Slavs, some of its subclades are rare in the population. The subclades that show zero frequencies, but are common in other ethnic groups attract particular interest. The property suggests their potential value as investigational markers. A set of such subclades includes Indian-Pakistani R-L657.1, which is found in Central Asia (1‒6%) and Kazakhstan (3%); Bashkir R-Y20746 (31%); R-Y934, which is widespread in Karachay-Cherkessia (58%), Kabardino-Balkaria (45%), Adygea (22%), and Bashkiria (22%); R-S23592, which is the most common in Kyrgyzstan (34%); and R-M73 and R2, which are common in Central Asia.

The other haplogroups and their subclades that were not detected in the study population are also of interest in this respect. The set includes the haplogroups B, D, H, and L, which are ethnospecific to Africans, East Asians, Romani, and ethnic groups of the Middle East and South Asia, respectively. The subclade C2a-F1906 is characteristic of the populations of Kalmykia (92%), Kazakhstan (27%), Kyrgyzstan (20%), Uzbekistan (19%), and Turkmenistan (6%). The subclade E-L677 occurs in the Middle East at a moderate frequency. The subclade G-M285 is most characteristic of Kazakhstan (6%).

It should be noted that two specimens from the Russian sample were assigned to the haplogroup O, which is not characteristic of Caucasians. One of the carriers had a patronymic suggesting a Chinese origin. The family name, first name, and patronymic of the other carrier were typical for a Slav. A carrier of the Vainakh subclade J2a1a1a2b2a3~ (Z7671), which is also unusual in Slavs, was additionally found in the sample. Because all donors self-identified themselves as Slavs in our samples, the findings might suggest child adoption, rape, adultery, incorrect self-identification, or complete assimilation for the genealogical history of the carriers.

Finally, consider the place that our method might occupy among the available means to directly establish the Y haplogroups by SNP markers. On the one hand, real-time PCR provides the most available genotyping method. Simple as it is, the method has a drawback of having a low multiplex capacity and, therefore, low resolution. On the other hand, methods of reading vast genetic information are currently available, the examples including massively parallel sequencing and high-density microarrays (Illumina, United States). These are excellent research tools and yield comprehensive information. However, the methods cannot be introduced in Russian forensics now because their protocols and bioinformatics data processing techniques are rather complex and stringent requirements are imposed on the amount and quality of DNA and skills of a researcher. We think that the place of our low-density biochip is between these two extreme approaches. Its operating protocol is maximally simplified and allows its broad implementation in routine testing, and its multiplex capacity is sufficient for solving routine forensic problems.

Because the technology has its limitations (up to 100 polymorphisms per biochip approximately), our test system almost reached the limit of its information capacity in the given format. We think that a two-step approach can be employed to method and to extend the range of haplogroups involved in genotyping. The first step is establishing the core haplogroup (e.g., with the Phenotype Expert test system [8]), and the subhaplogroup is addressed at the second step. Individual biochips should be designed for each core haplogroup in this case. This will increase the resolution of the method to 100 subclades per each core haplogroup and will fully meet the actual demand in the field. If the origin needs to be established to a deeper level (a descent group), genotyping at Y-chromosomal short tandem repeat (STR) loci is reasonable to include in the study. In that case, conserved Y-SNPs will ensure the data reliability by virtue of lacking homoplasy, while Y-STRs will provide for a more detailed analysis by virtue of their higher variability.

The panel of markers designed in our work makes it possible to assess the biogeographical patrilinear origin of a person. The panel is relevant for getting investigational information in the Russian Federation in view of migrations from former Soviet republics, Southeastern Asia, and the Middle East. The test system created on the basis of the panel includes a set of multiplex PCR primers and a biochip with specific DNA probes. A maximally simple genotyping protocol was developed to facilitate broad implementation of the method. Using the method, new data were obtained to characterize the haplogroup distribution in the Slavic population, most of which lives in central Russia. Our test system makes it possible to characterize the available ethnic and regional collections of DNA specimens within a short period of time and at minimum costs. The results will help to clarify the ethnic and biogeographical specificity of individual subclades and to establish it in populations not examined in this respect earlier.

Comments (0)

No login
gif