Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Massive sequence perturbation of the Raf ras binding domain reveals relationships between sequence conservation, secondary structure propensity, hydrophobic core organization and stability.

The contributions of specific residues to the delicate balance between function, stability and folding rates could be determined, in part by [corrected] comparing the sequences of structures having identical folds, but insignificant sequence homology. Recently, we have devised an experimental strategy to thoroughly explore residue substitutions consistent with a specific class of structure. Using this approach, the amino acids tolerated at virtually all residues of the c-Raf/Raf1 ras binding domain (Raf RBD), an exemplar of the common beta-grasp ubiquitin-like topology, were obtained and used to define the sequence determinants of this fold. Herein, we present analyses suggesting that more subtle sequence selection pressure, including propensity for secondary structure, the hydrophobic core organization and charge distribution are imposed on the Raf RBD sequence. Secondly, using the Gibbs free energies (DeltaG(F-U)) obtained for 51 mutants of Raf RBD, we demonstrate a strong correlation between amino acid conservation and the destabilization induced by truncating mutants. In addition, four mutants are shown to significantly stabilize Raf RBD native structure. Two of these mutations, including the well-studied R89L, are known to severely compromise binding affinity for ras. Another stabilized mutant consisted of a deletion of amino acid residues E104-K106. This deletion naturally occurs in the homologues a-Raf and b-Raf and could indicate functional divergence. Finally, the combination of mutations affecting five of 78 residues of Raf RBD results in stabilization of the structure by approximately 12 kJ mol(-1) (DeltaG(F-U) is -22 and -34 kJ mol(-1) for wt and mutant, respectively). The sequence perturbation approach combined with sequence/structure analysis of the ubiquitin-like fold provide a basis for the identification of sequence-specific requirements for function, stability and folding rate of the Raf RBD and structural analogues, highlighting the utility of conservation profiles as predictive tools of structural organization.

Amino Acid Sequence↗

Sequence conservation in families whose members have little or no sequence similarity: the four-helical cytokines and cytochromes.

Proteins for which there are good structural, functional and genetic similarities that imply a common evolutionary origin, can have sequences whose similarities are low or undetectable by conventional sequence comparison procedures. Do these proteins have sequence conservation beyond the simple conservation of hydrophobic and hydrophilic character at specific sites and if they do what is its nature? To answer these questions we have analysed the structures and sequences of two superfamilies: the four-helical cytokines and cytochromes c'-b(562). Members of these superfamilies have sequence similarities that are either very low or not detectable. The cytokine superfamily has within it a long chain family and a short chain family. The sequences of known representative structures of the two families were aligned using structural information. From these alignments we identified the regions that conserve the same main-chain conformation: the common core (CC). For members of the same family, the CC comprises some 50% of the individual structures; for the combination of both families it is 30%. We added homologous sequences to the structural alignment. Analysis of the residues occurring at sites within the CCs showed that 30% have little or no conservation, whereas about 40% conserve the polar/neutral or hydrophobic/neutral character of their residues. The remaining 30% conserve hydrophobic residues with strong or medium limitations on their volume variations. Almost all of these residues are found at sites that form the "buried spine" of each helix (at sites i, i+3, i+7, i+10, etc., or i, i+4, i+7, i+11, etc.) and they pack together at the centre of each structure to give a pattern of residue-residue contacts that is almost absolutely conserved. These CC conserved hydrophobic residues form only 10-15% of all the residues in the individual structures.A similar analysis of the cytochromes c'-b(562), which bind haem and have a very different function to that of the cytokines, gave very similar results. Again some 30% of the CC residues have hydrophobic residues with strong or medium conservation. Most of these form the buried spine of each helix and play the same role as those in the cytokines. The others, and some spine residues bind the haem co-factor.

Automation↗

AT-rich sequences flanking the 5'-end breakpoint of the 4977-bp deletion of human mitochondrial DNA are located between two bent-inducing DNA sequences that assume distorted structure in organello.

The 4977-bp deletion is the most common deletion among more than 90 large-scale deletions of human mitochondrial DNA (mtDNA) that are associated with aging and mitochondrial myopathies. The reason why the frequency of occurrence of this common deletion is so high in aged and myopathic human tissues is not clear. Since several studies proved that unusual DNA structures play very important roles in a number of recombination events, we hypothesized that some kind of unusual DNA structure may flank the breakpoints of the 4977-bp mtDNA deletion. We used two-dimensional (2-D) gel electrophoresis to assess the mobility abnormalities of the PCR-amplified DNA fragments encompassing the sequences of nucleotide position (np) 7901 to 9058 of human mtDNA. The results showed that the sequences of np 7901-8732 and np 8251-9058 exhibited retarded and increased mobilities, respectively, and that the sequence of np 8285-8676 showed normal mobility in the 2-D gel. This indicates that the 5'-end breakpoint of the 4977-bp deletion is located within the junction site of two flanking bent-inducing DNA sequences. We confirmed this notion by using osmium tetroxide (OsO4) to probe mtDNA in organello. The results showed that the two AT-rich sequences flanking the 5'-end breakpoint of the 4977-bp deletion are susceptible to OsO4 modification. These findings suggest that the DNA sequences of the 5'-end breakpoint of the common mtDNA deletion are rendered to assume a more distorted structure than B-DNA by these two flanking bent-inducing DNA sequences in organello and thereby render this region to be more vulnerable to attack by reactive oxygen species and free radicals.

Aging↗

Complete amino acid sequence of kaouthiagin, a novel cobra venom metalloproteinase with two disintegrin-like sequences.

The primary structure of kaouthiagin, a metalloproteinase from the venom of the cobra snake Naja kaouthia which specifically cleaves human von Willebrand factor (VWF), was determined by amino acid sequencing. Kaouthiagin is composed of 401 amino acid residues and one Asn-linked sugar chain. The sequence is highly similar to those of high-molecular mass snake venom metalloproteinases from viperid and crotalid venoms comprised of metalloproteinase, disintegrin-like, and Cys-rich domains. The metalloproteinase domain had a zinc-binding motif (HEXXHXXGXXH), which is highly conserved in the metzincin family. Kaouthiagin had an HDCD sequence in the disintegrin-like domain and uniquely had an RGD sequence in the Cys-rich domain. Metalloproteinase-inactivated kaouthiagin had no effect on VWF-induced platelet aggregation but still had an inhibitory effect on the collagen-induced platelet aggregation with an IC(50) of 0.2 microM, suggesting the presence of disintegrin-like activity in kaouthiagin. To examine the effects of these HDCD and RGD sequences, we prepared synthetic peptides cyclized by an S-S linkage. Both the synthetic cyclized peptides from the disintegrin-like domain and from the Cys-rich domain) had an inhibitory effect on collagen-induced platelet aggregation with IC(50) values of approximately 90 and approximately 4.5 microM, respectively. The linear peptide (RAAKHDCDLPELC) and the cyclized peptide had little effect on collagen-induced platelet aggregation. These results suggest that kaouthiagin not only inhibits VWF-induced platelet aggregation by cleaving VWF but also disturbs the agonist-induced platelet aggregation by both the disintegrin-like domain and the RGD sequence in the Cys-rich domain. Furthermore, our results imply that the corresponding part of the Cys-rich domain in other snake venom metalloproteinases also has a synergistic disturbing effect on platelet aggregation, serving as a second disintegrin-like domain. This is the first report of an elapid venom metalloproteinase with two disintegrin-like sequences.

Amino Acid Motifs↗

L1 family of repetitive DNA sequences in primates may be derived from a sequence encoding a reverse transcriptase-related protein.

Primate and rodent genomes contain a family of highly repetitive, long interspersed sequences, designated the L1 family or LINE-1. Characteristic features of the L1 family sequences such as an A-rich stretch at the 3' end, a truncated 5' end, the existence of significantly long open reading frames (ORFs) and the presence of L1 family transcripts in various types of cells, including pluripotential embryonic cells, suggest that the L1 family is derived from a sequence encoding a protein(s) and dispersed in the genome through an RNA-mediated process. These features of the L1 family are believed to be due to reverse transcription beginning at the 3' end of the L1 transcript and terminating prematurely and to the site duplication caused by the insertion of the complementary DNA. It is likely that this type of transcript is converted to cDNA and inserted into the chromosome through a process similar to that of the formation of processed pseudogenes. The above model, however, does not necessarily explain why the L1 family should produce the extraordinarily large number of copies (more than 10(4) per haploid genome) seen during evolution. It seems likely that the progenitor of the L1 family itself carries (or carried) a function which promotes the active dispersion of the L1 family sequence. We reasoned that such a function, if present, must be conserved during evolution and may be shown by comparative analysis of L1 family sequences from evolutionarily distant species. We show here that the L1 family sequence contains an ORF possessing significant sequence homology to several RNA-dependent DNA polymerases of viral and transposable element origins.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Map positions of 47 Arabidopsis sequences with sequence similarity to disease resistance genes.

Map positions have been determined for 42 non-redundant Arabidopsis expressed sequence tags (ESTs) showing similarity to disease resistance genes (R-ESTs), and for three Pto-like sequences that were amplified with degenerate primers. Employing a PCR-based strategy, yeast artificial chromosome (YAC) clones containing the EST sequences were identified. Since many YACs have been mapped, the locations of the R-ESTs could be inferred from the map positions of the YACs. R-EST clones that exhibited ambiguous map positions were mapped as either cleavable amplifiable polymorphic sequence (CAPS) or restriction fragment length polymorphism (RFLP) markers using F8 (Ler x Col-0) recombinant inbred (RI) lines. In all cases but two, the R-ESTs and Pto-like sequences mapped to single, unique locations. One R-EST and one Pto-like sequence each mapped to two locations. Thus, a total of 47 loci were identified in this study. Several R-ESTs occur in clusters suggesting that they may have arisen via gene duplication events. Interestingly, several R-ESTs map to regions containing genetically defined disease resistance genes. Thus, this collection of mapped R-ESTs may expedite the isolation of disease resistance genes. As the cDNA sequencing projects have identified an estimated 63% of Arabidopsis genes, a very large number of R-ESTs (approximately 95), and by inference disease resistance genes of the leucine-rich repeat-class probably occur in the Arabidopsis genome.

Arabidopsis↗

Improving quality of expressed sequence tag (EST) databases: recovery of reversed, antisense cDNA sequences.

Expressed sequence tag (EST) databases contain a significant number (5-20%) of reversed, antisense, cDNA sequences that can be recognized by the label "reversed clone: similarity on wrong strand" in the annotations to the sequence. Despite this high number of altered sequences, no attempt has been made to explain the alteration in molecular terms, or to evaluate their effect on the quality of the information curated in EST databases. In this paper we try to explain the way these altered sequences are originated, and propose a plausible mechanism: a "double priming" of the first strand oligo-dT primer at both ends of nascent cDNAs. In this way, a symmetrical cDNA intermediate is generated, an intermediate that can be cloned after partial digestion with the restriction enzyme used for the directional cloning. Furthermore, when "secondary" priming takes place inside the cDNA, the chain synthesized is prone to be truncated prematurely, with the subsequent loss of upstream information. One of the most subtle effects of this cloning alteration is the generation of virtual open reading frames (ORFs) in sequences with no homologues available for comparison. Nevertheless, and according to our model and our data, the "double priming mechanism" does not shift the ORF effected, so antisense sequences should be considered as normal ones after a simple transformation in their inverse-complementary forms.

Artifacts↗

VISA: Visual Sequence Analysis for the comparison of multiple amino acid sequences.

VISA (VIsual Sequence Analysis) is a software package that displays global similarities within a set of related protein sequences. The program identifies amino acid patterns that are common to many members of the set of sequences and displays them as a series of histograms. Individual peaks on the display can be assigned a color and analogous peaks in the other sequences are then automatically marked in the same color. This can be repeated for each significant peak and leads to a display in which major matching segments of multiple amino acid sequences appear as dominant peaks of the histograms with matching colors. These peaks usually correspond to the conserved sequence motifs that are characteristic of particular proteins. An extensive set of software tools is included to help the localization, visualization and analysis of the global similarities displayed. VISA provides a graphic overview of the sequence similarity that can help to understand the architecture of the protein family and can be helpful while designing experiments to probe function.

Algorithms↗

Sequence analysis by additive scales: DNA structure for sequences and repeats of all lengths.

MOTIVATION: DNA structure plays an important role in a variety of biological processes. Different di- and tri-nucleotide scales have been proposed to capture various aspects of DNA structure including base stacking energy, propeller twist angle, protein deformability, bendability, and position preference. Yet, a general framework for the computational analysis and prediction of DNA structure is still lacking. Such a framework should in particular address the following issues: (1) construction of sequences with extremal properties; (2) quantitative evaluation of sequences with respect to a given genomic background; (3) automatic extraction of extremal sequences and profiles from genomic databases; (4) distribution and asymptotic behavior as the length N of the sequences increases; and (5) complete analysis of correlations between scales. RESULTS: We develop a general framework for sequence analysis based on additive scales, structural or other, that addresses all these issues. We show how to construct extremal sequences and calibrate scores for automatic genomic and database extraction. We show that distributions rapidly converge to normality as Nincreases. Pairwise correlations between scales depend both on background distribution and sequence length and rapidly converge to an analytically predictable asymptotic value. For di- and tri-nucleotide scales, normal behavior and asymptotic correlation values are attained over a characteristic window length of about 10-15 bp. With a uniform background distribution, pairwise correlations between empirically-derived scales remain relatively small and roughly constant at all lengths, except for propeller twist and protein deformability which are positively correlated. There is a positive (resp. negative) correlation between dinucleotide base stacking (resp. propeller twist and protein deformability) and AT-content that increases in magnitude with length. The framework is applied to the analysis of various DNA tandem repeats. We derive exact expressions for counting the number of repeat unit classes at all lengths. Tandem repeats are likely to result from a variety of different mechanisms, a fraction of which is likely to depend on profiles characterized by extreme structural features.

Animals↗

A kinase sequence database: sequence alignments and family assignment.

UNLABELLED: The Kinase Sequence Database (KSD) located at http://kinase.ucsf.edu/ksd contains information on 290 protein kinase families derived by profile-based clustering of the non-redundant list of sequences obtained from a GenBank-wide search. Included in the database are a total of 5,041 protein kinases from over 100 organisms. Clustering into families is based on the extent of homology within the kinase catalytic domain (250-300 residues in length). Alignments of the families are viewed by interactive Excel-based sequence spreadsheets. In addition, KSD features evolutionary trees derived for each family and detailed information on each sequence as well as links to the corresponding GenBank entries. Sequence manipulation tools, such as evolutionary tree generation, novel sequence assignment, and statistical analysis, are also provided. AVAILABILITY: The kinase sequence database is a web-based service accessible at http://kinase.ucsf.edu/ksd CONTACT: buzko@cmp.ucsf.edu; shokat@cmp.ucsf.edu/ksd

Cluster Analysis↗

Prediction of the coding sequences of unidentified human genes. VI. The coding sequences of 80 new genes (KIAA0201-KIAA0280) deduced by analysis of cDNA clones from cell line KG-1 and brain.

In this series of projects of sequencing human cDNA clones which correspond to relatively long and nearly full-length transcripts, we newly determined the sequences of 80 clones, and predicted the coding sequences of the corresponding genes, named KIAA0201 to KIAA0280. Among the sequenced clones, 68 were obtained from human immature myeloid cell line KG-1 and 12 from human brain. The average size of the clones was 5.3 kb, and that of distinct ORFs in clones was 2.8 kb, corresponding to a protein of approximately 100 kDa. Computer search against the public databases indicated that the sequences of 22 genes were unrelated to any reported genes, while the remaining 58 genes carried sequences which show some similarities to known genes. Protein motifs that matched those in the PROSITE motif database were found in 25 genes and significant transmembrane domains were identified in 30 genes. Among the known genes to which significant similarity was shown, the genes that play key roles in regulation of developmental stages, apoptosis and cell-to-cell interaction were included. Taking into account of both the search data on sequence similarity and protein motifs, at least seven genes were considered to be related to transcriptional regulation and six genes to signal transduction. When the expression profiles of the cDNA clones were examined with different human tissues, about half of the clones from brain (5 of 11) showed significant tissue-specificity, while approximately 80% of the genes from KG-1 were expressed ubiquitously.

Amino Acid Sequence↗

Prediction of the coding sequences of unidentified human genes. VII. The complete sequences of 100 new cDNA clones from brain which can code for large proteins in vitro.

In this series of projects of sequencing human cDNA clones which correspond to relatively long transcripts, we newly determined the entire sequences of 100 cDNA clones which were screened on the basis of the potentiality of coding for large proteins in vitro. The cDNA libraries used were the fractions with average insert sizes from 5.3 to 7.0 kb of the size-fractionated cDNA libraries from human brain. The randomly sampled clones were single-pass sequenced from both the ends to select clones that are not registered in the public database. Then their protein-coding potentialities were examined by an in vitro transcription/translation system, and the clones that generated proteins larger than 60 kDa were entirely sequenced. Each clone gave a distinct open reading frame (ORF), and the length of the ORF was roughly coincident with the approximate molecular mass of the in vitro product estimated from its mobility on SDS-polyacrylamide gel electrophoresis. The average size of the cDNA clones sequenced was 6.1 kb, and that of the ORFs corresponded to 1200 amino acid residues. By computer-assisted analysis of the sequences with DNA and protein-motif databases (GenBank and PROSITE databases), the functions of at least 73% of the gene products could be anticipated, and 88% of them (the products of 64 clones) were assigned to the functional categories of proteins relating to cell signaling/communication, nucleic acid managing, and cell structure/motility. The expression profiles in a variety of tissues and chromosomal locations of the sequenced clones have been determined. According to the expression spectra, approximately 11 genes appeared to be predominantly expressed in brain. Most of the remaining genes were categorized into one of the following classes: either the expression occurs in a limited number of tissues (31 genes) or the expression occurs ubiquitously in all but a few tissues (47 genes).

Blotting, Northern↗

The L1 family (KpnI family) sequence near the 3' end of human beta-globin gene may have been derived from an active L1 sequence.

We previously reported that some L1 family (KpnI family) members are closely associated with the Alu family sequence. To understand the details of the L1-Alu association, the structure of a L1-Alu unit downstream from the beta-globin gene was compared between human and primates. The results revealed that the L1-Alu-associated sequence was formed by the insertion of the L1 sequence, T beta G41, into the 3' poly A tract of the preexisting Alu family sequence. It was estimated that the T beta G41 sequence was inserted after the divergence of Old World monkeys and hominoids and before the divergence of orang-utan and common ancestor of other higher hominoids. From the calculation of the mutation rates of L1 sequences, it was suggested that the T beta G41 was derived from an active L1 sequence which was able to encode reverse transcriptase-related protein.

Animals↗

Sequence organization and developmentally regulated transcription of a family of repetitive DNA sequences of Xenopus laevis.

Members of a family of DNA sequences of Xenopus laevis have been cloned and sequenced. Molecular analyses revealed that these sequences are moderately repetitive and dispersed throughout the genome. The sequences of seven clones were compared. Two of the clones lie in the globin gene cluster; one 5' to the adult alpha 1 gene, and the other in the first intron of the tadpole alpha 1 gene. In all clones, the homologous region begins at the same site, but the lengths of the common regions vary from 123 bp to over 320 bp due to heterogeneous 3' ends. Some of the repeats are bracketed by direct and/or inverted repeats, and relatively large palindromes were found 5' to the common region in some clones. These characteristics, and the presence of a repeat 5' to one of a pair of duplicated alpha genes suggests that some family members may be capable of transposition. A number of interesting features were found in the sequences, including multiple elements similar to the yeast autonomously replicating sequence, and a sequence which is about 80% homologous to the first 30 bases of the SV40 enhancer. Transcription studies revealed that homologous transcripts are detectable beginning at neurulation, increase in concentration up to stage 45, and disappear by metamorphosis. Implications of these data are discussed.

Aging↗

cDNA sequence and deduced amino acid sequence of human preprocolipase.

Complementary DNA clones for human pancreatic colipase were identified in human pancreatic cDNA libraries by hybridization with a pool of synthetic oligonucleotides containing all possible coding sequences for amino acids 75 to 80 of the partial human colipase protein sequence (Sternby, et al. Biochim Biophys Acta 1984;784:75). Alignment of overlapping cDNA clones yielded an mRNA sequence of 504 nucleotides [not including the poly(A) tail] encoding a polypeptide of 112 amino acids. The prepeptide comprised 17 amino acids, with an amino-terminal cluster of charged residues followed by a hydrophobic core of 12 residues typical of leader sequences. The deduced human procolipase sequence comprised 95 residues, including a propeptide of 5 residues. It was in complete agreement with the partial sequence previously obtained by protein sequencing. Northern blot analysis revealed that the polyadenylated preprocolipase transcript had a length of approximately 680 nucleotides.

Amino Acid Sequence↗

Integrated mapping and sequencing of a 115 kb DNA fragment from Bacillus subtilis: sequence analysis of a 21 kb segment containing the sigL locus.

A sequence strategy which combines a low redundancy shotgun approach and directed sequencing has been elaborated. Essentially, the sequences, as well as the size of the fragments utilized for a low coverage shotgun approach, were exploited for the construction of a physical map of the region to be sequenced. The latter considerably simplified the subsequent directed sequencing steps. We report the physical mapping of a 115 kb segment which covers nearly 100 kb of the hisA-cysB region of the Bacillus subtilis chromosome and contains previously sequenced genes sigL and sacB. Sequencing and analysis of a 21305 bp segment, which includes the sigL locus, revealed 21 ORFs, apparently belonging to at least seven transcription units. This segment has a G + C content greater than 47%, compared to 43% characteristic of the flanking regions, and mainly consists of genes whose products seem to be involved in the synthesis of an exopolysaccharide. These observations leave open the possibility that the analysed fragment has been acquired through horizontal transfer.

Bacillus subtilis↗

A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.

The expressed human genome is being sequenced and analyzed by disparate groups producing disparate data. The majority of the identified coding portion is in the form of expressed sequence tags (ESTs). The need to discover exonic representation and expression forms of full-length cDNAs for each human gene is frustrated by the partial and variable quality nature of this data delivery. A highly redundant human EST data set has been processed into integrated and unified expressed transcript indices that consist of hierarchically organized human transcript consensi reflecting gene expression forms and genetic polymorphism within an index class. The expression index and its intermediate outputs include cleaned transcript sequence, expression, and alignment information and a higher fidelity subset, SANIGENE. The STACK_PACK clustering system has been applied to dbEST release 121598 (GenBank version 110). Sixty-four percent of 1,313, 103 Homo sapiens ESTs are condensed into 143,885 tissue level multiple sequence clusters; linking through clone-ID annotations produces 68,701 total assemblies, such that 81% of the original input set is captured in a STACK multiple sequence or linked cluster. Indexing of alignments by substituent EST accession allows browsing of the data structure and its cross-links to UniGene. STACK metaclusters consolidate a greater number of ESTs by a factor of 1. 86 with respect to the corresponding UniGene build. Fidelity comparison with genome reference sequence AC004106 demonstrates consensus expression clusters that reflect significantly lower spurious repeat sequence content and capture alternate splicing within a whole body index cluster and three STACK v.2.3 tissue-level clusters. Statistics of a staggered release whole body index build of STACK v.2.0 are presented.

Algorithms↗

Complete subtyping of the HLA-A locus by sequence-specific amplification followed by direct sequencing or single-strand conformation polymorphism analysis.

A variety of reasons related to the HLA class I system has complicated the application of molecular approaches to HLA class I typing. Here we present a PCR-based HLA-A typing strategy considering the sequence variations of the two most polymorphic exons which allows complete subtyping of the HLA-A locus. The method is based on a sequence-specific amplification identifying the serologically defined HLA-A specificities. The PCR products generated by these group-specific primers bear the sequence information necessary for a postamplification specificity step. The primer pairs are located within one exon, either exon 2 or exon 3, which avoids amplification of polymorphic intron sequences allowing subsequent single-strand conformation polymorphism analysis and facilitating direct sequencing. Using this method we investigated 48 cell lines and 153 clinical samples. 23 PCR reactions are performed per individual for the assignment of the serological specificities A1-A80. The reproducibility was 100% in all cell lines and 85 clinical samples typed on two separate occasions. With the exception of 13 out of 231 possible serological combinations all homozygous and heterozygous combinations of A1-A80 can be distinguished by specific amplification patterns. Comparing the PCR based typing results with those of serology in 12% a discrepancy was found. Solid-phase sequencing or SSCP analysis of the group-specific PCR fragments allowed complete subtyping of the HLA-A locus. This strategy can identify all 48 HLA-A alleles based on the sequence variations of the 2nd and 3rd exon. 1128 homozygous and heterozygous allele combinations are possible for the HLA-A locus. Only 4 out of these 1128 allele combinations remained unresolved.

Alleles↗