Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA language model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

242 records · Page 14Linked to original sources

Parallel evolution of the genetic code in arthropod mitochondrial genomes.

The genetic code provides the translation table necessary to transform the information contained in DNA into the language of proteins. In this table, a correspondence between each codon and each amino acid is established: tRNA is the main adaptor that links the two. Although the genetic code is nearly universal, several variants of this code have been described in a wide range of nuclear and organellar systems, especially in metazoan mitochondria. These variants are generally found by searching for conserved positions that consistently code for a specific alternative amino acid in a new species. We have devised an accurate computational method to automate these comparisons, and have tested it with 626 metazoan mitochondrial genomes. Our results indicate that several arthropods have a new genetic code and translate the codon AGG as lysine instead of serine (as in the invertebrate mitochondrial genetic code) or arginine (as in the standard genetic code). We have investigated the evolution of the genetic code in the arthropods and found several events of parallel evolution in which the AGG codon was reassigned between serine and lysine. Our analyses also revealed correlated evolution between the arthropod genetic codes and the tRNA-Lys/-Ser, which show specific point mutations at the anticodons. These rather simple mutations, together with a low usage of the AGG codon, might explain the recurrence of the AGG reassignments.

Animals↗

Simplifying amino acid alphabets by means of a branch and bound algorithm and substitution matrices.

MOTIVATION: Protein and DNA are generally represented by sequences of letters. In a number of circumstances simplified alphabets (where one or more letters would be represented by the same symbol) have proved their potential utility in several fields of bioinformatics including searching for patterns occurring at an unexpected rate, studying protein folding and finding consensus sequences in multiple alignments. The main issue addressed in this paper is the possibility of finding a general approach that would allow an exhaustive analysis of all the possible simplified alphabets, using substitution matrices like PAM and BLOSUM as a measure for scoring. RESULTS: The computational approach presented in this paper has led to a computer program called AlphaSimp (Alphabet Simplifier) that can perform an exhaustive analysis of the possible simplified amino acid alphabets, using a branch and bound algorithm together with standard or user-defined substitution matrices. The program returns a ranked list of the highest-scoring simplified alphabets. When the extent of the simplification is limited and the simplified alphabets are maintained above ten symbols the program is able to complete the analysis in minutes or even seconds on a personal computer. However, the performance becomes worse, taking up to several hours, for highly simplified alphabets. AVAILABILITY: AlphaSimp and other accessory programs are available at http://bioinformatics.cribi.unipd.it/alphasimp

Algorithms↗

Tricross : using dot-plots in sequence-id space to detect uncataloged intergenic features.

MOTIVATION: The process of determining the functional sequence content of an organism is confounded by several factors. Large protein coding sequences are relatively easy to find by statistical methods. Smaller proteins however may escape detection due to their size falling below some arbitrary researcher-defined minimum cutoff, or the inability to precisely define a promoter, or translational start (Delcher et al., Nucleic Acids Res., 27, 4636-4641, 1999). Promoter and regulatory sequences themselves are difficult to define due to a significant amount of allowable sequence variation, as well as a probable lack of any completely accurate whole-organismal gene catalogs to date. Finally, certain genes coding functional RNAs may have insufficient structural or sequence constraints to be detectable by normal sequence structure/pattern searching methods (Eddy and Rivas, Bioinformatics, 16, 583-605, 2000). In those cases where there are multiple closely related organisms that have been sequenced, there is additional information that may be used in the investigation of sequence content-that being the possible conserved nature of functional sequences between the organisms. We present a method for the utilization of this conserved information to detect genes and other potentially functional sequences that may be missed by standard ORF-calling, RNA finding, and pattern matching software. The tricross programs produce a multi-way cross comparison of three sets of sequences, determine which are conserved in all three sets, and produce a graphical (Virtual Reality Modelling Language-VRML; (ISO/IEC 14772-1: 1997, VDC), 1997) representation as well as alignments of all sequence triples found. The software can also be applied to a pair of sequence sets, though the noise in the results increases. RESULTS: Tricross has been used to examine the intergenic-sequence content of the three archaeal Pyrococcus genomes to determine the most highly related sequences remaining between the annotated protein and RNA coding sequences. Set to relatively stringent similarity requirements for the search, tricross found 101 intergenic sequences conserved among the three organisms. Interestingly, 29 of these appear to contain members of a family of small RNA molecules (Kiss-Laszlo et al., EMBO J., 17, 797-807, 1998) only recently discovered in the Archaea (Armbruster, OSU, Diss., 1988; Omer et al., Science, 288, 517-522, 2000; Gaspin et al., J. Mol. Biol., 297, 895-906, 2000). While some of the remaining 72 appear to be individual highly conserved promoter sequences, others have no currently known biological significance. Although originally developed to facilitate the examination of intergenic sequences, none of the tricross logic is inherently specific to intergenic sequences. The software can also be applied to gene sequences, and has been used to produce inter-genomic gene order dot-plots for Haemophilus influenzae (Fleischmann et al., Science, 269, 496-512, 1995) versus H.ducreyi (unpublished data), and Neisseria meningiditis Z2491 (serogroup A) (Parkhill et al., Nature, 404, 502-506, 2000) versus Neisseria meningiditis Z58 (serogroup B) (Tettelin et al., Science, 287, 1809-1815, 2000) versus Neisseria gonorrhoeae (Lewis et al., http://micro-gen.ouhsc.edu/, 2000). AVAILABILITY: The tricross software package is available from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html. CONTACT: ray@biosci.ohio-state.edu; daniels.7@osu.edu; munsonr@pediatrics.ohio-state.edu SUPPLEMENTARY INFORMATION: Additional data from the cross-genomic comparisons examined in the discussion section are linked from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html.

Base Sequence↗

VI. Genome structure and cognitive map of Williams syndrome.

Williams syndrome (WMS) is a most compelling model of human cognition, of human genome organization, and of evolution. Due to a deletion in chromosome band 7q11.23, subjects have cardiovascular, connective tissue, and neurodevelopmental deficits. Given the striking peaks and valleys in neurocognition including deficits in visual-spatial and global processing, preserved language and face processing, hypersociability, and heightened affect, the goal of this work has been to identify the genes that are responsible, the cause of the deletion, and its origin in primate evolution. To do this, we have generated an integrated physical, genetic, and transcriptional map of the WMS and flanking regions using multicolor metaphase and interphase fluorescence in situ hybridization (FISH) of bacterial artificial chromosomes (BACs) and P1 artificial chromosomes (PACs), BAC end sequencing, PCR gene marker and microsatellite, large-scale sequencing, cDNA library, and database analyses. The results indicate the genomic organization of the WMS region as two nested duplicated regions flanking a largely single-copy region. There are at least two common deletion breakpoints, one in the centromeric and at least two in the telomeric repeated regions. Clones anchoring the unique to the repeated regions are defined along with three new pseudogene families. Primate studies indicate an evolutionary hot spot for chromosomal inversion in the WMS region. A cognitive phenotypic map of WMS is presented, which combines previous data with five further WMS subjects and three atypical WMS subjects with deletions; two larger (deleted for D7S489L) and one smaller, deleted for genes telomeric to FZD9, through LIMK1, but not WSCR1 or telomeric. The results establish regions and consequent gene candidates for WMS features including mental retardation, hypersociability, and facial features. The approach provides the basis for defining pathways linking genetic underpinnings with the neuroanatomical, functional, and behavioral consequences that result in human cognition.

Adolescent↗

Antiquity, geographic contiguity and genetic affinity among Tibeto-Burman populations of India: a microsatellite study.

BACKGROUND: The Tibeto-Burman (TB) populations are one of the four major linguistic population groups of India. They are considered belonging to different stocks and show wide variation in culture and language; however, their genetic relationship, antiquity and migration history among the regional populations has been little investigated. Molecular genetic studies are expected to clearly show the antiquity and genetic diversity of these populations. AIM: This paper seeks to understand the extent and magnitude of genetic affinities and diversities among 14 TB populations (12 Indian and two global groups), investigate the findings based on classical genetic markers and verify the historical accounts of their migration and genetic history based on 12 microsatellite markers. SUBJECTS AND METHODS: The allele frequency data for 12 STR loci of 13 Asian (Tibeto-Burman) populations were obtained from the literature and the Adi Pasi data was obtained by microsatellite typing of their blood samples. The 12 loci studied are D5S818, FGA, D8S1179, D21S11, D7S820, CSF1PO, D3S1358, THO1, D13S317, vWa, TPOX, D18S51. Three different distance measures, two phylogenetic trees and PCA plot have been employed to understand the genetic relationship of the studied populations. RESULTS: Average heterozygosity values range from 68 to 79% and the average G(ST) value is 4.7%. The dendrogram, based on the D(A) distance, shows the clustering of populations based on their diversities and geographical contiguity; the Mizoram and Arunachal Pradesh populations especially cluster together, populations from Sikkim form a separate subcluster and Manipur populations along with the Garo of West Bengal separate out from the other clusters. The Harpending and Ward regression model shows isolated populations positioned below the regression line and others, who experience external gene flow, placed above the line. The results support folklore migration accounts of their possible antiquity with the Tibetan and southern Chinese populations. CONCLUSIONS: Overall, geographic contiguity, punctuated by isolating barriers, is a major influencing factor of genetic affinity among the TB population; contiguous populations within a region show greater genetic relationship than with distant TB populations over a wide geographical area. The results of the microsatellite study also support the history of diverse routes of migration of these populations.

DNA Fingerprinting↗

Traces of archaic mitochondrial lineages persist in Austronesian-speaking Formosan populations.

Genetic affinities between aboriginal Taiwanese and populations from Oceania and Southeast Asia have previously been explored through analyses of mitochondrial DNA (mtDNA), Y chromosomal DNA, and human leukocyte antigen loci. Recent genetic studies have supported the "slow boat" and "entangled bank" models according to which the Polynesian migration can be seen as an expansion from Melanesia without any major direct genetic thread leading back to its initiation from Taiwan. We assessed mtDNA variation in 640 individuals from nine tribes of the central mountain ranges and east coast regions of Taiwan. In contrast to the Han populations, the tribes showed a low frequency of haplogroups D4 and G, and an absence of haplogroups A, C, Z, M9, and M10. Also, more than 85% of the maternal lineages were nested within haplogroups B4, B5a, F1a, F3b, E, and M7. Although indicating a common origin of the populations of insular Southeast Asia and Oceania, most mtDNA lineages in Taiwanese aboriginal populations are grouped separately from those found in China and the Taiwan general (Han) population, suggesting a prevalence in the Taiwanese aboriginal gene pool of its initial late Pleistocene settlers. Interestingly, from complete mtDNA sequencing information, most B4a lineages were associated with three coding region substitutions, defining a new subclade, B4a1a, that endorses the origin of Polynesian migration from Taiwan. Coalescence times of B4a1a were 13.2 +/- 3.8 thousand years (or 9.3 +/- 2.5 thousand years in Papuans and Polynesians). Considering the lack of a common specific Y chromosomal element shared by the Taiwanese aboriginals and Polynesians, the mtDNA evidence provided here is also consistent with the suggestion that the proto-Oceanic societies would have been mainly matrilocal.

Asian People↗

BAMarraytrade mark: Java software for Bayesian analysis of variance for microarray data.

BACKGROUND: DNA microarrays open up a new horizon for studying the genetic determinants of disease. The high throughput nature of these arrays creates an enormous wealth of information, but also poses a challenge to data analysis. Inferential problems become even more pronounced as experimental designs used to collect data become more complex. An important example is multigroup data collected over different experimental groups, such as data collected from distinct stages of a disease process. We have developed a method specifically addressing these issues termed Bayesian ANOVA for microarrays (BAM). The BAM approach uses a special inferential regularization known as spike-and-slab shrinkage that provides an optimal balance between total false detections and total false non-detections. This translates into more reproducible differential calls. Spike and slab shrinkage is a form of regularization achieved by using information across all genes and groups simultaneously. RESULTS: BAMarray is a graphically oriented Java-based software package that implements the BAM method for detecting differentially expressing genes in multigroup microarray experiments (up to 256 experimental groups can be analyzed). Drop-down menus allow the user to easily select between different models and to choose various run options. BAMarraycan also be operated in a fully automated mode with preselected run options. Tuning parameters have been preset at theoretically optimal values freeing the user from such specifications. BAMarray provides estimates for gene differential effects and automatically estimates data adaptive, optimal cutoff values for classifying genes into biological patterns of differential activity across experimental groups. A graphical suite is a core feature of the product and includes diagnostic plots for assessing model assumptions and interactive plots that enable tracking of prespecified gene lists to study such things as biological pathway perturbations. The user can zoom in and lasso genes of interest that can then be saved for downstream analyses. CONCLUSION: BAMarray is user friendly platform independent software that effectively and efficiently implements the BAM methodology. Classifying patterns of differential activity is greatly facilitated by a data adaptive cutoff rule and a graphical suite. BAMarray is licensed software freely available to academic institutions. More information can be found at http://www.bamarray.com.

Algorithms↗

FLAME, a novel fuzzy clustering method for the analysis of DNA microarray data.

BACKGROUND: Data clustering analysis has been extensively applied to extract information from gene expression profiles obtained with DNA microarrays. To this aim, existing clustering approaches, mainly developed in computer science, have been adapted to microarray data analysis. However, previous studies revealed that microarray datasets have very diverse structures, some of which may not be correctly captured by current clustering methods. We therefore approached the problem from a new starting point, and developed a clustering algorithm designed to capture dataset-specific structures at the beginning of the process. RESULTS: The clustering algorithm is named Fuzzy clustering by Local Approximation of MEmbership (FLAME). Distinctive elements of FLAME are: (i) definition of the neighborhood of each object (gene or sample) and identification of objects with "archetypal" features named Cluster Supporting Objects, around which to construct the clusters; (ii) assignment to each object of a fuzzy membership vector approximated from the memberships of its neighboring objects, by an iterative converging process in which membership spreads from the Cluster Supporting Objects through their neighbors. Comparative analysis with K-means, hierarchical, fuzzy C-means and fuzzy self-organizing maps (SOM) showed that data partitions generated by FLAME are not superimposable to those of other methods and, although different types of datasets are better partitioned by different algorithms, FLAME displays the best overall performance. FLAME is implemented, together with all the above-mentioned algorithms, in a C++ software with graphical interface for Linux and Windows, capable of handling very large datasets, named Gene Expression Data Analysis Studio (GEDAS), freely available under GNU General Public License. CONCLUSION: The FLAME algorithm has intrinsic advantages, such as the ability to capture non-linear relationships and non-globular clusters, the automated definition of the number of clusters, and the identification of cluster outliers, i.e. genes that are not assigned to any cluster. As a result, clusters are more internally homogeneous and more diverse from each other, and provide better partitioning of biological functions. The clustering algorithm can be easily extended to applications different from gene expression analysis.

Algorithms↗