Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comprehensive analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Recently integrated Alu elements and human genomic diversity.

A comprehensive analysis of two Alu Y lineage subfamilies was undertaken to assess Alu-associated genomic diversity and identify new Alu insertion polymorphisms for the study of human population genetics. Recently integrated Alu elements (283) from the Yg6 and Yi6 subfamilies were analyzed by polymerase chain reaction (PCR), and 25 of the loci analyzed were polymorphic for insertion presence/absence within the genomes of a diverse array of human populations. These newly identified Alu insertion polymorphisms will be useful tools for the study of human genomic diversity. Our screening of the Alu insertion loci also resulted in the recovery of several "young" Alu elements that resided at orthologous positions in nonhuman primate genomes. Sequence analysis demonstrated these "young" Alu insertions were the products of gene conversion events of older, preexisting Alu elements or independent parallel forward insertions of older Alu elements in the same short genomic region. The level of gene conversion between Alu elements suggests that it may have an influence on the single nucleotide polymorphism within Alu elements in the genome. We have also identified two genomic deletions associated with the retroposition and insertion of Alu Y lineage elements into the human genome. This type of Alu retroposition-mediated genomic deletion is a novel source of lineage-specific evolution within primate genomes.

Alu Elements↗

The ingi and RIME non-LTR retrotransposons are not randomly distributed in the genome of Trypanosoma brucei.

The ingi (long and autonomous) and RIME (short and nonautonomous) non--long-terminal repeat retrotransposons are the most abundant mobile elements characterized to date in the genome of the African trypanosome Trypanosoma brucei. These retrotransposons were thought to be randomly distributed, but a detailed and comprehensive analysis of their genomic distribution had not been performed until now. To address this question, we analyzed the ingi/RIME sequences and flanking sequences from the ongoing T. brucei genome sequencing project (TREU927/4 strain). Among the 81 ingi/RIME elements analyzed, 60% are complete, and 7% of the ingi elements (approximately 15 copies per haploid genome) appear to encode for their own transposition. The size of the direct repeat flanking the ingi/RIME retrotransposons is conserved (i.e., 12-bp), and a strong 11-bp consensus pattern precedes the 5'-direct repeat. The presence of a consensus pattern upstream of the retroelements was confirmed by the analysis of the base occurrence in 294 GSS containing 5'-adjacent ingi/RIME sequences. The conserved sequence is present upstream of ingis and RIMEs, suggesting that ingi-encoded enzymatic activities are used for retrotransposition of RIMEs, which are short nonautonomous retroelements. In conclusion, the ingi and RIME retroelements are not randomly distributed in the genome of T. brucei and are preceded by a conserved sequence, which may be the recognition site of the ingi-encoded endonuclease.

Amino Acid Sequence↗

Unexpected diversity and differential success of DNA transposons in four species of entamoeba protozoans.

We report the first comprehensive analysis of transposable element content in the compact genomes (approximately 20 Mb) of four species of Entamoeba unicellular protozoans for which draft sequences are now available. Entamoeba histolytica and Entamoeba dispar, two human parasites, have many retrotransposons, but few DNA transposons. In contrast, the reptile parasite Entamoeba invadens and the free-living Entamoeba moshkovskii contain few long interspersed elements but harbor diverse and recently amplified populations of DNA transposons. Representatives of three DNA transposase superfamilies (hobo/Activator/Tam3, Mutator, and piggyBac) were identified for the first time in a protozoan species in addition to a variety of members of a fourth superfamily (Tc1/mariner), previously reported only from ciliates and Trichomonas vaginalis among protozoans. The diversity of DNA transposons and their differential amplification among closely related species with similar compact genomes are discussed in the context of the biology of Entamoeba protozoans.

Amino Acid Sequence↗

A mitogenomic timescale for birds detects variable phylogenetic rates of molecular evolution and refutes the standard molecular clock.

Current understanding of the diversification of birds is hindered by their incomplete fossil record and uncertainty in phylogenetic relationships and phylogenetic rates of molecular evolution. Here we performed the first comprehensive analysis of mitogenomic data of 48 vertebrates, including 35 birds, to derive a Bayesian timescale for avian evolution and to estimate rates of DNA evolution. Our approach used multiple fossil time constraints scattered throughout the phylogenetic tree and accounts for uncertainties in time constraints, branch lengths, and heterogeneity of rates of DNA evolution. We estimated that the major vertebrate lineages originated in the Permian; the 95% credible intervals of our estimated ages of the origin of archosaurs (258 MYA), the amniote-amphibian split (356 MYA), and the archosaur-lizard divergence (278 MYA) bracket estimates from the fossil record. The origin of modern orders of birds was estimated to have occurred throughout the Cretaceous beginning about 139 MYA, arguing against a cataclysmic extinction of lineages at the Cretaceous/Tertiary boundary. We identified fossils that are useful as time constraints within vertebrates. Our timescale reveals that rates of molecular evolution vary across genes and among taxa through time, thereby refuting the widely used mitogenomic or cytochrome b molecular clock in birds. Moreover, the 5-Myr divergence time assumed between 2 genera of geese (Branta and Anser) to originally calibrate the standard mitochondrial clock rate of 0.01 substitutions per site per lineage per Myr (s/s/l/Myr) in birds was shown to be underestimated by about 9.5 Myr. Phylogenetic rates in birds vary between 0.0009 and 0.012 s/s/l/Myr, indicating that many phylogenetic splits among avian taxa also have been underestimated and need to be revised. We found no support for the hypothesis that the molecular clock in birds "ticks" according to a constant rate of substitution per unit of mass-specific metabolic energy rather than per unit of time, as recently suggested. Our analysis advances knowledge of rates of DNA evolution across birds and other vertebrates and will, therefore, aid comparative biology studies that seek to infer the origin and timing of major adaptive shifts in vertebrates.

Animals↗

A roadmap of tandemly arrayed genes in the genomes of human, mouse, and rat.

Tandemly arrayed genes (TAGs) play an important functional and physiological role in the genome. Most previous studies have focused on individual TAG families in a few species, yet a broad characterization of TAGs is not available. Here we identified all TAGs in the genomes of humans, mouse, and rat and performed a comprehensive analysis of TAG distribution, TAG sizes, TAG orientations and intergenic distances, and TAG functions. TAGs account for about 14-17% of all genes in the genome and nearly one-third of all duplicated genes, highlighting the predominant role that tandem duplication plays in gene duplication. For all species, TAG distribution is highly heterogeneous along chromosomes and some chromosomes are enriched with TAG forests, whereas others are enriched with TAG deserts. The majority of TAGs are of size 2 for all genomes, similar to the previous findings in Caenorhabditis elegans, Arabidopsis thaliana, and Oryza sativa, suggesting that it is a rather general phenomenon in eukaryotes. The comparison with the genome patterns shows that TAG members have a significantly higher proportion of parallel gene orientation in all species, corroborating Graham's claim that parallel orientation is the preferred form of orientation in TAGs. Moreover, TAG members with parallel orientation tend to be closer to each other than all neighboring genes in the genome with parallel orientation. The analyses of Gene Ontology function indicate that genes with receptor or binding activities are significantly overrepresented by TAGs. Computer simulation reveals that random gene rearrangements have little effect on the statistics of TAGs for all genomes. Finally, the average proportion of TAGs shows a trend of increase with the increase of family sizes, although the correlation between TAG proportions in individual families and family sizes is not significant.

Animals↗

The world according to Maf.

Maf family proteins are so named because of their structural similarity to the founding member, the oncoprotein v-Maf. The small Maf proteins (MafF, MafG and MafK), as do all family members, include a characteristic basic region linked to a leucine zipper (b-Zip) domain which mediate DNA binding and subunit dimerization respectively. The small Maf proteins form homodimers or heterodimers with other b-Zip proteins present in the cell and bind to Maf recognition elements (MARE) in DNA. Since they lack known transcriptional activation domains, the small Maf proteins function either as obligatory heterodimeric partner molecules with numerous large subunits, discussed below, or alternatively as homo- or heterodimeric transcriptional repressors. The three small Maf proteins are expressed in a number of overlapping tissues, but their expression profiles nonetheless appear to be under meticulous tissue- and developmental stage-specific control. The MARE bears a striking resemblance to the NF-E2 binding sequence. NF-E2 binding sites in the human beta-globin locus control region have been directly implicated as integral components in the circuitry required for eliciting changes in chromatin structure that precede globin gene activation. While the NF-E2 DNA sequence has been shown to be important for erythroid-specific gene regulation, a growing list of other genes may also be regulated through the same, or very similar, cis elements in non-erythroid cells. Taken together, these observations argue that comprehensive analysis of the activities of the small Maf proteins may provide a unique perspective for expanding our understanding of transcriptional regulation that can be elicited through interacting transcription factor networks.

Animals↗

GXD: a gene expression database for the laboratory mouse. The Gene Expression Database Group.

The Gene Expression Database (GXD) is a community resource that stores and integrates expression information for the laboratory mouse, with a particular emphasis on mouse development, and makes these data freely available in formats appropriate for comprehensive analysis. GXD is implemented as a relational database and integrated with the Mouse Genome Database (MGD) to enable global analysis of genotype, expression and phenotype information. Interconnections with sequence databases and with databases from other species further extend GXD's utility for the analysis of gene expression data. GXD is available through the Mouse Genome Informatics Web Site at http://www.informatics.jax.org/

Animals↗

Identification of genes regulated by muscarinic acetylcholine receptors: application of an improved and statistically comprehensive mRNA differential display technique.

In order to identify genes that are regulated by muscarinic acetylcholine receptors, we developed an mRNA differential display technique (DD) approach. By increasing redundancy and by evaluating optimised reagents and conditions for reverse transcription of total RNA, PCR and separation of PCR products, we generated a DD protocol that yields highly consistent results. A set of 64 distinct random primers was specifically designed in order to approach a statistically comprehensive analysis of all mRNA species in a defined cell population. This modified DD protocol was applied to total RNA of HEK293 cells stably expressing muscarinic m1 acetylcholine receptors and cells stimulated with the receptor agonist carbachol were compared to identical but non-stimulated cells. In 81 of 192 possible PCR experiments, 38 differential bands were identified. Sequence analysis followed by northern blot analyses confirmed differentially expressed genes in 19 of 23 bands analysed. These represented 10 distinct immediate-early genes that were up-regulated by m1AChR activation: Egr-1, Egr-2, Egr-3, NGFi-B, ETR101, c- jun, jun -D, Gos-3 and hcyr61, as well as the unknown gene Gig-2. These data show that this improved DD protocol can be readily applied to reliably identify differentially expressed genes.

Base Sequence↗

Axeldb: a Xenopus laevis database focusing on gene expression.

Axeldb is a database storing and integrating gene expression patterns and DNA sequences identified in a large-scale in situ hybridization study in Xenopus laevis embryos. The data are organised in a format appropriate for comprehensive analysis, and enable comparison of images of expression pattern for any given set of genes. Information on literature, cDNA clones and their availability, nucleotide sequences, expression pattern and accompanying pictures are available. Current developments are aimed toward the interconnection with other databases and the integration of data from the literature. Axeldb is implemented using an ACEDB database system, and available through the web at http://www.dkfz-heidelberg.de/abt0135/axeldb.htm

Animals↗

Proteomics of Mycoplasma genitalium: identification and characterization of unannotated and atypical proteins in a small model genome.

We present the results of a comprehensive analysis of the proteome of Mycoplasma genitalium (MG), the smallest autonomously replicating organism that has been completely sequenced. Our aim was to identify and characterize all soluble proteins in MG that are structurally and functionally uncharacterized. We were particularly interested in identifying proteins that differed significantly from typical globular proteins, for example, proteins which are unstructured in the absence of a 'partner' molecule or those that exhibit unusual thermodynamic properties. This work is complementary to other structural genomics projects whose primary aim is to determine the three-dimensional structures of proteins with unknown folds. We have identified all the full-length open reading frames (ORFs) in MG that have no homologs of known structure and are of unknown function. Twenty-five of the total 483 ORFs fall into this category and we have expressed, purified and characterized 11 of them. We have used circular dichroism (CD) to rapidly investigate their biophysical properties. Our studies reveal that these proteins have a wide range of structures varying from highly helical to partially structured to unfolded or random coil. They also display a variety of thermodynamic properties ranging from cooperative unfolding to no detectable unfolding upon thermal denaturation. Several of these proteins are highly conserved from mycoplasma to man. Further information about target selection and CD results is available at http://bioinfo.mbb.yale.edu/genome

Amino Acid Sequence↗

The srs2 suppressor of UV sensitivity acts specifically on the RAD5- and MMS2-dependent branch of the RAD6 pathway.

The SRS2 gene encodes a helicase that affects recombination, gene conversion and DNA damage repair in the yeast Saccharomyces cerevisiae. Loss-of-function mutations in srs2 suppress the extreme sensitivity towards UV radiation of rad6 and rad18 mutants, both of which are impaired in post-replication DNA repair and damage-induced mutagenesis. A sub-branch within the RAD6 pathway is mediated by RAD5, UBC13 and MMS2, and a comprehensive analysis of the srs2 effect on other known members of the RAD6 pathway reported here now demonstrates that suppression by srs2 is specific for mutants within this RAD5-dependent sub-system. Further evidence for the concerted action of RAD5 with UBC13 and MMS2 in DNA damage repair is given by examination of the effects of cell cycle stage as well as deletion of other repair systems on the activity of post-replication repair. Finally, it is shown that MMS2, like UBC13 and many other repair genes, is transcriptionally up-regulated in response to DNA damage. The data presented here support the notion that RAD5, UBC13 and MMS2 encode an ensemble of genetically and physically interacting repair factors within the RAD6 pathway that is coordinately affected by SRS2.

Adenosine Triphosphatases↗

Leaky ribosomal scanning in mammalian genomes: significance of histone H4 alternative translation in vivo.

Like alternative splicing, leaky ribosomal scanning (LRS), which occurs at suboptimal translational initiation codons, increases the physiological flexibility of the genome by allowing alternative translation. Comprehensive analysis of 22 208 human mRNAs indicates that, although the most important positions relative to the first nucleotide of the initiation codon, -3 and +4, are usually such that support initiation (A-3 = 42%, G-3 = 36% and G+4 = 47%), only 37.4% of the genes adhere to the purine (R)-3/G+4 rule at both positions simultaneously, suggesting that LRS may occur in some of the remaining (62.6%) genes. Moreover, 12.5% of the genes lack both R-3 and G+4, potentially leading to sLRS. Compared with 11 genes known to undergo LRS, 10 genes with experimental evidence for high fidelity A+1T+2G+3 initiation codons adhered much more strongly to the R-3/G+4 rule. Among the intron-less histone genes, only the H3 genes adhere to the R-3/G+4 rule, while the H1, H2A, H2B and H4 genes usually lack either R-3 or G+4. To address in vivo the significance of the previously described LRS of H4 mRNAs, which results in alternative translation of the osteogenic growth peptide, transgenic mice were engineered that ubiquitously and constitutively express a mutant H4 mRNA with an A+1T+1 mutation. These transgenic mice, in particular the females, have a high bone mass phenotype, attributable to increased bone formation. These data suggest that many genes may fulfill cryptic functions by LRS.

Animals↗

T-STAG: resource and web-interface for tissue-specific transcripts and genes.

T-STAG (tissue-specific transcripts and genes) is a resource and web-interface, designated to analyze tissue/tumor-specific expression patterns in human and mouse transcriptomes. It integrates our refined prediction of specific expression patterns both in genes as well as in individual isoforms with man-mouse orthology data. In combination with the features for combining/contrasting the genes expressed in different tissues, T-STAG implicates important biological applications, such as the detection of differentially expressed genes in tumors, the retrieval of orthologs with significant expression in the same tissue etc. Additionally, our refined categorization of expressed sequence tags (ESTs) according to the normalization of cDNA libraries allows searching for putative low-abundant transcripts. The results are tightly linked to our visualization tools, GeneNest (expression patterns of genes) and SpliceNest (gene structure and alternative splicing). The user-friendly interface of T-STAG offers a platform for comprehensive analysis of tissue and/or tumor-specific expression patterns revealed by the EST data. T-STAG is freely accessible at http://tstag.molgen.mpg.de.

Animals↗

The Diamond STING server.

Diamond STING is a new version of the STING suite of programs for a comprehensive analysis of a relationship between protein sequence, structure, function and stability. We have added a number of new functionalities by both providing more structure parameters to the STING Database and by improving/expanding the interface for enhanced data handling. The integration among the STING components has also been improved. A new key feature is the ability of the STING server to handle local files containing protein structures (either modeled or not yet deposited to the Protein Data Bank) so that they can be used by the principal STING components: (Java)Protein Dossier ((J)PD) and STING Report. The current capabilities of the new STING version and a couple of biologically relevant applications are described here. We have provided an example where Diamond STING identifies the active site amino acids and folding essential amino acids (both previously determined by experiments) by filtering out all but those residues by selecting the numerical values/ranges for a set of corresponding parameters. This is the fundamental step toward a more interesting endeavor-the prediction of such residues. Diamond STING is freely accessible at http://sms.cbi.cnptia.embrapa.br and http://trantor.bioc.columbia.edu/SMS.

Acid Anhydride Hydrolases↗

Clustering and conservation patterns of human microRNAs.

MicroRNAs (miRNAs) are approximately 22 nt-long non-coding RNA molecules, believed to play important roles in gene regulation. We present a comprehensive analysis of the conservation and clustering patterns of known miRNAs in human. We show that human miRNA gene clustering is significantly higher than expected at random. A total of 37% of the known human miRNA genes analyzed in this study appear in clusters of two or more with pairwise chromosomal distances of at most 3000 nt. Comparison of the miRNA sequences with their homologs in four other organisms reveals a typical conservation pattern, persistent throughout the clusters. Furthermore, we show enrichment in the typical conservation patterns and other miRNA-like properties in the vicinity of known miRNA genes, compared with random genomic regions. This may imply that additional, yet unknown, miRNAs reside in these regions, consistent with the current recognition that there are overlooked miRNAs. Indeed, by comparing our predictions with cloning results and with identified miRNA genes in other mammals, we corroborate the predictions of 18 additional human miRNA genes in the vicinity of the previously known ones. Our study raises the proportion of clustered human miRNAs that are <3000 nt apart to 42%. This suggests that the clustering of miRNA genes is higher than currently acknowledged, alluding to its evolutionary and functional implications.

Base Sequence↗

Engineering a rare-cutting restriction enzyme: genetic screening and selection of NotI variants.

Restriction endonucleases (REases) with 8-base specificity are rare specimens in nature. NotI from Nocardia otitidis-caviarum (recognition sequence 5'-GCGGCCGC-3') has been cloned, thus allowing for mutagenesis and screening for enzymes with altered 8-base recognition and cleavage activity. Variants possessing altered specificity have been isolated by the application of two genetic methods. In step 1, variant E156K was isolated by its ability to induce DNA-damage in an indicator strain expressing M.EagI (to protect 5'-NCGGCCGN-3' sites). In step 2, the E156K allele was mutagenized with the objective of increasing enzyme activity towards the alternative substrate site: 5'-GCTGCCGC-3'. In this procedure, clones of interest were selected by their ability to eliminate a conditionally toxic substrate vector and induce the SOS response. Thus, specific DNA cleavage was linked to cell survival. The secondary substitutions M91V, F157C and V348M were each found to have a positive effect on specific activity when paired with E156K. For example, variant M91V/E156K cleaves 5'-GCTGCCGC-3' with a specific activity of 8.2 x 10(4) U/mg, a 32-fold increase over variant E156K. A comprehensive analysis indicates that the cleavage specificity of M91V/E156K is relaxed to a small set of 8 bp substrates while retaining activity towards the NotI sequence.

Base Sequence↗

A reliable method to display authentic DNase I hypersensitive sites at long-ranges in single-copy genes from large genomes.

The study of eukaryotic gene transcription depends on methods to discover distal cis-acting control sequences. Comparative bioinformatics is one powerful strategy to reveal these domains, but still requires conventional wet-bench techniques to elucidate their specificity and function. The DNase I hypersensitivity assay (DHA) is also a method to identify regulatory domains, but can also suggest their function. Technically however, the classical DHA is constrained to mapping gene loci in small increments of approximately 20 kb. This limitation hinders efficient and comprehensive analysis of distal gene regions. Here, we report an improved method termed mega-DHA that extends the range of existing DHAs to facilitate assaying intervals that approach 100 kb. We demonstrate its feasibility for efficient analysis of single-copy genes within a large and complex genome by assaying 230 kb of the human ADAMTS14-perforin-paladin gene cluster in four experiments. The results identify distinct networks of regulatory domains specific to expression of perforin and its two neighboring genes.

Animals↗