Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Molecular genetic analyses of potential beta-galactosidase genes in Xanthomonas campestris.

Xanthomonas campestris pv. campestris, which displays no significant beta-1,4-D-galactopyranosidase activity, has three annotated beta-galactosidase genes in the sequenced genome, designated galA, galB and galC herein. GalA and GalB are similar to glycosyl hydrolase (GH) family 2 enzymes, including Escherichia coli LacZ. galA and galB cannot express detectable activity even after being cloned in-frame and driven by the vector's promoter. GalC is a GH35 enzyme homologous to the Xanthomonas axonopodis pv. manihotis Bga. The latter cleaves beta,1-3-linked galactose 1,000 times faster than beta,1-4-linked galactose and is not responsible for lactose utilization. In X. campestris pv. campestris cells, GalC is readily detectable by Western blotting, and the levels can be increased by cloning the gene under the control of the vector's promoter. Results of insertional mutation, transcriptional fusion assay and Western blotting indicated that galC, clustered with several GH genes, is cotranscribed with the upstream gene(s) and is expressed constitutively. Xc17L is a previously isolated mutant with elevated beta-galactosidase activity and a greatly improved ability to grow on lactose. Results of DNA sequencing of Xc17L galA, galB and galC, enzyme assays of galA, galB and galC mutants derived from Xc17L, and Western blotting of GalC in Xc17L indicated that the three beta-galactosidase genes do not encode the elevated beta-galactosidase activity in Xc17L. The presence of a fourth beta-galactosidase gene is proposed.

Amino Acid Motifs↗

FISH analysis of Drosophila melanogaster heterochromatin using BACs and P elements.

The heterochromatin of chromosomes 2 and 3 of Drosophila melanogaster contains about 30 essential genes defined by genetic analysis. In the last decade only a few of these genes have been molecularly characterized and found to correspond to protein-coding genes involved in important cellular functions. Moreover, several predicted genes have been identified by annotation of genomic sequence that are associated with polytene chromosome divisions 40, 41 and 80 but their locations on the cytogenetic map of the heterochromatin are still uncertain. To expand our current knowledge of the genetic functions located in heterochromatin, we have performed fluorescence in situ hybridization (FISH) mapping to mitotic chromosomes of nine bacterial artificial chromosomes (BACs) carrying several predicted genes and of 13 P element insertions assigned to the proximal regions of 2R and 3L. We found that 22 predicted genes map to the h46 region of 2R and eight map to the h47 regions of 3L. This amounts to at least 30 predicted genes located in these heterochromatic regions, whereas previous studies detected only seven vital genes. Finally, another 58 genes localize either in the euchromatin-heterochromatin transition regions or in the proximal euchromatin of 2R and 3L.

Animals↗

Quantitative molecular cartography of emergency myelopoiesis reveals conserved modules of hematopoietic activation.

Hematopoietic stem and progenitor cells (HSPCs) respond to infections, inflammation, and regenerative challenges using emergency myelopoiesis (EM) pathways to amplify myeloid cell production. However, it remains unclear how various EM inducers regulate HSPCs using shared or distinct molecular mechanisms. Here, we generate a comprehensive and generalizable cell annotation method (HemaScribe) and a refined quantitative model of hematopoietic differentiation (HemaScape) using single-cell RNA sequencing (scRNA-seq) of murine HSPCs, which we apply to a broad range of EM modalities. We uncover multiple strategies for enhancing myelopoiesis that act at different levels of the HSPC hierarchy and are associated with both unique and shared transcriptional response modules. In particular, we identify a myeloid progenitor-based EM activation module across diverse inflammatory challenges that is conserved in humans and informs outcomes in adult and pediatric acute myeloid leukemia. Our work illuminates fundamental regulatory mechanisms in hematopoietic regeneration that have direct translational applications in disease contexts.

Animals↗

Cis-regulatory variations: a study of SNPs around genes showing cis-linkage in segregating mouse populations.

BACKGROUND: Changes in gene expression are known to be responsible for phenotypic variation and susceptibility to diseases. Identification and annotation of the genomic sequence variants that cause gene expression changes is therefore likely to lead to a better understanding of the cause of disease at the molecular level. In this study we investigate the pattern of single nucleotide polymorphisms (SNPs) in genes for which the mRNA levels show cis-genetic linkage (gene expression quantitative trait loci mapping in cis, or cis-eQTLs) in segregating mouse populations. Such genes are expected to have polymorphisms near their physical location (cis-variations) that affect their mRNA levels by altering one or more of the cis-regulatory elements. This led us to characterize the SNPs in promoter (5 Kb upstream) and non-coding gene regions (introns and 5 Kb downstream) (cis-SNPs) and the effects they may have on putative transcription factor binding sites. RESULTS: We demonstrate that the cis-eQTL genes (CEGs) have a significantly higher frequency of cis-SNPs compared to non-CEGs (when both sets are taken from the non-IBD regions, i.e. regions not identical by descent). Most CEGs having cis-SNPs do not contain these SNPs in the phylogenetically conserved regions. In those CEGs that contain cis-SNPs in the phylogenetically conserved regions, enrichment of cis-SNPs occurs both within and outside of the conserved sequences. A higher fraction of CEGs are also seen to harbor cis-SNP that affect predicted transcription factor binding sites, a likely consequence of the higher cis-SNPs density in these genes. CONCLUSION: This present study provides the first genome-wide investigation of the putative cis-regulatory variations in a large set of genes whose levels of expression give rise to cis-linkage in segregating mammalian populations. Our results provide insights into the challenges that exist in identifying polymorphisms regulating gene expression using bioinformatic sequence analysis approaches. The data provided herein should benefit future investigations in this area.

Adipose Tissue↗

Protein molecular function prediction by Bayesian phylogenomics.

We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.

Adenosine Deaminase↗

Genomewide function conservation and phylogeny in the Herpesviridae.

The Herpesviridae are a large group of well-characterized double-stranded DNA viruses for which many complete genome sequences have been determined. We have extracted protein sequences from all predicted open reading frames of 19 herpesvirus genomes. Sequence comparison and protein sequence clustering methods have been used to construct herpesvirus protein homologous families. This resulted in 1692 proteins being clustered into 243 multiprotein families and 196 singleton proteins. Predicted functions were assigned to each homologous family based on genome annotation and published data and each family classified into seven broad functional groups. Phylogenetic profiles were constructed for each herpesvirus from the homologous protein families and used to determine conserved functions and genomewide phylogenetic trees. These trees agreed with molecular-sequence-derived trees and allowed greater insight into the phylogeny of ungulate and murine gammaherpesviruses.

Animals↗

Predicting ligand-binding function in families of bacterial receptors.

The three-dimensional fold of a new protein sequence can often be inferred directly from sequence homology to a protein of known structure. The function of a new protein sequence is more difficult to predict, however, since homologues can have different molecular and cellular functions. To develop and automate computational methods for determining molecular function, we have analyzed ligand-binding specificity in two related families of binding proteins. One of these families includes Escherichia coli lactose repressor and ribose-binding protein, and the other includes E. coli sulfate- and phosphate-binding proteins. These proteins have similar folds but varying specificity, binding many different small molecules, including mono- and disaccharides, purines, oxyanions, ferric iron, and polyamines. Starting from template structural alignments, alignments of over 90 sequences per family were generated by iterative database searches with hidden Markov models. Phylogenetic trees were made of full-length sequences and of subsets of residues lining the binding cleft, to determine whether subbranches of the trees correlate with ligand-binding preference. Automated analyses of residues in the binding pocket were also used to predict ligand-binding function for many uncharacterized database sequences and to identify specific side chain-ligand contacts in proteins without solved structures. Our results demonstrate the utility of anchoring functional annotation within a protein family context.

Amino Acid Sequence↗

A scan for positively selected genes in the genomes of humans and chimpanzees.

Since the divergence of humans and chimpanzees about 5 million years ago, these species have undergone a remarkable evolution with drastic divergence in anatomy and cognitive abilities. At the molecular level, despite the small overall magnitude of DNA sequence divergence, we might expect such evolutionary changes to leave a noticeable signature throughout the genome. We here compare 13,731 annotated genes from humans to their chimpanzee orthologs to identify genes that show evidence of positive selection. Many of the genes that present a signature of positive selection tend to be involved in sensory perception or immune defenses. However, the group of genes that show the strongest evidence for positive selection also includes a surprising number of genes involved in tumor suppression and apoptosis, and of genes involved in spermatogenesis. We hypothesize that positive selection in some of these genes may be driven by genomic conflict due to apoptosis during spermatogenesis. Genes with maximal expression in the brain show little or no evidence for positive selection, while genes with maximal expression in the testis tend to be enriched with positively selected genes. Genes on the X chromosome also tend to show an elevated tendency for positive selection. We also present polymorphism data from 20 Caucasian Americans and 19 African Americans for the 50 annotated genes showing the strongest evidence for positive selection. The polymorphism analysis further supports the presence of positive selection in these genes by showing an excess of high-frequency derived nonsynonymous mutations.

Animals↗

TransportDB: a comprehensive database resource for cytoplasmic membrane transport systems and outer membrane channels.

TransportDB (http://www.membranetransport.org/) is a comprehensive database resource of information on cytoplasmic membrane transporters and outer membrane channels in organisms whose complete genome sequences are available. The complete set of membrane transport systems and outer membrane channels of each organism are annotated based on a series of experimental and bioinformatic evidence and classified into different types and families according to their mode of transport, bioenergetics, molecular phylogeny and substrate specificities. User-friendly web interfaces are designed for easy access, query and download of the data. Features of the TransportDB website include text-based and BLAST search tools against known transporter and outer membrane channel proteins; comparison of transporter and outer membrane channel contents from different organisms; known 3D structures of transporters, and phylogenetic trees of transporter families. On individual protein pages, users can find detailed functional annotation, supporting bioinformatic evidence, protein/DNA sequences, publications and cross-referenced external online resource links. TransportDB has now been in existence for over 10 years and continues to be regularly updated with new evidence and data from newly sequenced genomes, as well as having new features added periodically.

Bacterial Outer Membrane Proteins↗

Pfam: a comprehensive database of protein domain families based on seed alignments.

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Amino Acid Sequence↗

Mitochondrial voltage-dependent anion channel gene family in Drosophila melanogaster: complex patterns of evolution, genomic organization, and developmental expression.

Voltage-dependent anion channels (VDACs), also known as mitochondrial porins, are a family of small pore-forming proteins of the mitochondrial outer membrane found in all eukaryotes. VDACs play important roles in the regulated flux of metabolites between the cytosolic and mitochondrial compartments, energy metabolism, and apoptosis. Annotation of the genome sequence of Drosophila melanogaster revealed three genes (CG17137, CG31722-A, and CG31722-B) with homology to porin, the previously described Drosophila VDAC. Molecular analysis reveals a complex pattern of organization and expression. The genomic organization of these four genes and sequence comparisons with other insect VDAC homologs indicate that this gene family evolved through a mechanism of duplication and divergence from an ancestral VDAC gene during the radiation of the genus Drosophila. CG17137, CG31722-A, and CG31722-B are expressed in a male-specific pattern on both transcriptional and translational levels, while porin is equally expressed in both male and female flies. Additionally, CG31722-A and CG31722-B are expressed as a dicistronic transcript. Western blot analysis and immunofluorescence microscopy confirm that these proteins localize to the mitochondrion. Further expression analysis showed that CG17137 and CG31722-B are abundant in testes, while porin is ubiquitously expressed. While porin, CG17137, and CG31722-B are expressed to different degrees during embryogenesis, all of these proteins are dramatically reduced relative to cytochrome c content during larvogenesis. These studies illustrate a complex genomic organization and spatiotemporal pattern of expression for Drosophila VDACs as well as an evolutionary history consistent with either a partitioning of VDAC functions or an acquisition of novel functions among isoforms.

Amino Acid Sequence↗

Go molecular function terms are predictive of subcellular localization.

A protein's function is closely linked to its subcellular localization. Use of Gene Ontology (GO) molecular function terms to extend sequence-based subcellular localization prediction has been previously shown to improve predictive performance. Here, we explore directly the relationship between GO function annotations and localization information, identifying both highly predictive single terms, and terms with large information gain with respect to location. The results identify a number of predictive and informative GO terms with respect to subcellular location, particularly nucleus, extracellular space, membrane, mitochondrion, endoplasmic reticulum and Golgi. There are several clear examples illustrating why the addition of function information provides additional predictive power over sequence alone. Other interesting phenomena can also be seen in the results. Most predictive or informative terms are imperfect, and incorrect prediction may often call out significant biological phenomena. Finally, these results may be useful in the GO annotation process.

Amino Acid Sequence↗

New knowledge from old: in silico discovery of novel protein domains in Streptomyces coelicolor.

BACKGROUND: Streptomyces coelicolor has long been considered a remarkable bacterium with a complex life-cycle, ubiquitous environmental distribution, linear chromosomes and plasmids, and a huge range of pharmaceutically useful secondary metabolites. Completion of the genome sequence demonstrated that this diversity carried through to the genetic level, with over 7000 genes identified. We sought to expand our understanding of this organism at the molecular level through identification and annotation of novel protein domains. Protein domains are the evolutionary conserved units from which proteins are formed. RESULTS: Two automated methods were employed to rapidly generate an optimised set of targets, which were subsequently analysed manually. A final set of 37 domains or structural repeats, represented 204 times in the genome, was developed. Using these families enabled us to correlate items of information from many different resources. Several immediately enhance our understanding both of S. coelicolor and also general bacterial molecular mechanisms, including cell wall biosynthesis regulation and streptomycete telomere maintenance. DISCUSSION: Delineation of protein domain families enables detailed analysis of protein function, as well as identification of likely regions or residues of particular interest. Hence this kind of prior approach can increase the rate of discovery in the laboratory. Furthermore we demonstrate that using this type of in silico method it is possible to fairly rapidly generate new biological information from previously uncorrelated data.

Amino Acid Motifs↗

Systematic gene targeting on the X chromosome of Drosophila melanogaster.

The genome of the model organism Drosophila melanogaster has been sequenced and annotated. Based on this groundwork, we performed a systematic genetic screen of the D. melanogaster X chromosome, which carries about one sixth of the genes of the organism. We generated a collection of single P-element insertions to provide genetic and molecular access to virtually all X-chromosomal genes. The study complements earlier work designed to systematically identify vital genes on the X chromosome by targeting transcription units which are phenotypically silent. We describe single UAS sequence-bearing P-element insertions throughout the X chromosome, which allows one to express the tagged genes under control of tissue/organ-directed GAL4 activity. In addition, the present collection of single insertion lines provides a tool to generate chromosomal deletions which are on average less than 33 kb in size.

Animals↗

Overproduction and characterization of recombinant UDP-glucose pyrophosphorylase from Escherichia coli K-12.

Using oligonucleotide probes synthesized on the basis of partial amino acid sequences, we have cloned and sequenced the gene of Escherichia coli K-12 encoding UDP-glucose pyrophosphorylase. The gene consists of 906 base pairs and encodes a polypeptide of 302 amino acid residues with a calculated molecular weight of 32,941. Its nucleotide sequence was found to be identical with that recently registered (EMBL, X59940) for a gene coding for an unknown 33-kDa protein, which was later annotated as UDP-glucose pyrophosphorylase on the basis of genetic studies. The UDP-glucose pyrophosphorylase gene, mapped at 27.3 min in the E. coli chromosome, complemented the galU mutation, which renders the bacterium unable to ferment galactose. The recombinant enzyme overproduced in E. coli cells and purified to homogeneity catalyzed the synthesis and pyrophosphorolysis of UDP-glucose by a sequential mechanism. The enzyme required Mg2+ for maximal activity and was inhibited by free UTP and pyrophosphate. The E. coli enzyme shows significant sequence similarities with the enzymes from Acetobacter xylinum and Salmonella typhimurium. However, little or no similarity was found with the eukaryotic enzymes that are involved in the biosynthesis of storage carbohydrates, or with other enzymes acting on similar sugar nucleotides. Thus, UDP-glucose pyrophosphorylases participating in diverse metabolic pathways can be classified structurally into the prokaryotic and eukaryotic groups, even though they have almost identical catalytic properties.

Amino Acid Sequence↗

Application of physiological genomics to the study of hearing disorders.

Although the biophysical principles of how the ear operates are reasonably well understood, little is known about the specific genes that confer normal function to the inner ear. Nevertheless, the recent implementation of genomic tools has led to extraordinary progress in the identification of mutated genes that cause non-syndromic and syndromic forms of deafness. Part of this success is directly related to the sequencing of the human and mouse genomes and improved gene annotation methods. This review discusses how physiological genomic tools, such as genomic databases, expressed sequence tag databases and DNA arrays have been applied to find candidate genes for important molecular processes in the inner ear. It also illustrates, using the discovery of genes encoding essential components of cochlear K+ homeostasis as an example, how the combination of physiological genomic tools with physiological and morphological information has led to an in-depth understanding of cochlear ion homeostasis. Finally, it discusses how the use of applied genomic tools, such as gene arrays, will further advance our knowledge of how the inner ear works, develops, ages and regenerates.

Cochlea↗

A mutation in the receptor binding site of GDF5 causes Mohr-Wriedt brachydactyly type A2.

BACKGROUND: Brachydactyly type A2 (OMIM 112600) is characterised by hypoplasia/aplasia of the second middle phalanx of the index finger and sometimes the little finger. BDA2 was first described by Mohr and Wriedt in a large Danish/Norwegian kindred and mutations in BMPR1B were recently demonstrated in two affected families. METHODS: We found and reviewed Mohr and Wriedt's original unpublished annotations, updated the family pedigree, and examined 37 family members clinically, and radiologically by constructing the metacarpo-phalangeal profile (MCPP) pattern in nine affected subjects. Molecular analyses included sequencing of BMPR1B, linkage analysis for STS markers flanking GDF5, sequencing of GDF5, confirmation of the mutation by a restriction enzyme assay, and localisation of the mutation inferred from the very recently reported GDF5 crystal structure, and by superimposing the GDF5 protein sequence onto the crystal structure of BMP2 bound to Bmpr1a. RESULTS: A short middle phalanx of the index finger was found in all affected individuals, but other fingers were occasionally involved. The fourth finger was characteristically spared. This distinguishes Mohr-Wriedt type BDA2 from BDA2 caused by mutations in BMPR1B. An MCPP analysis most efficiently detected mutation carrier status. We identified a missense mutation, c.1322T>C, causing substitution of a leucine with a proline at amino acid residue 441 within the active signalling domain of GDF5. The mutation was predicted to reside in the binding site for BMP type 1 receptors. CONCLUSION: GDF5 is a novel BDA2 causing gene. It is suggested that impaired activity of BMPR1B is the molecular mechanism responsible for the BDA2 phenotype.

Binding Sites↗

The Complete Chloroplast Genome and the Phylogenetic Analysis of Panicum bisulcatum (Thumb.) (Poaceae).

The chloroplast (cp) genome of Panicum bisulcatum (Thumb.), a significant agricultural weed, was sequenced and characterized to elucidate its genomic architecture, evolutionary dynamics, and phylogenetic relationships. The complete cp genome was assembled as a circular DNA molecule of 138,489 bp, exhibiting a typical quadripartite structure comprising a large single-copy (LSC, 82,260 bp), a small single-copy (SSC, 12,569 bp), and a pair of inverted repeats (IR, 21,830 bp each) regions. It encodes 135 genes, including 89 protein-coding genes, 49 tRNAs, and 8 rRNAs. Functional annotation revealed that most genes are involved in photosynthesis and genetic system. A total of 51 simple sequence repeats (SSRs) and 62 long repeats (LRs) were identified, providing potential molecular markers. Comparative analysis of IR boundaries highlighted both conserved features and species-specific expansion/contraction events among Panicum species. Phylogenomic analysis robustly placed P. bisulcatum within the genus Panicum, showing a closest relationship with P. incomtum and confirming the monophyly of the genus. Furthermore, single nucleotide polymorphism (SNP) analysis with its closest relative, P. incomtum, revealed 4659 SNPs, with a dominance of synonymous substitutions, indicating the action of purifying selection. This study provides the first comprehensive cp genomic resource for P. bisulcatum, which will facilitate future studies in species identification, phylogenetic reconstruction, population genetics, and the development of sustainable management strategies for this weed.

Phylogeny↗