Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein↗

The Mouse Genome Database (MGD): expanding genetic and genomic resources for the laboratory mouse. The Mouse Genome Database Group.

The Mouse Genome Database (MGD) is a comprehensive public database of mouse genomic, genetic and phenotypic information (http://www. informatics.jax.org). This community database provides information about genes, serves as a mapping resource of the mouse genome, details mammalian orthologs, integrates experimental data, represents standardized mouse nomenclature for genes and alleles, incorporates links to other genomic resources such as sequence data, and includes a variety of additional information about the laboratory mouse. MGD scientists and annotators work cooperatively with the research community to provide an integrated, consensus view of the mouse genome while also providing experimental data including data conflicting with the consensus representation. Recent improvements focus on the representation of phenotypic information and the enhancement of gene and allele descriptions.

Animals↗

The Zebrafish Information Network (ZFIN): a resource for genetic, genomic and developmental research.

The Zebrafish Information Network, ZFIN, is a WWW community resource of zebrafish genetic, genomic and developmental research information (http://zfin.org). ZFIN provides an anatomical atlas and dictionary, developmental staging criteria, research methods, pathology information and a link to the ZFIN relational database (http://zfin. org/ZFIN/). The database, built on a relational, object-oriented model, provides integrated information about mutants, genes, genetic markers, mapping panels, publications and contact information for the zebrafish research community. The database is populated with curated published data, user submitted data and large dataset uploads. A broad range of data types including text, images, graphical representations and genetic maps supports the data. ZFIN incorporates links to other genomic resources that provide sequence and ortholog data. Zebrafish nomenclature guidelines and an automated registration mechanism for new names are provided. Extensive usability testing has resulted in an easy to learn and use forms interface with complex searching capabilities.

Animals↗

UniProt: the Universal Protein knowledgebase.

To provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information, the Swiss-Prot, TrEMBL and PIR protein database activities have united to form the Universal Protein Knowledgebase (UniProt) consortium. Our mission is to provide a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references and query interfaces. The central database will have two sections, corresponding to the familiar Swiss-Prot (fully manually curated entries) and TrEMBL (enriched with automated classification, annotation and extensive cross-references). For convenient sequence searches, UniProt also provides several non-redundant sequence databases. The UniProt NREF (UniRef) databases provide representative subsets of the knowledgebase suitable for efficient searching. The comprehensive UniProt Archive (UniParc) is updated daily from many public source databases. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). The scientific community is encouraged to submit data for inclusion in UniProt.

Animals↗

The Universal Protein Resource (UniProt).

The Universal Protein Resource (UniProt) provides the scientific community with a single, centralized, authoritative resource for protein sequences and functional information. Formed by uniting the Swiss-Prot, TrEMBL and PIR protein database activities, the UniProt consortium produces three layers of protein sequence databases: the UniProt Archive (UniParc), the UniProt Knowledgebase (UniProt) and the UniProt Reference (UniRef) databases. The UniProt Knowledgebase is a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase with extensive cross-references. This centrepiece consists of two sections: UniProt/Swiss-Prot, with fully, manually curated entries; and UniProt/TrEMBL, enriched with automated classification and annotation. During 2004, tens of thousands of Knowledgebase records got manually annotated or updated; we introduced a new comment line topic: TOXIC DOSE to store information on the acute toxicity of a toxin; the UniProt keyword list got augmented by additional keywords; we improved the documentation of the keywords and are continuously overhauling and standardizing the annotation of post-translational modifications. Furthermore, we introduced a new documentation file of the strains and their synonyms. Many new database cross-references were introduced and we started to make use of Digital Object Identifiers. We also achieved in collaboration with the Macromolecular Structure Database group at EBI an improved integration with structural databases by residue level mapping of sequences from the Protein Data Bank entries onto corresponding UniProt entries. For convenient sequence searches we provide the UniRef non-redundant sequence databases. The comprehensive UniParc database stores the complete body of publicly available protein sequence data. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). New releases are published every two weeks.

Amino Acid Sequence↗

The Universal Protein Resource (UniProt): an expanding universe of protein information.

The Universal Protein Resource (UniProt) provides a central resource on protein sequences and functional annotation with three database components, each addressing a key need in protein bioinformatics. The UniProt Knowledgebase (UniProtKB), comprising the manually annotated UniProtKB/Swiss-Prot section and the automatically annotated UniProtKB/TrEMBL section, is the preeminent storehouse of protein annotation. The extensive cross-references, functional and feature annotations and literature-based evidence attribution enable scientists to analyse proteins and query across databases. The UniProt Reference Clusters (UniRef) speed similarity searches via sequence space compression by merging sequences that are 100% (UniRef100), 90% (UniRef90) or 50% (UniRef50) identical. Finally, the UniProt Archive (UniParc) stores all publicly available protein sequences, containing the history of sequence data with links to the source databases. UniProt databases continue to grow in size and in availability of information. Recent and upcoming changes to database contents, formats, controlled vocabularies and services are described. New download availability includes all major releases of UniProtKB, sequence collections by taxonomic division and complete proteomes. A bibliography mapping service has been added, and an ID mapping service will be available soon. UniProt databases can be accessed online at http://www.uniprot.org or downloaded at ftp://ftp.uniprot.org/pub/databases/.

Databases, Protein↗

Genetic variability at the human FMO1 locus: significance of a basal promoter yin yang 1 element polymorphism (FMO1*6).

The flavin-containing monooxygenases (FMOs) are important for the disposition of a variety of toxicants, therapeutics, and dietary components. Although FMO1 is the dominant isoform in fetal liver and adult kidney and intestine and despite up to a 10-fold intersubject variation in expression, a paucity of information is available on FMO1 genetic variability. To address this issue, 24 samples from the Coriell DNA Polymorphism Discovery Resource Panel were sequenced revealing 10 common single nucleotide polymorphisms (SNPs): four located upstream of the structural gene; three within exonic sequences; one within the intron 1 splice donor site; and two with the 3'-untranslated region. Six of these variants are novel. Compared with other FMO loci within the chromosome 1q23-25 cluster, FMO1 seems more highly conserved. Of the identified FMO1 SNPs, only a C>A transversion 9536 base pairs upstream of the exon 2 ATG start codon (g.-9536C>A) would likely affect function, because it lies within the conserved core binding sequence for the yin yang 1 (YY1) transcription factor. Electrophoretic mobility shift assays demonstrated that the g.-9536C>A transversion eliminated YY1 binding. Furthermore, data from transient expression assays in HepG2 cells suggested this SNP could account for a 2- to 3-fold loss of FMO1 promoter activity. Genotype analysis revealed a g.-9,536A allele (FMO1*6) frequency of 13 and 11% in African- and northern European-Americans, respectively, but a significantly higher frequency of 30% in Hispanic-Americans. Thus, the FMO1*6 variant may account for some of the observed interindividual variation in FMO1 expression.

Base Sequence↗

Bioinformatics tools for whole genomes.

The advent of whole-genome data resources--not only sequence but also other genome-scale data collections such as gene expression, protein interaction, and genetic variation--is having two marked, complementary effects on the relatively new discipline of bioinformatics. First, the veritable flood of data is creating a need and demand for new tools for dealing adequately with the deluge, and, second, the unprecedented extent, diversity, and impending completeness of the data sets are creating opportunities for new approaches to discovery based on computational methods.

Algorithms↗

Alternate promoter and 5'-untranslated exon usage of the mouse adrenocorticotropin receptor gene in adipose tissue.

Mouse adrenocorticotropin receptor (ACTH-R/MC2R) messenger ribonucleic acid (mRNA) is expressed predominantly in the adrenal gland and, to a lesser extent, in adipose tissue. In this study, we found a novel 135-bp exon 1 (exon 1f) of the ACTH-R gene transcribed in mouse adipose tissue by RNA ligase-mediated rapid amplification of cDNA ends, which was located 1.4 kb downstream in the genome of previously-reported exon 1 (exon 1a) transcribed in the adrenal gland. The novel promoter region, 1.4 kb upstream of exon 1f contained three CCAAT boxes. RT-PCR analysis revealed that ACTH-R mRNA from adipose tissue and differentiated 3T3-L1 adipocytes exclusively contained exon 1f. Thus, the promoter region flanking to exon 1f is thought to be essential for adipose tissue, while that flanking to exon 1a is specific for the adrenal gland. A search for a similar sequence of mouse ACTH-R exon 1f and its flanking region in the human genome sequence database of GenBank Human Genome Resources did not reveal such a sequence in the region of the human ACTH-R gene. This may explain the absence of ACTH-R expression in human adipose tissue.

5' Untranslated Regions↗

A YAC contig spanning the blepharophimosis-ptosis-epicanthus inversus syndrome and propionic acidemia loci.

Blepharophimosis-ptosis-epicanthus inversus syndrome (BPES) is an autosomal dominant condition consisting of congenital dysplasia of the eyelids with a reduced horizontal diameter of the palpebral fissures, droopy eyelids and epicanthus inversus. Two clinical entities have been described: type I and type II. The former is distinguished by female infertility, whereas the latter presents without other symptoms. Both type I and type II were recently mapped on the long arm of chromosome 3 (3q22-q23), suggesting a common gene may be affected. The centromeric and the telomeric limits of this region are well defined between loci D3S1316 and D3S1615, which reside approximately 5 cM apart. Here, we present the construction of a YAC contig spanning the entire BPES locus using 17 polymorphic markers, 2 STS and 28 ESTs. This region of approximately 5 Mb was covered by 31 YACs, and was supported by detailed FISH analysis. In addition, we have precisely mapped the propionyl-CoA carboxylase beta polypeptide (PCCB), the gene mutated in propionic acidemia, within this contig. Apart from providing a framework for the identification of the BPES gene, this contig will also be useful for the future identification of defects and genes mapped to this region, and for developing template resources for genomic sequencing.

Amino Acid Metabolism, Inborn Errors↗

Enhancement of achievement and attitudes through individualized learning-style presentations of two allied health courses.

This investigation analyzed the effects of the instructional resource Programmed Learning Sequence (PLS) on the achievement and attitudes of college students and correlated the findings with the individuals' learning styles. The subjects were enrolled in Sonography I and Cross-Sectional Anatomy in a college of health-related professions. Both classes were administered the Productivity Environmental Preference Survey to identify learning-style strengths, and alternately presented with lessons using a PLS in a book format and traditional lectures. The sonography class also was exposed to a PLS in multimedia computer format. The Semantic Differential Scale measured the students' attitudes comparing the instructional methods experienced, and class examinations measured content mastery. In both classes, examination scores were significantly higher (effect size for the sonography class was 1.42; for the anatomy class, 0.63) and students' attitude scores were significantly higher when PLS rather than the traditional method was used. In the sonography class, achievement was significantly higher with the book PLS than with the computer PLS (effect size, 1.11). Significant correlations emerged between learning-style elements and achievement: students who preferred learning with the book PLS required more quiet in the environment than did those who preferred the computer PLS; students who preferred learning traditionally and with the computer PLS required more light than those preferring the book PLS; and students who preferred learning with an authority figure favored the traditional method. Examination of the data for other correlations between learning-style preferences and attitudes using the book PLS also revealed many other significant findings, demonstrating its ability to accommodate diverse styles.

Adolescent↗

Seasonal branch nutrient dynamics in two Mediterranean woody shrubs with contrasted phenology.

BACKGROUND AND AIMS: Mediterranean woody plants have a wide variety of phenological strategies. Some authors have classified the Mediterranean phanaerophytes into two broad phenological categories: phenophase-overlappers (that overlap resource-demanding activities in a short period of the year) and phenophase-sequencers (that protract resource-demanding activities throughout the year). In this work the impact of both phenological strategies on leaf nutrient accumulation and retranslocation dynamics at the level of leaves and branches was evaluated. Phenophase-overlappers were expected to accumulate nutrients in leaves throughout most of the year and withdraw them efficiently in a short period. Phenophase-sequencers were expected to withdraw nutrients progressively throughout the year, without long accumulation periods. METHODS: To test this hypothesis, variations in phenology and leaf NPK in the crown of a phenophase-overlapper Cistus laurifolius and a phenophase-sequencer Bupleurum fruticosum were monitored monthly during 2 years. KEY RESULTS: Changes in nutrient concentration at the leaf level were not clearly related with the different phenologies. Nitrogen and phosphorous resorption efficiencies were lower in the phenophase-overlapper, and accumulation-retranslocation seasonality was similar in both species. Changes in the branch nutrient pool agreed with the hypothesis that the phenophase-overlapper accumulated nutrients from summer until the bud burst of the following spring, recovering a large nutrient pool during massive leaf shedding. The phenophase-sequencer did not accumulate nutrients from autumn until early spring, achieving lower nutrient recovery during spring leaf shedding. CONCLUSIONS: It is concluded that phenological demands influence branch nutrient cycling. This effect is easier to detect by assessing changes in the branch nutrient pool rather than changes in the leaf nutrient concentration.

Bupleurum↗

IMGT databases, web resources and tools for immunoglobulin and T cell receptor sequence analysis, http://imgt.cines.fr.

IMGT, the international ImMunoGeneTics database((R)) (http://imgt.cines.fr), is a high-quality integrated information system specializing in immunoglobulins (IG), T cell receptors (TR) and major histocompatibility complex (MHC) of human and other vertebrates, created in 1989, by LIGM, at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes several databases (IMGT/LIGM-DB, IMGT/3Dstructure-DB, IMGT/HLA-DB), Web resources ('IMGT Marie-Paule page') and interactive tools (IMGT/V-QUEST, IMGT/JunctionAnalysis). IMGT expertly annotated data and tools described in this paper are particularly useful for the analysis of the IG and TR rearrangements in leukemia, lymphoma and myeloma, and in translocations involving the antigen receptor loci. IMGT is freely available at http://imgt.cines.fr.

Amino Acid Sequence↗

Chicken genomics charts a path to the genome sequence.

In this paper, the current status of chicken genomics is reviewed. This is timely given the current intense activity centred on sequencing the complete genome of this model species. The genome project is based on a decade of map building by genetic linkage and cytogenetic methods, which are now being replaced by high-resolution radiation hybrid and bacterial artificial chromosome (BAC) contig maps. Markers for map building have generally depended on labour-intensive screening procedures, but in recent years this has changed with the availability of almost 500,000 chicken expressed sequence tags (ESTs). These resources and tools will be critical in the coming months when the chicken genome sequence is being assembled (eg cross-checked with other maps) and annotated (eg gene structures based on ESTs). The future for chicken genome and biological research is an exciting one, through the integration of these resources. For example, through the proposed chicken Ensembl database, it will be possible to solve challenging scientific questions by exploiting the power of a chicken model. One area of interest is the study of developmental mechanisms and the discovery of regulatory networks throughout the genome. Another is the study of the molecular nature of quantitative genetic variation. No other animal species have been phenotyped and selected so intensively as agricultural animals and thus there is much to be learned in basic and medical biology from this research.

Animals↗

BN phenome: detailed characterization of the cardiovascular, renal, and pulmonary systems of the sequenced rat.

The postgenome era has provided resources to link disease phenotypes to the genomic sequence, i.e., creating a disease "phenome." Our detailed characterization of the sequenced BN rat strain (BN/NHsdMcwi) provides the first concerted effort in creating a direct link between a sequenced genome and its resulting biology. For the BN sequence to be of broad value to investigators, these measures need to be put into the context of the spectrum of the laboratory rats, so that their physiology can be benchmarked against the sequenced BN. As a major step in generating a comprehensive cardiovascular and pulmonary disease phenome, we measured 281 traits related to diseases of the heart, lung, and blood (http://pga.mcw.edu) in the sequenced BN. We compared these data with those of the same traits measured across multiple genetic backgrounds, both genders, and differing environments. We show that no single strain, inbred or outbred, can be considered a physiological control strain; what is normal depends on what trait is being measured and the strains' genome backgrounds. We find vast differences between the genders, also dependent on genome background. By combining the values across all strains studied, we generated a "population" mean and normal range of values for each of these traits, which are more genetically representative than the measured values in any single inbred or outbred strain. These data provide a baseline for physiological comparison of traits related to cardiovascular, lung, blood, and renal function in the sequenced BN rats relative to the major strains of rats studied in biomedical research.

Animals↗

Partial amino acid sequence determination of bovine corneal protein 54 K (BCP 54)

The most abundant soluble protein of bovine cornea, BCP 54 (Bovine Corneal Protein, molecular weight 54 kD) was isolated and digested under both limited and complete digestion conditions with Staphylococcus aureus V8 protease. The fragments resulting from limited digestion were separated by one-dimensional sodium dodecyl sulfate polyacrylamide gel electrophoresis, transferred to a polyvinylidene difluoride membrane, visualized by Coomassie Blue staining, cut out, and submitted to N-terminal protein sequence analysis. Complete digestion fragments were separated on a Vydac C18 reverse-phase HPLC, collected, and concentrated prior to sequencing. Using this method, we obtained amino acid sequence data from three internal V8 protease derived fragments of BCP 54 and a number of HPLC fragments. Comparison of these amino acid sequences, corresponding to 30% of the BCP 54 molecule, to those sequences contained within release 22 of the National Biomedical Research Foundation Protein Identification Resource revealed no extended sequence similarity of known proteins to BCP54.

Aldehyde Dehydrogenase↗

Gene3D: structural assignment for whole genes and genomes using the CATH domain structure database.

We present a novel web-based resource, Gene3D, of precalculated structural assignments to gene sequences and whole genomes. This resource assigns structural domains from the CATH database to whole genes and links these to their curated functional and structural annotations within the CATH domain structure database, the functional Dictionary of Homologous Superfamilies (DHS) and PDBsum. Currently Gene3D provides annotation for 36 complete genomes (two eukaryotes, six archaea, and 28 bacteria). On average, between 30% and 40% of the genes of a given genome can be structurally annotated. Matches to structural domains are found using the profile-based method (PSI-BLAST). and a novel protocol, DRange, is used to resolve conflicts in matches involving different homologous superfamilies.

Animals↗

ASTRAL compendium enhancements.

The ASTRAL compendium provides several databases and tools to aid in the analysis of protein structures, particularly through the use of their sequences. It is partially derived from the SCOP database of protein domains, and it includes sequences for each domain as well as other resources useful for studying these sequences and domain structures. Several major improvements have been made to the ASTRAL compendium since its initial release 2 years ago. The number of protein domain sequences included has doubled from 15 190 to 30 867, and additional databases have been added. The Rapid Access Format (RAF) database contains manually curated mappings linking the biological amino acid sequences described in the SEQRES records of PDB entries to the amino acid sequences structurally observed (provided in the ATOM records) in a format designed for rapid access by automated tools. This information is used to derive sequences for protein domains in the SCOP database. In cases where a SCOP domain spans several protein chains, all of which can be traced back to a single genetic source, a 'genetic domain' sequence is created by concatenating the sequences of each chain in the order found in the original gene sequence. Both the original-style library of SCOP sequences and a new library including genetic domain sequences are available. Selected representative subsets of each of these libraries, based on multiple criteria and degrees of similarity, are also included. ASTRAL may be accessed at http://astral.stanford.edu/.

Amino Acid Sequence↗