Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

An accurate, residue-level, pair potential of mean force for folding and binding based on the distance-scaled, ideal-gas reference state.

Structure prediction on a genomic scale requires a simplified energy function that can efficiently sample the conformational space of polypeptide chains. A good energy function at minimum should discriminate native structures against decoys. Here, we show that a recently developed, residue-specific, all-atom knowledge-based potential (167 atomic types) based on distance-scaled, finite ideal-gas reference state (DFIRE-all-atom) can be substantially simplified to 20 residue types located at side-chain center of mass (DFIRE-SCM) without a significant change in its capability of structure discrimination. Using 96 standard multiple decoy sets, we show that there is only a small reduction (from 80% to 78%) in success rate of ranking native structures as the top 1. The success rate is higher than two previously developed, all-atom distance-dependent statistical pair potentials. Applied to structure selections of 21 docking decoys without modification, the DFIRE-SCM potential is 29% more successful in recognizing native complex structures than an all-atom statistical potential trained by a database of dimeric interfaces. The potential also achieves 92% accuracy in distinguishing true dimeric interfaces from artificial crystal interfaces. In addition, the DFIRE potential with the C(alpha) positions as the interaction centers recognizes 123 native structures out of a comprehensive 125-protein TOUCHSTONE decoy set in which each protein has 24,000 decoys with only C(alpha) positions. Furthermore, the performance by DFIRE-SCM on newly established 25 monomeric and 31 docking Rosetta-decoy sets is comparable to (or better than in the case of monomeric decoy sets) that of a recently developed, all-atom Rosetta energy function enhanced with an orientation-dependent hydrogen bonding potential.

Amino Acids↗

A homozygous diploid subset of commercial wine yeast strains.

Genetic analysis was performed on 45 commercial yeasts which are used in winemaking because of their superior fermentation properties. Genome sizes were estimated by propidium iodide fluorescence and flow cytometry. Forty strains had genome sizes consistent with their being diploid, while five had a range of aneuploid genome sizes that ranged from 1.2 to 1.8 times larger. The diploid strains are all Saccharomyces cerevisiae, based on genetic analysis of microsatellite and minisatellite markers and on DNA sequence analysis of the internal transcribed spacer (ITS) region of nuclear ribosomal DNA of four strains. Four of the five aneuploid strains appeared to be interspecific hybrids between Saccharomyces kudriavzevii and Saccharomyces cerevisiae, with the fifth a hybrid between two S. cerevisiae strains. An identification fingerprint was constructed for the commercial yeast strains using 17 molecular markers. These included six published trinucleotide microsatellites, seven new dinucleotide microsatellites, and four published minisatellite markers. The markers provided unambiguous identification of the majority of strains; however, several had identical or similar patterns, and likely represent the same strain or mutants derived from it. The combined use of all 17 polymorphic loci allowed us to identify a set of eleven commercial wine yeast strains that appear to be genetically homozygous. These strains are presumed to have undergone inbreeding to maintain their homozygosity, a process referred to previously as 'genome renewal'.

Base Sequence↗

Genome conservation between the bovine and human interleukin-8 receptor complex: improper annotation of bovine interleukin-8 receptor b identified.

Interleukin (IL)-8 and its receptors, CXCR1 and CXCR2, are key regulators of inflammation. However, knowledge of these receptors at the genomic level is limiting or absent in cattle. Therefore, our objective was to identify bovine orthologs of human CXCR1 and CXCR2. Alignment of bovine CXCR2 reference mRNA to the bovine genome revealed two regions of similarity on BTA2 approximately 20 kb apart and on opposite strands. Comparison with the human genome suggested the more centromeric region to be CXCR2 and the more telomeric region to be CXCR1 which contradicts the current annotation of the bovine CXCR2 reference mRNA. This observation was verified by sequencing RT-PCR products of specific regions within each predicted IL-8 receptor and comparing with human sequences using ClustalW. Further examination of coding and non-coding regions within the IL-8 receptor genome complex revealed that both bovine and canine CXCR1 and CXCR2 genes had more conserved sequences in common with the human genes than either mouse or rat, and may offer more suitable animal models for certain applications. This molecular information provides a stepping stone for greater understanding of the role each IL-8 receptor plays in inflammation and will enhance our ability to develop strategies against inflammatory based diseases.

Animals↗

Detection of single nucleotide polymorphisms.

Single nucleotide polymorphism (SNP) detection technologies are used to scan for new polymorphisms and to determine the allele(s) of a known polymorphism in target sequences. SNP detection technologies have evolved from labor intensive, time consuming, and expensive processes to some of the most highly automated, efficient, and relatively inexpensive methods. Driven by the Human Genome Project, these technologies are now maturing and robust strategies are found in both SNP discovery and genotyping areas. The nearly completed human genome sequence provides the reference against which all other sequencing data can be compared. Global SNP discovery is therefore only limited by the amount of funding available for the activity. Local, target, SNP discovery relies mostly on direct DNA sequencing or on denaturing high performance liquid chromatography (dHPLC). The number of SNP genotyping methods has exploded in recent years and many robust methods are currently available. The demand for SNP genotyping is great, however, and no one method is able to meet the needs of all studies using SNPs. Despite the considerable gains over the last decade, new approaches must be developed to lower the cost and increase the speed of SNP detection.

Alleles↗

Genetic reclassification of porcine enteroviruses.

The genetic diversity of porcine teschoviruses (PTVs; previously named porcine enterovirus 1) and most serotypes of porcine enteroviruses (PEVs) was studied. Following the determination of the major portion of the genomic sequence of PTV reference strain Talfan, the nucleotide and derived amino acid sequences of the RNA-dependent RNA polymerase (RdRp) region, the capsid VP2 region and the 3' non-translated region (3'-NTR) were compared among PTVs and PEVs and with other picornaviruses. The sequences were obtained by RT-PCR and 3'-RACE with primers based on the sequences of Talfan and available PEV strains. Phylogenetic analysis of RdRp/VP2 and analysis of the predicted RNA secondary structure of the 3'-NTR indicated that PEVs should be reclassified genetically into at least three groups, one that should be assigned to PTVs and two PEV subspecies represented by strain PEV-8 V13 and strain PEV-9 UKG410/73.

Animals↗

Comparative genome mapping of Pseudomonas aeruginosa PAO with P. aeruginosa C, which belongs to a major clone in cystic fibrosis patients and aquatic habitats.

A physical and genetic map was constructed for Pseudomonas aeruginosa C. Mainly, two-dimensional methods were used to place 47 SpeI, 8 PacI, 5 SwaI, and 4 I-CeuI sites onto the 6.5-Mb circular chromosome. A total of 21 genes, including the rrn operons and the origin of replication, were located on the physical map. Comparison of the physical and genetic map of strain C with that of the almost 600-kb-smaller genome of P. aeruginosa reference strain PAO revealed conservation of gene order between the two strains. A large-scale mosaic structure which was due to insertions of blocks of new genetic elements which had sizes of 23 to 155 kb and contained new SpeI sites was detected in the strain C chromosome. Most of these insertions were concentrated in three locations: two are congruent with the ends of the region rich in biosynthetic genes, and the third is located in the proposed region of the replication terminus. In addition, three insertions were scattered in the region rich in biosynthetic genes. The arrangement of the rrn operons around the origin of replication was conserved in C, PAO, and nine other examined independent strains.

Chromosome Mapping↗

Ontology annotation: mapping genomic regions to biological function.

With numerous whole genomes now in hand, and experimental data about genes and biological pathways on the increase, a systems approach to biological research is becoming essential. Ontologies provide a formal representation of knowledge that is amenable to computational as well as human analysis, an obvious underpinning of systems biology. Mapping function to gene products in the genome consists of two, somewhat intertwined enterprises: ontology building and ontology annotation. Ontology building is the formal representation of a domain of knowledge; ontology annotation is association of specific genomic regions (which we refer to simply as 'genes', including genes and their regulatory elements and products such as proteins and functional RNAs) to parts of the ontology. We consider two complementary representations of gene function: the Gene Ontology (GO) and pathway ontologies. GO represents function from the gene's eye view, in relation to a large and growing context of biological knowledge at all levels. Pathway ontologies represent function from the point of view of biochemical reactions and interactions, which are ordered into networks and causal cascades. The more mature GO provides an example of ontology annotation: how conclusions from the scientific literature and from evolutionary relationships are converted into formal statements about gene function. Annotations are made using a variety of different types of evidence, which can be used to estimate the relative reliability of different annotations.

Animals↗

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface↗

Cloning, sequencing, and expression of the apa gene coding for the Mycobacterium tuberculosis 45/47-kilodalton secreted antigen complex.

Effective protection against a virulent challenge with Mycobacterium tuberculosis is induced mainly by previous immunization with living attenuated mycobacteria, and it has been hypothesized that secreted proteins serve as major targets in the specific immune response. To identify and purify molecules present in culture medium filtrate which are dominant antigens during effective vaccination, a two-step selection procedure was used to select antigens able to interact with T lymphocytes and/or antibodies induced by immunization with living bacteria and to counterselect antigens interacting with the immune effectors induced by immunization with dead bacteria. A Mycobacterium bovis BCG 45/47-kDa antigen complex, present in BCG culture filtrate, has been previously identified and isolated (F. Romain, A. Laqueyrerie, P. Militzer, P. Pescher, P. Chavarot, M. Lagranderie, G. Auregan, M. Gheorghiu, and G. Marchal, Infect. Immun. 61:742-750, 1993). Since the cognate antibodies recognize the very same antigens present in M. tuberculosis culture medium filtrates, a project was undertaken to clone, express, and sequence the corresponding gene of M. tuberculosis. An M. tuberculosis shuttle cosmid library was transferred in Mycobacterium smegmatis and screened with a competitive enzyme-linked immunosorbent assay to detect the clones expressing the proteins. A clone containing a 40-kb DNA insert was selected, and by means of subcloning in Escherichia coli, a 2-kb fragment that coded for the molecules was identified. An open reading frame in the 2,061-nucleotide sequence codes for a secreted protein with a consensus signal peptide of 39 amino acids and a predicted molecular mass of 28,779 Da. The gene was referred to as apa because of the high percentages of proline (21.7%) and alanine (19%) in the purified protein. Southern hybridization analysis of digested total genomic DNA from M. tuberculosis (reference strains H37Rv and H37Ra) indicated that the apa gene was present as a single copy on the genome. The N-terminal identity or homology of the M. tuberculosis and M. bovis BCG purified molecules and their similar global and deduced amino acid compositions demonstrated the perfect correspondence between the molecular and chemical analyses. The presence of a high percentage of proline (21.7%) was confirmed and explained the apparent higher molecular mass (45/47 kDa) determined by sodium dodecyl sulfate-polyacrylamide gel electrophoresis resulting from the increased rigidity of molecules due to proline residues.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

Comparative genomic hybridization: an overview.

Comparative genomic hybridization (CGH) is a newly described molecular-cytogenetic assay that globally assays for chromosomal gains and losses in a genomic complement. In this assay, normal human metaphase chromosomes are competitively hybridized with two differentially labeled genomic DNAs (test and reference), which upon fluorescence microscopy, reveal the chromosomal locations of copy number changes in DNA sequences between the two complements. Application of CGH to DNAs extracted from fresh frozen specimens and cell lines of various tumor types has revealed a number of recurring chromosomal gains and losses that were undetected by traditional cytogenetic analysis. Few previously known sites were found to be in higher copy number, or lost by CGH, while many novel amplified regions were identified. These regions warrant further molecular genetic studies aimed at isolating the perturbed genes. Since CGH can also be performed on DNA extracted from formalin-fixed paraffin-embedded archived tumor specimens with few modifications, gains and losses of genetic material can be determined for specimens that would otherwise be unanalyzable. Prospective and retrospective application of CGH to tumor specimens would permit correlative studies to be performed, possibly identifying diagnostic and prognostic indicators of disease. CGH may also have a future role in detection and identification of chromosomal abnormalities in prenatal diagnosis and in dysmorphic anomalies.

Chromosome Mapping↗

Assignment of Rfp-Y to the chicken major histocompatibility complex/NOR microchromosome and evidence for high-frequency recombination associated with the nucleolar organizer region.

Rfp-Y is a second region in the genome of the chicken containing major histocompatibility complex (MHC) class I and II genes. Haplotypes of Rfp-Y assort independently from haplotypes of the B system, a region known to function as a MHC and to be located on chromosome 16 (a microchromosome) with the single nucleolar organizer region (NOR) in the chicken genome. Linkage mapping with reference populations failed to reveal the location of Rfp-Y, leaving Rfp-Y unlinked in a map containing >400 markers. A possible location of Rfp-Y became apparent in studies of chickens trisomic for chromosome 16 when it was noted that the intensity of restriction fragments associated with Rfp-Y increased with increasing copy number of chromosome 16. Further evidence that Rfp-Y might be located on chromosome 16 was obtained when individuals trisomic for chromosome 16 were found to transmit three Rfp-Y haplotypes. Finally, mapping of cosmid cluster III of the molecular map of chicken MHC genes (containing a MHC class II gene and two rRNA genes) to Rfp-Y validated the assignment of Rfp-Y to the MHC/NOR microchromosome. A genetic map can now be drawn for a portion of chicken chromosome 16 with Rfp-Y, encompassing two MHC class I and three MHC class II genes, separated from the B system by a region containing the NOR and exhibiting highly frequent recombination.

Animals↗

Evolution of paralogous genes: Reconstruction of genome rearrangements through comparison of multiple genomes within Staphylococcus aureus.

Analysis of evolution of paralogous genes in a genome is central to our understanding of genome evolution. Comparison of closely related bacterial genomes, which has provided clues as to how genome sequences evolve under natural conditions, would help in such an analysis. With species Staphylococcus aureus, whole-genome sequences have been decoded for seven strains. We compared their DNA sequences to detect large genome polymorphisms and to deduce mechanisms of genome rearrangements that have formed each of them. We first compared strains N315 and Mu50, which make one of the most closely related strain pairs, at the single-nucleotide resolution to catalogue all the middle-sized (more than 10 bp) to large genome polymorphisms such as indels and substitutions. These polymorphisms include two paralogous gene sets, one in a tandem paralogue gene cluster for toxins in a genomic island and the other in a ribosomal RNA operon. We also focused on two other tandem paralogue gene clusters and type I restriction-modification (RM) genes on the genomic islands. Then we reconstructed rearrangement events responsible for these polymorphisms, in the paralogous genes and the others, with reference to the other five genomes. For the tandem paralogue gene clusters, we were able to infer sequences for homologous recombination generating the change in the repeat number. These sequences were conserved among the repeated paralogous units likely because of their functional importance. The sequence specificity (S) subunit of type I RM systems showed recombination, likely at the homology of a conserved region, between the two variable regions for sequence specificity. We also noticed novel alleles in the ribosomal RNA operons and suggested a role for illegitimate recombination in their formation. These results revealed importance of recombination involving long conserved sequence in the evolution of paralogous genes in the genome.

Amino Acid Sequence↗

CGHPRO -- a comprehensive data analysis tool for array CGH.

BACKGROUND: Array CGH (Comparative Genomic Hybridisation) is a molecular cytogenetic technique for the genome wide detection of chromosomal imbalances. It is based on the co-hybridisation of differentially labelled test and reference DNA onto arrays of genomic BAC clones, cDNAs or oligonucleotides, and after correction for various intervening variables, loss or gain in the test DNA can be indicated from spots showing aberrant signal intensity ratios. Now that this technique is no longer confined to highly specialized laboratories and is entering the realm of clinical application, there is a need for a user-friendly software package that facilitates estimates of DNA dosage from raw signal intensities obtained by array CGH experiments, and which does not depend on a sophisticated computational environment. RESULTS: We have developed a user-friendly and versatile tool for the normalization, visualization, breakpoint detection and comparative analysis of array-CGH data. CGHPRO is a stand-alone JAVA application that guides the user through the whole process of data analysis. The import option for image analysis data covers several data formats, but users can also customize their own data formats. Several graphical representation tools assist in the selection of the appropriate normalization method. Intensity ratios of each clone can be plotted in a size-dependent manner along the chromosome ideograms. The interactive graphical interface offers the chance to explore the characteristics of each clone, such as the involvement of the clones sequence in segmental duplications. Circular Binary Segmentation and unsupervised Hidden Markov Model algorithms facilitate objective detection of chromosomal breakpoints. The storage of all essential data in a back-end database allows the simultaneously comparative analysis of different cases. The various display options facilitate also the definition of shortest regions of overlap and simplify the identification of odd clones. CONCLUSION: CGHPRO is a comprehensive and easy-to-use data analysis tool for array CGH. Since all of its features are available offline, CGHPRO may be especially suitable in situations where protection of sensitive patient data is an issue. It is distributed under GNU GPL licence and runs on Linux and Windows.

Algorithms↗

Scalable production of pectinases from Bacillus licheniformis SMIA-2 using agro-Industrial by-products with genomic insights.

UNLABELLED: The study re-analyzed the draft genome of Bacillus licheniformis SMIA-2 and generated a reference-guided pseudo-scaffold. Cross-validated genome annotation identified five candidate loci associated with pectin degradation, including putative pectate lyases, polygalacturonase, and downstream uronate-catabolic genes. Submerged fermentation with passion fruit peel flour and corn steep liquor yielded crude enzymatic extracts, which were spray-dried at 110 °C using maltodextrin and microcrystalline cellulose as stabilizers. The dried formulation retained pectinase activity for 180 days at 5 °C and showed additional cellulase, amylase, xylanase, and protease activities. Pectinase displayed optimal activity at pH 8.5 and 70 °C, with stability between pH 8.0-8.5 and 65-70 °C. Despite not using a reference strain and the absence of some omics analyses, with genomic and industrial claims presented as evidence of biotechnological potential rather than definitive functional validation of individual genes, these results support a sustainable, scalable, and alkaline-tolerant enzyme platform based on agro-industrial residues. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s10068-026-02252-3.

Agro-industrial residues↗

p53 expression in normal versus transformed mammalian cells.

To answer the question whether the level of p53 expression also reflects the status of a cell, with reference to transformation and genome stability, we have examined, by immunocytochemistry, the presence of p53 protein in a number of cell types including human diploid cells, Chinese hamster embryonal cells at different passages and gene amplified and/or transformed Chinese hamster cell lines. Primary human fibroblasts at early passage (LEO) and an established, non transformed, Chinese hamster cell line at early passage (CHEF/18) did not show any detectable p53 expression, either nuclear or cytoplasmic. All transformed human (Raji) and Chinese hamster cell lines (CHO, V79, V79/B7) showed a nuclear expression of p53, although at different intensities. Two cell lines selected from V79/B7 for their resistance to phosphonacetyl-L-aspartate or methotrexate and previously shown to bear gene amplification, showed p53 expression. In PALA L cells p53 expression was nuclear as in other positive cell lines tested, while in MTX M cells it was cytoplasmic. CHEF/18 cells at late passage in culture showed the typical behaviour of transformed cells and p53 was detected in several cells. Moreover, when transformed CHO cells were treated with compounds known to induce reverse transformation, both the disappearance of hallmarks of transformed phenotype and p53 reduction were observed. These results indicate a strong association within the same cell type between p53 expression and transformed status.

Animals↗

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗

Implementation of SMA carrier testing in genetic laboratories: comparison of two methods for quantifying the SMN1 gene.

The degeneration and loss of motor neurons of the anterior horn characterize children affected with spinal muscular atrophy (SMA). Mutations in the survival motor neuron gene (SMN1) are determinant for the development of the disease whereas the number of copies of SMN2, the highly homologous copy of SMN1, plays a role as a phenotypic modifier factor. The detection of SMN1 homozygous deletions is the typical test for SMA diagnosis. Owing to the limitation of this test for carrier and heterozygous deletion analysis, the demand of SMN1 quantitative tests is permanently growing. The high incidence of SMA, the notable carrier frequency, the severity of the disease, and the lack of effective treatment may justify the implementation of such an analysis in DNA diagnostic labs. The advantages and disadvantages of two reliable quantitative methods were evaluated. One of these is a competitive PCR protocol using internal standards and a genomic sequence as a reference. The other method is a real-time PCR employing an external standard as a reference. Both methods present sufficient advantages for incorporation into molecular genetic diagnostic labs. The possibility of studying samples from different labs, the versatility and reproducibility of the analysis, and cost-benefit calculations must be considered in the final choice.

Cyclic AMP Response Element-Binding Protein↗

Variability in molecular typing of Coxsackie A viruses by RFLP analysis and sequencing.

The aim of the present study was to develop an assay capable of classifying the Coxsackie A virus (CAV) prototype strains on the basis of restriction fragment length polymorphism (RFLP) analysis of 5'-UTR-derived reverse transcription polymerase chain reaction (RT-PCR) amplicons, and to determine how these data could be used for typing wild-type CAV isolates. Moreover, sequencing of the amplified genomic fragments of the clinical isolates, and comparison with all the published sequences of the respective genomic region of enterovirus reference and wild-type strains were attempted for typing of the isolates. Twenty-four prototype CAV strains from the 23 currently recognized serotypes were studied; most of them were successfully differentiated with the aid of four restriction endonucleases: HaeIII, HpaII, DdeI, and StyI. It was not possible to differentiate between CAV5, 7, and 16, or between CAV15 and 18 in this way, but the members of each of these two groups were satisfactorily differentiated with the aid of single-strand conformational polymorphism (SSCP) analysis of their RT-PCR amplicons. Fifteen clinical isolates, 13 of them of known CAV serotype, were also studied with the same four endonucleases and the results were compared with the data obtained from the RFLP analysis of the reference strains. The experimental results showed that only two clinical samples of previously known identity had an identical restriction pattern with the respective prototype strains. The sequences of the amplicons of the clinical isolates had the greatest percentage of alignment with enterovirus strains of a different serotype, indicating variability in the 5'-UTR and the inability to use the whole sequence of the amplicons for typing CAVs. The significance of the findings in relation to the possible usefulness of the RFLP-based method is discussed.

Cells, Cultured↗