Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

CAFTAN: a tool for fast mapping, and quality assessment of cDNAs.

BACKGROUND: The German cDNA Consortium has been cloning full length cDNAs and continued with their exploitation in protein localization experiments and cellular assays. However, the efficient use of large cDNA resources requires the development of strategies that are capable of a speedy selection of truly useful cDNAs from biological and experimental noise. To this end we have developed a new high-throughput analysis tool, CAFTAN, which simplifies these efforts and thus fills the gap between large-scale cDNA collections and their systematic annotation and application in functional genomics. RESULTS: CAFTAN is built around the mapping of cDNAs to the genome assembly, and the subsequent analysis of their genomic context. It uses sequence features like the presence and type of PolyA signals, inner and flanking repeats, the GC-content, splice site types, etc. All these features are evaluated in individual tests and classify cDNAs according to their sequence quality and likelihood to have been generated from fully processed mRNAs. Additionally, CAFTAN compares the coordinates of mapped cDNAs with the genomic coordinates of reference sets from public available resources (e.g., VEGA, ENSEMBL). This provides detailed information about overlapping exons and the structural classification of cDNAs with respect to the reference set of splice variants. The evaluation of CAFTAN showed that is able to correctly classify more than 85% of 5950 selected "known protein-coding" VEGA cDNAs as high quality multi- or single-exon. It identified as good 80.6 % of the single exon cDNAs and 85 % of the multiple exon cDNAs. The program is written in Perl and in a modular way, allowing the adoption of this strategy to other tasks like EST-annotation, or to extend it by adding new classification rules and new organism databases as they become available. We think that it is a very useful program for the annotation and research of unfinished genomes. CONCLUSION: CAFTAN is a high-throughput sequence analysis tool, which performs a fast and reliable quality prediction of cDNAs. Several thousands of cDNAs can be analyzed in a short time, giving the curator/scientist a first quick overview about the quality and the already existing annotation of a set of cDNAs. It supports the rejection of low quality cDNAs and helps in the selection of likely novel splice variants, and/or completely novel transcripts for new experiments.

Chromosome Mapping↗

Structural variation of the human genome.

There is growing appreciation that the human genome contains significant numbers of structural rearrangements, such as insertions, deletions, inversions, and large tandem repeats. Recent studies have defined approximately 5% of the human genome as structurally variant in the normal population, involving more than 800 independent genes. We present a detailed review of the various structural rearrangements identified to date in humans, with particular reference to their influence on human phenotypic variation. Our current knowledge of the extent of human structural variation shows that the human genome is a highly dynamic structure that shows significant large-scale variation from the currently published genome reference sequence.

Genetic Variation↗

Systematic analysis of embryonic expression profiles of zinc finger genes in Ciona intestinalis.

The recent decoding of a number of animal genomes has provided unprecedented information regarding evolution and gene structures, but this information must be supplemented with precise gene annotations and the temporal and spatial expression patterns of individual genes. In the present study, we systematically identified and characterized 566 zinc finger genes in the genome of Ciona intestinalis, an emerging model system for genome-wide studies of development and evolution. Of these genes, 356 genes encoded a potential transcription factor based on putative nucleic acid binding activity or domains of unknown function. We further examined the expression patterns of 225 genes during embryogenesis, and, when considered with a previous study [Imai, K.S., Hino, K., Yagi, K., Satoh, N., Satou, Y., 2004. Gene expression profiles of transcription factors and signaling molecules in the ascidian embryo: towards a comprehensive understanding of gene networks. Development 131, 4047-4058], we have characterized the developmental expression patterns of nearly 85% of the potential zinc finger-containing transcription factors. Overall, zinc finger genes are preferentially maternally expressed with little larval expression during development. The present study provides a valuable reference for genome-wide studies in this species and for future studies wishing to examine zinc finger gene expression patterns in other animals.

Animals↗

Genomic fingerprinting and development of a dendrogram for Brucella spp. isolated from seals, porpoises, and dolphins.

Genomic DNA from reference strains and biovars of the genus Brucella was analyzed using pulsed-field gel electrophoresis (PFGE). Fingerprints were compared to estimate genetic relatedness among the strains and to obtain information on evolutionary relationships. Electrophoresis of DNA digested with the restriction endonuclease XbaI produced fragment profiles for the reference type strains that distinguished these strains to the level of species. Included in this study were strains isolated from marine mammals. The PFGE profiles from these strains were compared with those obtained from the reference strains and biovars. Isolates from dolphins had similar profiles that were distinct from profiles of Brucella isolates from seals and porpoises. Distance matrix analyses were used to produce a dendrogram. Biovars of B. abortus were clustered together in the dendrogram; similar clusters were shown for biovars of B. melitensis and for biovars of B. suis. Brucella ovis, B. canis, and B. neotomae differed from each other and from B. abortus, B. melitensis, and B. suis. The relationship between B. abortus strain RB51 and other Brucella biovars was compared because this strain has replaced B. abortus strain 19 for use as a live vaccine in cattle and possibly in bison and elk. These results support the current taxonomy of Brucella species and the designation of an additional genomic group(s) of Brucella. The PFGE analysis in conjunction with distance matrix analysis was a useful tool for calculating genetic relatedness among the Brucella species.

Animals↗

Aligning multiple genomic sequences with the threaded blockset aligner.

We define a "threaded blockset," which is a novel generalization of the classic notion of a multiple alignment. A new computer program called TBA (for "threaded blockset aligner") builds a threaded blockset under the assumption that all matching segments occur in the same order and orientation in the given sequences; inversions and duplications are not addressed. TBA is designed to be appropriate for aligning many, but by no means all, megabase-sized regions of multiple mammalian genomes. The output of TBA can be projected onto any genome chosen as a reference, thus guaranteeing that different projections present consistent predictions of which genomic positions are orthologous. This capability is illustrated using a new visualization tool to view TBA-generated alignments of vertebrate Hox clusters from both the mammalian and fish perspectives. Experimental evaluation of alignment quality, using a program that simulates evolutionary change in genomic sequences, indicates that TBA is more accurate than earlier programs. To perform the dynamic-programming alignment step, TBA runs a stand-alone program called MULTIZ, which can be used to align highly rearranged or incompletely sequenced genomes. We describe our use of MULTIZ to produce the whole-genome multiple alignments at the Santa Cruz Genome Browser.

Animals↗

Proteome reference map of Pseudomonas putida strain KT2440 for genome expression profiling: distinct responses of KT2440 and Pseudomonas aeruginosa strain PAO1 to iron deprivation and a new form of superoxide dismutase.

The genome sequence of Pseudomonas putida strain KT2440, a nutritionally versatile, saprophytic and plant root-colonizing Gram-negative soil bacterium, was recently determined by K. E. Nelson et al. (2002, Environ Microbiol 4: 799-808). Here, we present a two-dimensional gel protein reference map of KT2440 cells grown in mineral salts medium with glucose as carbon source. Proteins were identified by matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) analysis, in conjunction with an in-house database developed from the genome sequence of KT2440, and approximately 200 two-dimensional gel spots were assigned. The map was used to assess the genomic response of KT2440 to iron limitation stress and to compare this response with that of the closely related facultative human pathogen Pseudomonas aeruginosa strain PAO1. The synthesis of about 25 proteins was affected in both strains, including four prominent upregulated ferric uptake regulator (Fur) protein-dependent proteins, but there were also striking differences in their proteome responses, for example in the expression of superoxide dismutases (Sod), which may indicate important roles of iron-responsive functions in the adaptation of these two bacteria to different lifestyles. The Sod enzyme of KT2440 was shown to be a novel heterodimer of the SodA and SodB polypeptides.

Adaptation, Biological↗

The application of molecular phylogenetics to the analysis of viral genome diversity and evolution.

DNA sequencing and molecular phylogenetics are increasingly being used in virology laboratories to study the transmission of viruses. By reconstructing the evolutionary history of viral genomes the behaviour of viral populations can be modelled, and the future of epidemics may be forecast. The manner in which such viral DNA sequences are analysed is the focus of this review. Many researchers resort to the often-quoted 'black box' approach because phylogenetics theory can be daunting, and phylogenetics software packages can appear to be difficult to use. However, because phylogenetic analyses are often used in important and sensitive arenas, for example to provide evidence indicating transmission between persons, it is vital that appropriate care is taken to estimate reliably true relationships. In this review, we discuss how a molecular phylogenetics study should be approached, give an overview of the methods and programs for analysing DNA sequence data, and point readers to appropriate texts for further details. The aim of this review, therefore, is to provide researchers with an easy to understand guide to molecular phylogenetics, with special reference to viral genomes.

Amino Acid Sequence↗

Mapping chicken genes using preferential amplification of specific alleles.

To map the chicken genome, an international reference population was developed at our laboratory (East Lansing, MI) using an F2 backcross between inbred jungle fowl (JF) and inbred white leghorns (WL). To augment the number of type I genes on the East Lansing (E) map, segregation of the JF-specific allele was followed using preferential amplification of specific alleles (PASA) in polymerase chain reactions (PCR). Among 15 functional genes that were added to the E map, agrin and mannose-6-phosphate receptor genes were found to occur in conserved syntenic groups. Using this PCR-based approach, six conserved groups spanning more than 243 centimorgans (cM) in the chicken were syntenic with human and mouse.

Alleles↗

Comparison of the Escherichia coli K-12 genome with sampled genomes of a Klebsiella pneumoniae and three salmonella enterica serovars, Typhimurium, Typhi and Paratyphi.

The Escherichia coli K-12 genome (ECO) was compared with the sampled genomes of the sibling species Salmonella enterica serovars Typhimurium, Typhi and Paratyphi A (collectively referred to as SAL) and the genome of the close outgroup Klebsiella pneumoniae (KPN). There are at least 160 locations where sequences of >400 bp are absent from ECO but present in the genomes of all three SAL and 394 locations where sequences are present in ECO but close homologs are absent in all SAL genomes. The 394 sequences in ECO that do not occur in SAL contain 1350 (30.6%) of the 4405 ECO genes. Of these, 1165 are missing from both SAL and KPN. Most of the 1165 genes are concentrated within 28 regions of 10-40 kb, which consist almost exclusively of such genes. Among these regions were six that included previously identified cryptic phage. A hypothetical ancestral state of genomic regions that differ between ECO and SAL can be inferred in some cases by reference to the genome structure in KPN and the more distant relative Yersinia pestis. However, many changes between ECO and SAL are concentrated in regions where all four genera have a different structure. The rate of gene insertion and deletion is sufficiently high in these regions that the ancestral state of the ECO/SAL lineage cannot be inferred from the present data. The sequencing of other closely related genomes, such as S.bongori or Citrobacter, may help in this regard.

Databases as Topic↗

Effectiveness of mass spectrometry and genomic analysis in the surveillance of nontuberculous Mycobacterium in Taiwan.

Nontuberculous mycobacteria (NTM) are diverse, and species-level identification remains challenging in routine diagnostics. We analyzed NTM isolates collected at three regional centers of the National Taiwan University Hospital (NTUH) from 2019 to 2024 to assess geographic variation and identification performance after implementation of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS). Among 3,188 cases meeting the microbiological criteria for probable pulmonary NTM disease, the species distribution differed by region: Mycobacterium avium complex predominated in central Taiwan (Yunlin, 47.3%), whereas M. abscessus complex (Taipei, 26.5%) and M. kansasii (Hsinchu, 12.4%) were more common in northern Taiwan. In 2019, 14.5% of isolates were reported to be unidentified by MALDI-TOF MS; with workflow optimization and database updates, this percentage decreased but plateaued at 4.5-4.8%. Whole-genome sequencing (WGS) of 61 randomly selected persistently unidentified isolates revealed eight average nucleotide identity (ANI)-defined clusters; 55 isolates (90.2%) could not be assigned to known species using current reference databases. Two clusters detected only in Hsinchu were phylogenetically closest to M. kyorinense, with ANI values below the species demarcation threshold. Overall, we observed marked regional heterogeneity of NTM in Taiwan and a persistent identification gap that remained after MALDI-TOF MS optimization and follow-up WGS.IMPORTANCEThis study characterized regional differences in the NTM species distribution across Taiwan, and the results highlight the limitations of current identification approaches. MALDI-TOF MS identifies most isolates, but locally circulating lineages represent a persistent gap in global reference libraries. Even with whole-genome sequencing (WGS), 90.2% (55/61) of persistently unresolved isolates could not be assigned to known species in the current reference databases despite the formation of clear ANI- and phylogeny-defined clusters. These findings show that both proteomic and genomic reference resources for clinical NTM remain incomplete. Expanding regionally representative databases and performing WGS for isolates that remain unresolved by MALDI-TOF MS will be necessary to improve species-level resolution for surveillance and clinical interpretation.

Taiwan↗

Two distinct upstream regulatory domains containing multicopy cellular transcription factor binding sites provide basal repression and inducible enhancer characteristics to the immediate-early IES (US3) promoter from human cytomegalovirus.

The US3 gene of human cytomegalovirus (HCMV) is expressed at immediate-early (IE) times in permissive HF cells, but not in nonpermissive rodent cells, and encodes several proteins that have been reported to have regulatory characteristics, although they are dispensable for growth in cell culture. Both spliced and unspliced forms of US3 IE transcripts are associated with the second of only two known large and complex upstream enhancer domains within the 229-kb HCMV genome, which we refer to as the IES cis-acting control region. Only the 260-bp proximal segment (from -313 to -55) of the 600-bp IES control domain, which contains multicopy NF-kappaB binding sites, proved to be necessary to transfer both high basal expression plus phorbol ester- and okadaic acid-inducible characteristics to heterologous promoters in transient assays in U-937 and K-562 cells. However, the IES control region contains a distinctive 280-bp distal domain, characterized by the presence of seven interspersed repeats of a 10-bp TGTCGCGACA palindromic consensus motif that encompasses a NruI site. This far upstream Nru repeat region (from -596 to -314) imparted up to 20-fold down-regulation effects onto strong basal heterologous promoters as well as onto the IES enhancer plus minimal promoter region in both U-937 and K-562 cells. Functional Nru repressor elements (NREs) could not be generated by multimerizing either the palindromic (P) Nru motifs alone or adjacent degenerate interrupted (I Nru motifs alone. However, multimerized forms of the combined P plus I elements reconstituted the full 20-fold cis-acting down-regulation phenotype of the intact NRE domain. The P and I forms of the Nru elements each bound independently and specifically to related cellular DNA-binding factors to form differently migrating A or B complexes, respectively, whereas the combined P plus I elements bound cooperatively to both the A and B complexes with high affinity. Interestingly, nuclear extracts from U-937, K-562, HeLa, and Vero cells all formed both the A and B NRE binding factor complexes, whereas those from HF cells produced only A complexes, and Raji, HL60, and BALB/c 3T3 cells lacked both types of binding factor complexes. The core pentameric CGACA and CGATA half sites present in both the P and I Nru motifs are related to recently described Drosophila chromosomal insulator binding sites. Therefore, in addition to its cis-repression or silencer characteristics, the NRE domain appears likely to act to shield adjacent segments of the viral genome from the chromatin-reorganizing effects of the IES-inducible enhancer. We speculate that differential expression and regulation of the IES enhancer-controlled US3 protein, either in concert with or separately from the major IE (MIE) enhancer-controlled IE1 and IE2 transactivator proteins, may play a critical role in determining HCMV permissiveness in some cell types and perhaps also in the establishment of or reactivation from latency.

Animals↗

Amplification of oncogenes in human cancer cells.

Gene amplification refers to a genomic change that results in an increased dosage of the gene(s) affected. Amplification represents one of the major molecular pathways through which the oncogenic potential of proto-oncogenes is activated during tumorigenesis. The architecture of amplified genomic structures is simple in some tumor types, involving in the vast majority of cases only one gene, such as MYCN in neuroblastomas. On the other hand, it can be complex and discontinuous, involving several syntenic co-amplified genes, such as in the 11q13 amplification in breast cancer, although in many of these cases there may be a single target gene. The presence of different nonsyntenic amplified genes raises the possibility that cells of certain tumors are susceptible to independent amplification events. In general, the amplified genes do not undergo additional damage by mutations. The data indicate that it is the enhanced level of a wild-type protein that contributes to tumorigenesis.

Breast Neoplasms↗

Monozygotic twinning and Wiedemann-Beckwith syndrome.

Monozygotic (MZ) twinning occurs with relatively high frequency in Wiedemann-Beckwith syndrome (WBS). Ten sets of MZ twins with WBS have been reported. Nine of these have been female and in each case the twins were discordant for the WBS phenotype. The tenth set was male. They were concordant for WBS and both had a duplication of chromosome 15 which they shared in common with their phenotypically normal mother. The WBS gene has been assigned to the locus 11p15 and there appear to be several different genetic mechanisms involving this locus which all give rise to WBS. An imprinting effect for the WBS gene has been proposed because of the transmission of the gene preferentially through the maternal line in some large pedigrees. We describe two further sets of female MZ twins with WBS. One pair is concordant and one discordant for the condition. The possible genetic mechanisms involved in the expression of WBS are discussed, with particular reference to twinning, genomic imprinting and X-inactivation which is thought to be associated with the occurrence of MZ twinning in females.

Beckwith-Wiedemann Syndrome↗

An on-line two-dimensional polyacrylamide gel electrophoresis protein database of adult Drosophila melanogaster.

An annotated two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) protein database of adult Drosophila melanogaster has been constructed, based on the protein patterns of heads, thoraces and abdomens of adult male and female Drosophila melanogaster. About 1200 major protein spots are catalogued. Common proteins, found in all body parts, as well as bodypart- and sex-specifically expressed proteins are reported. Of the major proteins, 91, or 7.5%, are differentially expressed in the two sexes or in different body parts, at least in part reflecting specific functional requirements. At the present time 43 proteins, or about 3.5% of the detected proteins, have been identified. These data can be accessed interactively from our World Wide Web (WWW) server through clickable inline gel images and hypertext links. Identified protein spots are cross-referenced, through hypertext links, to the SWISS-PROT annotated database of protein primary sequences and the Fly-Base database of Drosophila genomic data. Our reference gels can be used to gain immediate access to protein spot identify and to the pattern of differentially expressed proteins in Drosophila melanogaster. The work presented in this article ties together information from protein 2-D PAGE, molecular biology and genetics and offers a uniform way to access this large volume of data.

Abdomen↗

Nebulin cDNAs detect a 25-kilobase transcript in skeletal muscle and localize to human chromosome 2.

By virtue of the protein's size, myofibrillar localization, and proposed functional role, the gene encoding the giant sarcomere matrix protein nebulin represents a possible site for myopathic mutations. Using polyclonal anti-nebulin antisera to screen a cDNA expression library, we have isolated and characterized two separate human fetal muscle cDNA clones. By recovering fusion polypeptide-bound portions of our polyclonal antiserum and reutilizing them to probe Western blots, we further demonstrate that the expressed cDNAs encode polypeptide epitopes unique to the protein nebulin. Both cDNAs detect a 25-kb skeletal muscle RNA transcript and localize to human chromosome 2. The identification of nebulin cDNA clones enables the complete analysis of this enormous mRNA by transcript walking through muscle cDNA libraries. Here we report a restriction map of the 3' end of the human nebulin transcript, with reference to the genomic fragments identified by the cDNA.

Chromosome Mapping↗

The nuclear envelope: form and reformation.

The membrane system that encloses genomic DNA is referred to as the nuclear envelope. However, with emerging roles in signaling and gene expression, these membranes clearly serve as more than just a physical barrier separating the nucleus and cytoplasm. Recent progress in our understanding of nuclear envelope architecture and composition has also revealed an intriguing connection between constituents of the nuclear envelope and human disease, providing further impetus to decipher this cellular structure and the dramatic remodeling process it undergoes with each cell division.

Animals↗

National epidemic of Lordsdale Norovirus in the UK.

BACKGROUND: In early 2002 reports of outbreaks of gastroenteritis reached unprecedented levels in the UK. Forty five Norovirus outbreaks were reported in January 2002. OBJECTIVES: The objective of the study was to determine whether the outbreaks were Noroviral in origin and if so whether they represented a homogeneous or heterogeneous collection of Noroviruses by applying EIA and sequence analysis to representative faecal samples. STUDY DESIGN: Faecal specimens were collected during the week of highest incidence from 21 outbreaks in a variety of health care settings including hospitals and nursing homes. The outbreaks occurred in geographically distinct regions of the UK and samples were collected by reference laboratories in Glasgow, Manchester, Bristol and Southampton. RESULTS: The samples were all positive for Noroviruses by negative stain electron microscopy (EM) and Lordsdale virus (LV) EIA, therefore reverse transcriptase polymerase chain reaction (RT-PCR) amplification and nucleotide sequencing of the Norovirus RNA polymerase gene was performed on amplicons from samples of each of the 21 outbreaks to investigate the nature and extent of diversity. All samples were very closely related to the reference Lordsdale virus genome sequence. LV was first discovered during an hospital outbreak of gastroenteritis in Southampton General Hospital in March 1993. CONCLUSIONS: Noroviruses are a major cause of outbreaks of gastroenteritis in health care settings. LV is the predominant Norovirus in the UK and was detected in outbreaks that occurred during the national peak of gastroenteritis reports in January 2002.

Amino Acid Sequence↗

Detection of spring viremia of carp virus isolates by hybridization with non-radioactive probes and amplification by polymerase chain reaction.

For detection of spring viremia of carp virus (SVCV) DNA probes have been constructed using the reverse transcription-polymerase chain reaction (RT-PCR) amplification technique and cDNA cloning in plasmid and phage vectors. The specific primers for amplification of SVCV M and G genes were chosen and synthesized. Studies were carried out to establish the sensitivity and specificity of viral RNA detection in infected cell culture and pathogenic material from fish by the use of non-radioactive probes and RT-PCR. The efficiency of amplification with primers, complementary to the genome of the reference Fijan strain, was estimated in RT-PCR experiments with two SVCV strains. Under the same conditions, the quantity of PCR products amplified from the M2 strain was less than that from the ZL4 strain, which implies that the latter is more similar to the reference European SVCV Fijan isolate. Using DNA probes and dot-blot hybridization, SVCV was tested in samples taken from different organs of artificially infected carp with clinical signs of acute disease. The virus could be detected most reliably in fish brain. In most cases the hybridization signal was positive with samples having a viral titer of not less than 10(5) TCID50/g.

Animals↗