Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Consistent and idiosyncratic pleiotropy in shaping genetic correlations.

Pleiotropy, the phenomenon where a single mutation influences multiple phenotypic traits, creates genetic correlations that can constrain evolutionary trajectories. Yet genetic correlations differ in their persistence: some remain stable over long evolutionary timescales, whereas others change rapidly across generations or environments. One explanation is that similar values of genetic correlation, rG, can arise from different pleiotropic architectures: broadly aligned effects across many loci, or disproportionate covariance contributions from a few large effect loci. Motivated by the distinction between vertical and horizontal pleiotropy, here, we develop a bivariate marker effect framework for recombinant mapping populations that separates candidate large covariance contributors from the polygenic background correlation, rD. We define rD as the correlation among marker effects after trimming markers with unusually large covariance contributions. rD is a trait-pair summary of how consistently small and moderate effect markers align across the genome; high rD is expected when many perturbations propagate through shared developmental, physiological, causal, or geometric structure. Applying this framework to high-dimensional yeast single-cell morphology, we show that trait pairs with similar rG can differ substantially in rD, and that a small number of candidate outlier regions can strongly influence some marker effect correlations. We then test whether rD predicts the environmental stability of genetic correlations under geldanamycin-mediated Hsp90 perturbation. Trait pairs with stronger rD show smaller absolute changes in rG. These results suggest that genetic correlations supported by a strong polygenic marker effect background are more environmentally stable than correlations shaped primarily by a few large covariance contributors.

Genetic Pleiotropy↗

Isolation, structure and expression of mammalian genes for histidyl-tRNA synthetase.

A full length cDNA clone that codes for human histidyl-tRNA synthetase (HRS) and cDNA clones that span the full length transcript of hamster HRS have been isolated. The full length human HRS cDNA was expressed after transfection into Cos 1 cells and a CHO ts mutant defective in the gene for HRS. The complete nucleotide sequence of the hamster and human gene were obtained and extensive homologies were observed in three regions on comparing these sequences between themselves and with the sequence of HRS derived from yeast. These results provide unequivocal evidence that we have indeed cloned the hamster and human gene for HRS. Three overlapping phage recombinants containing the complete hamster chromosomal gene for HRS have also been isolated. The genomic HRS is divided into 13 exons. The precise locations of each of the 5' and 3' exon-intron boundaries were defined by sequencing the appropriate regions of the cloned genomic DNA and aligning them with the sequence of HRS cDNAs. These studies provide the basis for future structural and functional analysis of the gene for HRS. In particular, it will be of interest to examine if different exons of HRS correlate to different domains of the HRS polypeptide.

Amino Acyl-tRNA Synthetases↗

The gene coding for small ribosomal subunit RNA in the basidiomycete Ustilago maydis contains a group I intron.

The nucleotide sequence of the gene coding for small ribosomal subunit RNA in the basidiomycete Ustilago maydis was determined. It revealed the presence of a group I intron with a length of 411 nucleotides. This is the third occurrence of such an intron discovered in a small subunit rRNA gene encoded by a eukaryotic nuclear genome. The other two occurrences are in Pneumocystis carinii, a fungus of uncertain taxonomic status, and Ankistrodesmus stipitatus, a green alga. The nucleotides of the conserved core structure of 101 group I intron sequences present in different genes and genome types were aligned and their evolutionary relatedness was examined. This revealed a cluster including all group I introns hitherto found in eukaryotic nuclear genes coding for small and large subunit rRNAs. A secondary structure model was designed for the area of the Ustilago maydis small ribosomal subunit RNA precursor where the intron is situated. It shows that the internal guide sequence pairing with the intron boundaries fits between two helices of the small subunit rRNA, and that minimal rearrangement of base pairs suffices to achieve the definitive secondary structure of the 18S rRNA upon splicing.

Base Sequence↗

Utility of the white gene in estimating phylogenetic relationships among mosquitoes (Diptera: Culicidae).

The utility of a nuclear protein-coding gene for reconstructing phylogenetic relationships within the family Culicidae was explored. Relationships among 13 species representing three subfamilies and nine genera of Culicidae were analyzed using a 762-bp fragment of coding sequence from the eye color gene, white. Outgroups for the study were two species from the sister group Chaoboridae. Sequences were determined from clone PCR products amplified from genomic DNA, and aligned following conceptual intron splicing and amino acid translation. Third codon positions were characterized by high levels of divergence and biased nucleotide composition, the intensity and direction of which varied among taxa. Equal weighting of all characters resulted in parsimony and neighboring-joining trees at odds with the generally accepted phylogenetic hypothesis based on morphology and rDNA sequences. The application of differential weighting schemes recovered the traditional hypothesis, in which the subfamily Anophelinae formed the basal clade. The subfamily Toxorhynchitinae occupied an intermediate position, and was a sister group to the subfamily Culicinae. Within Culicinae, the genera Sabethes and Tripteroides formed an ancestral clade, while the Culex-Deinocerites and Aedes-Haemagogus clades occupied increasingly derived positions in the molecular phylogeny. An intron present in the Culicinae-Toxorhynchitinae lineage and one outgroup taxon was absent in the basal Anophelinae lineage and the second outgroup taxon, suggesting that intron insertions or deletions may not always be reliable systematic characters.

Amino Acid Sequence↗

Clustering of DNA sequences in human promoters.

We have determined the distribution of each of the 65,536 DNA sequences that are eight bases long (8-mer) in a set of 13,010 human genomic promoter sequences aligned relative to the putative transcription start site (TSS). A limited number of 8-mers have peaks in their distribution (cluster), and most cluster within 100 bp of the TSS. The 156 DNA sequences exhibiting the greatest statistically significant clustering near the TSS can be placed into nine groups of related sequences. Each group is defined by a consensus sequence, and seven of these consensus sequences are known binding sites for the transcription factors (TFs) SP1, NF-Y, ETS, CREB, TBP, USF, and NRF-1. One sequence, which we named Clus1, is not a known TF binding site. The ninth sequence group is composed of the strand-specific Kozak sequence that clusters downstream of the TSS. An examination of the co-occurrence of these TF consensus sequences indicates a positive correlation for most of them except for sequences bound by TBP (the TATA box). Human mRNA expression data from 29 tissues indicate that the ETS, NRF-1, and Clus1 sequences that cluster are predominantly found in the promoters of housekeeping genes (e.g., ribosomal genes). In contrast, TATA is more abundant in the promoters of tissue-specific genes. This analysis identified eight DNA sequences in 5082 promoters that we suggest are important for regulating gene expression.

Base Sequence↗

The ubiquitin-specific protease family from Arabidopsis. AtUBP1 and 2 are required for the resistance to the amino acid analog canavanine.

Ubiquitin-specific proteases (UBPs) are a family of unique hydrolases that specifically remove polypeptides covalently linked via peptide or isopeptide bonds to the C-terminal glycine of ubiquitin. UBPs help regulate the ubiquitin/26S proteolytic pathway by generating free ubiquitin monomers from their initial translational products, recycling ubiquitins during the breakdown of ubiquitin-protein conjugates, and/or by removing ubiquitin from specific targets and thus presumably preventing target degradation. Here, we describe a family of 27 UBP genes from Arabidopsis that contain both the conserved cysteine (Cys) and histidine boxes essential for catalysis. They can be clustered into 14 subfamilies based on sequence similarity, genomic organization, and alignments with their closest relatives from other organisms, with seven subfamilies having two or more members. Recombinant AtUBP2 functions as a bona fide UBP: It can release polypeptides attached to ubiquitins via either alpha- or epsilon-amino linkages by an activity that requires the predicted active-site Cys within the Cys box. From the analysis of T-DNA insertion mutants, we demonstrate that the AtUBP1 and 2 subfamily helps confer resistance to the arginine analog canavanine. This phenotype suggests that the AtUBP1 and 2 enzymes are needed for abnormal protein turnover in Arabidopsis.

Amino Acid Sequence↗

Annotation and evolutionary relationships of a small regulatory RNA gene micF and its target ompF in Yersinia species.

BACKGROUND: micF RNA, a small regulatory RNA found in bacteria, post-transcriptionally regulates expression of outer membrane protein F (OmpF) by interaction with the ompF mRNA 5'UTR. Phylogenetic data can be useful for RNA/RNA duplex structure analyses and aid in elucidation of mechanism of regulation. However micF and associated genes, ompF and ompC are difficult to annotate because of either similarities or divergences in nucleotide sequence. We report by using sequences that represent "gene signatures" as probes, e.g., mRNA 5'UTR sequences, closely related genes can be accurately located in genomic sequences. RESULTS: Alignment and search methods using NCBI BLAST programs have been used to identify micF, ompF and ompC in Yersinia pestis and Yersinia enterocolitica. By alignment with DNA sequences from other bacterial species, 5' start sites of genes and upstream transcriptional regulatory sites in promoter regions were predicted. Annotated genes from Yersinia species provide phylogenetic information on the micF regulatory system. High sequence conservation in binding sites of transcriptional regulatory factors are found in the promoter region upstream of micF and conservation in blocks of sequences as well as marked sequence variation is seen in segments of the micF RNA gene. Unexpected large differences in rates of evolution were found between the interacting RNA transcripts, micF RNA and the 5' UTR of the ompF mRNA. micF RNA/ompF mRNA 5' UTR duplex structures were modeled by the mfold program. Functional domains such as RNA/RNA interacting sites appear to display a minimum of evolutionary drift in sequence with the exception of a significant change in Y. enterocolitica micF RNA. CONCLUSIONS: Newly annotated Yersinia micF and ompF genes and the resultant RNA/RNA duplex structures add strong phylogenetic support for a generalized duplex model. The alignment and search approach using 5' UTR signatures may be a model to help define other genes and their start sites when annotated genes are available in well-defined reference organisms.

5' Untranslated Regions↗

Introgression among maternal lineages inferred from complete mitogenomes and molecular dating helps resolve phylogeography of European roe deer.

BACKGROUND: The European roe deer (Capreolus capreolus) is one of the most widespread ungulates in Europe, with a phylogeographic structure mainly shaped by Pleistocene glacial cycles and secondary contacts with the Siberian roe deer (C. pygargus). METHODS: We sequenced 52 complete mitogenomes of C. capreolus from Slovenia, Poland and France, and combined them with 24 publicly available sequences of C. capreolus and C. pygargus, yielding an alignment of 76 genomes representing 59 haplotypes (42 from C. capreolus and 17 from C. pygargus). Phylogeographic structure was assessed using a median-joining network, and divergence times were estimated using a time-calibrated Bayesian phylogeny based on mitochondrial coding regions, incorporating published ancient C. pygargus mitogenomes. We additionally screened mitochondrial protein-coding genes for selection. RESULTS: The haplotype network recovered the three major European roe deer clades (Eastern, Central, and Western) and detected Central-clade haplotypes in France. Two Polish haplotypes (Cp9 and Cp10), detected in C. capreolus, clustered within the C. pygargus mitochondrial lineage, supporting mitochondrial introgression. Time-calibrated phylogenies placed introgressed haplotypes within established C. pygargus lineages. Selection analyses provided limited evidence for episodic positive selection restricted to a small number of codons. CONCLUSIONS: Whole mitogenomes improve resolution of roe deer phylogeography and reveal introgressed maternal lineages, while time-calibrated phylogenies and selection tests add evolutionary context for interpreting mtDNA diversity in genus Capreolus.

Animals↗

Phylogenic analysis of rickettsial patatin-like protein with conserved phospholipase A2 active sites.

Genome analysis of Rickettsia felis highlighted the presence of three patatin-like protein (PLP) genes (pat1, pat2A, and pat2B), whereas only one PLP gene (pat1) is found in the other sequenced rickettsial genomes. Here, we aligned the rickettsial PLPs with characterized patatins from plants and found that they possess all the conserved amino acid residues identified as important for phospholipase A(2) activity. We also carried out a phylogenic analysis of the rickettsial PLPs together with bacterial and eukaryotic homologs. The phylogeny of the rickettsial Pat1 proteins is in conflict with the currently recognized Rickettsia phylogeny. Possible scenarios that might explain this incongruence are discussed and involve either gene conversion or gene duplication events.

Amino Acid Sequence↗

Development and epidemiological investigation of a TaqMan-based multiplex real-time quantitative PCR assay for simultaneous detection of five bovine viruses (BVDV, AKAV, BNoV, BEV, and BCoV).

INRODUCTION: Infectious diseases caused by bovine viral diarrhea virus (BVDV), Akabane virus (AKAV), bovine norovirus (BNoV), bovine enterovirus (BEV), and bovine coronavirus (BCoV) significantly threaten the cattle industry, resulting in substantial economic losses. These pathogens often present similar clinical signs, such as diarrhea, vomiting, and reproductive disorders in pregnant cattle, and frequent covert or mixed infections further complicate accurate diagnosis. Therefore, rapid, sensitive, and field‑deployable diagnostic methods are essential for effective disease surveillance and control in the cattle industry. METHODS: In this study, we report for the first time the establishment of a TaqMan‑based real‑time quantitative PCR (qPCR) assay that enables simultaneous detection of these five bovine viruses. Multiple sequence alignment of conserved genomic regions was performed, and virus‑specific primers and probes were designed and optimized using Beacon Designer 7 software. Subsequently, a TaqMan‑based multiplex real‑time qPCR assay was established for simultaneous detection of BVDV, AKAV, BNoV, BEV, and BCoV. The established detection method was applied to 200 clinical samples collected from 10 farms in multiple regions of Jilin Province. RESULTS: The results showed that the detection rates for BVDV, AKAV, BNoV, BEV, and BCoV were 33.50%, 0.50%, 4.50%, 7.50%, and 12.00%, respectively. Mixed infections were detected in 9 samples co‑infected with two of the five pathogens, with an overall mixed infection rate of 4.50%. Compared with conventional PCR, coincidence rates were 100% for BVDV, AKAV, BNoV, BEV, and BCoV. DISCUSSION: These findings indicate that the TaqMan multiplex real‑time qPCR assay developed here demonstrates favorable specificity, sensitivity, and reproducibility. This assay enables efficient detection and surveillance of bovine viruses, offering a reliable technical tool for the diagnosis and control of corresponding viral diseases in cattle.

Akabane virus (AKAV)↗

Hybrid 'Sinta' papaya exhibits unique ACC synthase 1 cDNA isoforms.

Five ripening-related ACC synthase cDNA isoforms were cloned from 80% ripe papaya cv. 'Sinta' by reverse transcription-PCR using gene-specific primers. Clone 2 had the longest transcript and contained all common exons and three alternative exons. Clones 3 and 4 contained common exons and one alternative exon each, while clone 1, the most common transcript, contained only the common exons. Clone 5 could be due to cloning artifacts and might not be a unique cDNA fragment. Thus, there are only four isoforms of ACC synthase mRNA. Southern blot analysis indicates that all five clones came from only one gene existing as a single copy in the 'Sinta' papaya genome. Multiple sequence alignment indicates that the four isoforms arise from a single gene, possibly through alternative splicing mechanisms. All the putative alternative exons were present at the 5'-end of the gene comprising the N-terminal region of the protein. 'Sinta' ACC synthase cDNAs were of the capacs 1 type and are most closely related to a 1.4 kb capacs 1-type DNA (AJ277160) from Eksotika papaya. No capacs 2-type cDNAs were cloned from 'Sinta' by RT-PCR. This is the first report of possible alternative splicing mechanism in ripening-related ACC synthase genes in hybrid papaya, possibly to modulate or fine-tune gene expression relevant to fruit ripening.

Amino Acid Sequence↗

Mining colon cancer specific alternative splicing in EST database.

Among 75218 splicing sites, 137 colon cancer specific alternative splicing isoforms were found by mining EST database. Alternative splicing database were first constructed by aligning EST to genomic sequence. Numbers of ESTs from normal or cancer colon tissue supporting splicing isoform at each splicing site were then queried and analyzed with Fisher exact test. There were 53 3' splicing, 42 5' splicing, 40 exon skipping and 2 mutual exclusive cancer specific splicing isoforms.

Alternative Splicing↗

Toward an integrated resource for pharmacogenomics (PGx): Survey findings from the genomic medicine communities.

PURPOSE: Pharmacogenomics (PGx) is a critical component of precision health care that aims to improve drug efficacy and reduce adverse events. Terminologies and standards have not always aligned between PGx and broader genomic medicine communities, which is a barrier to PGx implementation. An updated assessment of community barriers, needs, and perspectives is critical to enable more standardized terminologies and interpretation frameworks. METHODS: The Clinical Genome Resource's PGx Interpretation Committee (PGxIC, formerly referred to as the PGx Working Group, PGxWG) conducted 2 surveys targeting the PGx and genomic medicine communities (n = 508) to evaluate perspectives on PGx clinical validity and actionability frameworks, as well as other barriers to PGx implementation. Surveys were tailored toward self-reported familiarity with PGx. Data primarily consisted of free text, which were analyzed using qualitative content analysis methods. RESULTS: Survey responses indicated conflation of terminology across disciplines, including confusion around differing definitions of terms in PGx and non-PGx contexts. Data also indicated broad support for leveraging existing PGx guidelines and framework structures alongside the standardization of approaches and centralization of resources. CONCLUSION: These novel survey results demonstrate broad consensus on the importance of integrating PGx into clinical practice, including support for development of gene-drug response clinical validity and actionability frameworks aligned with Clinical Genome Resource's frameworks for gene-disease relationships.

Humans↗

A physical map of the mouse genome.

A physical map of a genome is an essential guide for navigation, allowing the location of any gene or other landmark in the chromosomal DNA. We have constructed a physical map of the mouse genome that contains 296 contigs of overlapping bacterial clones and 16,992 unique markers. The mouse contigs were aligned to the human genome sequence on the basis of 51,486 homology matches, thus enabling use of the conserved synteny (correspondence between chromosome blocks) of the two genomes to accelerate construction of the mouse map. The map provides a framework for assembly of whole-genome shotgun sequence data, and a tile path of clones for generation of the reference sequence. Definition of the human-mouse alignment at this level of resolution enables identification of a mouse clone that corresponds to almost any position in the human genome. The human sequence may be used to facilitate construction of other mammalian genome maps using the same strategy.

Animals↗

The Legume Information System (LIS): an integrated information resource for comparative legume biology.

The Legume Information System (LIS) (http://www.comparative-legumes.org), developed by the National Center for Genome Resources in cooperation with the USDA Agricultural Research Service (ARS), is a comparative legume resource that integrates genetic and molecular data from multiple legume species enabling cross-species genomic and transcript comparisons. The LIS virtual plant interface allows simplified and intuitive navigation of transcript data from Medicago truncatula, Lotus japonicus, Glycine max and Arabidopsis thaliana. Transcript libraries are represented as images of plant organs in different developmental stages, which are selected to query the analyzed and annotated data. Complex queries can be accomplished by adding modifiers, keywords and sequence names. The LIS also contains annotated genomic data featuring transcript alignments to validate gene predictions as well as motif and similarity analyses. The genomic browser supports comparative analysis via novel dynamic functional annotation comparisons. CMap, developed as part of the GMOD project (http://www.gmod.org/cmap/index.shtml), has been incorporated to support comparative analyses of community linkage and physical map data. LIS is being expanded to incorporate gene expression and biochemical pathways which will be seamlessly integrated forming a knowledge discovery framework.

Arabidopsis↗

Rice bioinformatics. analysis of rice sequence data and leveraging the data to other plant species.

Rice (Oryza sativa) is a model species for monocotyledonous plants, especially for members in the grass family. Several attributes such as small genome size, diploid nature, transformability, and establishment of genetic and molecular resources make it a tractable organism for plant biologists. With an estimated genome size of 430 Mb (Arumuganathan and Earle, 1991), it is feasible to obtain the complete genome sequence of rice using current technologies. An international effort has been established and is in the process of sequencing O. sativa spp. japonica var "Nipponbare" using a bacterial artificial chromosome/P1 artificial chromosome shotgun sequencing strategy. Annotation of the rice genome is performed using prediction-based and homology-based searches to identify genes. Annotation tools such as optimized gene prediction programs are being developed for rice to improve the quality of annotation. Resources are also being developed to leverage the rice genome sequence to partial genome projects such as expressed sequence tag projects, thereby maximizing the output from the rice genome project. To provide a low level of annotation for rice genomic sequences, we have aligned all rice bacterial artificial chromosome/P1 artificial chromosome sequences with The Institute of Genomic Research Gene Indices that are a set of nonredundant transcripts that are generated from nine public plant expressed sequence tag projects (rice, wheat, sorghum, maize, barley, Arabidopsis, tomato, potato, and barrel medic). In addition, we have used data from The Institute of Genomic Research Gene Indices and the Arabidopsis and Rice Genome Projects to identify putative orthologues and paralogues among these nine genomes.

Base Sequence↗

Beyond Blacklists: A Critical Assessment of Exclusion Set Generation Strategies and Alternative Approaches.

Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Tools like the widely used Blacklist software facilitate this process, but their algorithmic details and parameter choices are not always clearly documented, affecting reproducibility and biological relevance. We examined the Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to variability in input data, aligner choice, and read length. We also identified and addressed a coding issue that led to over-annotation of high-signal regions. We further explored the use of "sponge" sequences-unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA-as an alternative approach. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets and suggest that sponge sequences offer a flexible, alignment-guided strategy for reducing artifacts and improving functional genomics analyses.

Journal Article↗

Consensus promoter identification in the human genome utilizing expressed gene markers and gene modeling.

Deciphering the human genome includes locating the promoters that initiate transcription and identifying the exons of genes. Many promoter prediction programs have been proposed, but when they are applied to extended regions of the genome, most of their predictions are false-positives. The extensive collection of gene transcript sequences is an important new source of information, which has not been used previously in promoter predictions. Our approach is to enhance the specificity of predictions by restricting the genomic regions that are searched using gene transcript alignments as anchors in the genome for gene modeling. We developed a consensus promoter prediction method combining previously developed algorithms with the GENSCAN gene modeling program. Our method, CONPRO (CONsensus PROmoter), identifies promoters with very high confidence, and the predicted promoters are guaranteed to be associated with genes. On our test data set, the method correctly detects promoters for approximately half of all human genes (37%-71%), and most predictions are true promoters (85%-90%). Applying our method to the human genome and human genes from the Unigene data set, we find the promoters for 13,744 genes. Of these, 6440 are genes with a functionally cloned mRNA, and 7304 are novel genes for which only expressed sequence tags (ESTs) are available. Candidate promoters for many novel genes will be a useful resource in elucidating complex biological response mechanisms.

5' Untranslated Regions↗