Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Evolutionary analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

[Evolutionary and prognostic studies of patients with primary liver cancer. II. Multifactorial analysis using a stepped-regression mathematical model and graphics].

The subject of this study is the evolution and prognosis of 63 patients with primary liver carcinoma assessed bu multifactor regression mathematical model realized with the aid of the program 2R of the statistical package VMDR and by graphic expression of the functions of survival, mortality, speed of mortality function growth and graphic assessment of survival according to Okuda's method. The step regression analysis in 2 steps of the regression mathematical model of the clinical indices pointed out the factors age, edema and liver encephalopathy. By 7 steps in the regression mathematical model of the combined clinical and clinico-chemical indices as basic prognostic factors were selected: prothrombin time, liver encephalopathy, direct bilirubin, age, GGTP and sex.

Adult↗

Genome-wide characterization of the bZIP gene family in Rattus norvegicus and expression profiling analysis during brain development.

BACKGROUND: The brown rat (Rattus norvegicus) serves as a cornerstone model organism in biomedical research, particularly for understanding physiological homeostasis and stress responses. The basic leucine zipper (bZIP) transcription factor family is a pivotal regulatory network involved in growth, organogenesis, and neurodevelopment. Despite its importance, a systematic characterization of the bZIP gene family in rats has remained elusive. RESULTS: In this study, we performed a genome-wide identification of 61 RnbZIP genes, which were categorized into 10 distinct subfamilies based on phylogenetic relationships and chromosomal localization. Structural analysis revealed conserved motif arrangements within subfamilies, while collinearity analysis identified significant gene duplication events-predominantly tandem and segmental duplications-that have driven the evolutionary expansion of the RnbZIP family. Quantitative analysis showed that members within the same subfamily shared 45%-92% sequence similarity (calculated using the BLOSUM62 scoring matrix), and all duplicated gene pairs underwent strong purifying selection (Ka/Ks&#x2009;<&#x2009;1). Comparative genomics across seven rodent species further underscored the evolutionary conservation and divergence of these factors. Expression profiling across diverse organs and brain developmental stages indicated that RnbZIP genes exhibit high tissue specificity. Notably, 10 candidate genes, including RnbZIP01, RnbZIP02, and RnbZIP08, demonstrated dynamic expression patterns during brain maturation, suggesting their essential roles in neurodevelopmental processes. CONCLUSIONS: Our findings provide a comprehensive structural and evolutionary framework for the RnbZIP gene family, highlighting their potential regulatory functions in rat organogenesis and brain development. This study establishes a valuable resource for further functional characterization of specific bZIP members in mammalian neurological systems.

Animals↗

A new method to model membrane protein structure based on silent amino acid substitutions.

The importance of accurately modeling membrane proteins cannot be overstated, in lieu of the difficulties in solving their structures experimentally. Often, however, modeling procedures (e.g., global searching molecular dynamics) generate several possible candidates rather then pointing to a single model. Herein we present a new approach to select among candidate models based on the general hypothesis that silent amino acid substitutions, present in variants identified from evolutionary conservation data or mutagenesis analysis, do not affect the stability of a native structure but may destabilize the non-native structures also found. The proof of this hypothesis has been tested on the alpha-helical transmembrane domains of two homodimers, human glycophorin A and human CD3-zeta, a component of the T-cell receptor. For both proteins, only one structure was identified using all the variants. For glycophorin A, this structure is virtually identical to the structure determined experimentally by NMR. We present a model for the transmembrane domain of CD3-zeta that is consistent with predictions based on mutagenesis, homology modeling, and the presence of a disulfide bond. Our experiments suggest that this method allows the prediction of transmembrane domain structure based only on widely available evolutionary conservation data.

Amino Acid Substitution↗

CADp44: a novel regulatory subunit of the 26S proteasome and the mammalian homolog of yeast Sug2p.

We have identified a novel protein, CADp44, based on the analysis of cDNAs derived from the brainstem of the 13-lined ground squirrel, Spermophilus tridecemlineatus. CADp44 has an unmodified molecular mass of 44,178 Da and contains multiple functional domains, including a conserved ATPase domain (CAD) and a leucine zipper motif. We show that distinct regions of the CADp44 sequence are identical to a set of peptides prepared from a recently identified bovine protein, referred to as p42, which is found in the PA700 regulatory complex of the 26S proteasome (DeMartino et al., 1996). We also show that CADp44 is the functional homolog of the newly characterized Sug2 protein from the budding yeast, Saccharomyces cerevisiae (Russell et al., 1996). Consistent with its role as a component of the 26S proteasome, CADp44 mRNA is found in all ground squirrel tissues examined. Evolutionary relationships based on sequence analysis show that both CADp44 and yeast Sug2p are distinct from the other five CAD ATPases found in the PA700, and together comprise the sixth and newest CAD subunit of the regulatory complex of the 26S proteasome.

ATPases Associated with Diverse Cellular Activitie↗

Transcription analysis of Streptococcus thermophilus phages in the lysogenic state.

The transcription of prophage genes was studied in two lysogenic Streptococcus thermophilus cells by Northern blot and primer-extension experiments. In the lysogen containing the cos-site phage Sfi21 only two gene regions of the prophage were transcribed. Within the lysogeny module an 1.6-kb-long mRNA started at the promoter of the phage repressor gene and covered also the next two genes, including a superinfection exclusion (sie) gene. A second, quantitatively more prominent 1-kb-long transcript was initiated at the promoter of the sie gene. Another prophage transcript of 1.6-kb length covered a group of genes without database matches that were located between the lysin gene and the right attachment site. The rest of the prophage genome was transcriptionally silent. A very similar transcription pattern was observed for a S. thermophilus lysogen containing the pac-site phage O1205 as a prophage. Prophages from pathogenic streptococci encode virulence genes downstream of the lysin gene. We speculate that temperate phages from lactic streptococci also encode nonessential phage genes ("lysogenic conversion genes") in this region that increase the ecological fitness of the lysogen to further their own evolutionary success. A comparative genome analysis revealed that many temperate phages from low GC content Gram-positive bacteria encode a variable number of genes in that region and none was linked to known phage-related function. Prophages from pathogenic streptococci encode toxin genes in this region. In accordance with theoretical predictions on prophage-host genome interactions a prophage remnant was detected in S. thermophilus that had lost most of the prophage DNA while transcribed prophage genes were spared from the deletion process.

Base Sequence↗

Analysis of the type 1 pilin gene cluster fim in Salmonella: its distinct evolutionary histories in the 5' and 3' regions.

The type 1 pilin encoded by fim is present in both Escherichia coli and Salmonella natural isolates, but several lines of evidence indicate that similarities at the fim locus may be an example of independent acquisition rather than common ancestry. For example, the fim gene cluster is found at different chromosomal locations and with distinct gene orders in these closely related species. In this work we examined the fim gene cluster of Salmonella, the genes of which show high nucleotide sequence divergence from their E. coli counterparts, as well as a different G+C content and codon usage. DNA hybridization analysis revealed that, among the salmonellae, the fim gene cluster is present in all isolates of S. enterica but is absent from S. bongori. Molecular phylogenetic analyses of the fimA and fimI genes yield an estimate of phylogeny that is in satisfactory congruence with housekeeping and other virulence genes examined in this species. In contrast, phylogenetic analyses of the fimZ, fimY, and fimW genes indicate that horizontal transfer of this region has occurred more than once. There is also size variation in the fimZ, fimY, and fimW intergenic regions in the 3' region, and these genes are absent in isolate S2983 of subspecies IIIa. Interestingly, the G+C contents of the fimZ, fimY, and fimW genes are less than 46%, which is considerably lower than those of the other six genes of the fim cluster. This study demonstrates that horizontal transmission of all or part of the same gene cluster can occur repeatedly, with the result that different regions of a single gene cluster may have different evolutionary histories.

Adhesins, Escherichia coli↗

wgatools: an ultrafast toolkit for manipulating whole-genome alignments.

SUMMARY: With the rapid development of long-read sequencing technologies, the era of individual complete genomes is approaching. We have developed wgatools, a cross-platform, ultrafast toolkit that supports a range of whole-genome alignment formats, offering practical tools for conversion, processing, evaluation, and visualization of alignments, thereby facilitating population-level genome analysis and advancing functional and evolutionary genomics. AVAILABILITY AND IMPLEMENTATION: wgatools supports diverse formats and can process, filter, and statistically evaluate alignments, perform alignment-based variant calling, and visualize alignments both locally and genome-wide. Built with Rust for efficiency and safe memory usage, it ensures fast performance and can handle large datasets consisting of hundreds of genomes. wgatools is published as free software under the MIT open-source license, and its source code is freely available at https://github.com/wjwei-handsome/wgatools and https://zenodo.org/records/14882797.

Software↗

Codon usage in Homo sapiens: evidence for a coding pattern on the non-coding strand and evolutionary implications of dinucleotide discrimination.

This study reports the analysis of codon usage in 35 complete Homo sapiens genes. Both codon frequency and inter-codon interference exhibit patterns of evolutionary interest. There is a significant positive correlation between the frequency with which a given codon is used and the frequency with which its complement is used. Since the frequency of appearance of the complementary codon on the coding strand is equal to the frequency of appearance of the original codon on the non-coding strand, in the same phase, the non-coding strand is found to resemble the coding strand in triplet composition. The same effect has been observed in Escherichia coli. This preference for the use of certain complementary triplets as codons suggests that the evolution of the use of the genetic code depended to some extent upon the double-stranded nature of the coding material. In addition, the effect of discrimination against the use of two dinucleotides, CpG and UpA, is observed in codon usage and also in adjacent codon interference. Codons beginning with G, or A, are unlikely to be preceded by codons ending in C, or U, respectively. Consideration of codon assignment in the genetic code together with the observed CpG infrequency suggests that the evolution of the code may have been influenced by conditions in which the use of CpG dinucleotides was unfavorable. The infrequent use of UpA dinucleotides can be explained as the result of frameshift mutation during gene evolution.

Base Sequence↗

Phenotypic and molecular characterizations of Yersinia pestis isolates from Kazakhstan and adjacent regions.

Recent interest in characterizing infectious agents associated with bioterrorism has resulted in the development of effective pathogen genotyping systems, but this information is rarely combined with phenotypic data. Yersinia pestis, the aetiological agent of plague, has been well defined genotypically on local and worldwide scales using multi-locus variable number tandem repeat analysis (MLVA), with emphasis on evolutionary patterns using old isolate collections from countries where Y. pestis has existed the longest. Worldwide MLVA studies are largely based on isolates that have been in long-term laboratory culture and storage, or on field material from parts of the world where Y. pestis has potentially circulated in nature for thousands of years. Diversity in these isolates suggests that they may no longer represent the wild-type organism phenotypically, including the possibility of altered pathogenicity. This study focused on the phenotypic and genotypic properties of 48 Y. pestis isolates collected from 10 plague foci in and bordering Kazakhstan. Phenotypic characterization was based on diagnostic tests typically performed in reference laboratories working with Y. pestis. MLVA was used to define the genotypic relationships between the central-Asian isolates and a group of North American isolates, and to examine Kazakh Y. pestis diversity according to predefined plague foci and on an intermediate geographical scale. Phenotypic properties revealed that a large portion of this collection lacks one or more plasmids necessary to complete the blocked flea/mammal transmission cycle, has lost Congo red binding capabilities (Pgm-), or both. MLVA analysis classified isolates into previously identified biovars, and in some cases groups of isolates collected within the same plague focus formed a clade. Overall, MLVA did not distinguish unique phylogeographical groups of Y. pestis isolates as defined by plague foci and indicated higher genetic diversity among older biovars.

Animals↗

Modulation of RNA polymerase core functions by basal transcription factor TFB/TFIIB.

The archaeal basal transcriptional machinery consists of TBP (TATA-binding protein), TFB (transcription factor B; a homologue of eukaryotic TFIIB) and an RNA polymerase that is structurally very similar to eukaryotic RNA polymerase II. This constellation of factors is sufficient to assemble specifically on a TATA box-containing promoter and to initiate transcription at a specific start site. We have used this system to study the functional interaction between basal transcription factors and RNA polymerase, with special emphasis on the post-recruitment function of TFB. A bioinformatics analysis of the B-finger of archaeal TFB and eukaryotic TFIIB reveals that this structure undergoes rapid and apparently systematic evolution in archaeal and eukaryotic evolutionary domains. We provide a detailed analysis of these changes and discuss their possible functional implications.

Amino Acid Sequence↗

The secretin G-protein-coupled receptor family: teleost receptors.

Twenty-one members of the secretin family (family 2) of G-protein-coupled receptors (GPCRs) were identified via directed cloning and data-mining of the Fugu Genome Consortium database, representing the most comprehensive description of secretin GPCRs in a teleost fish to date. Duplicated genes were identified for many of the family members, namely the receptors for pituitary adenylate cyclase-activating polypeptide (PACAP)/vasoactive intestinal peptide (VIP), calcitonin, calcitonin gene-related peptide (CGRP), growth hormone releasing hormone (GHRH), glucagon receptor/glucagon-like peptide (GLP) and parathyroid hormone-related peptide (PTHrP)/PTH. Mining of other teleost genomes (zebrafish and Tetraodon) revealed that the duplicated genes identified in the Takifugu genome were also present in these fish. Additional database searching of the Escherichia coli, yeast, Drosophila, Caenorhabditis elegans and Ciona genomes revealed that the family 2 of GPCRs were only present in the multicellular organisms. Orthologues of all the human secretin receptors were identified with the exception of secretin itself. Additional database searches in the Fugu Genome Consortium database also failed to reveal a secretin ligand and so it is hypothesised that both the receptor and the ligand evolved after the divergence of teleost/tetrapod lineages. Phylogenetic analysis at both the protein and the DNA level provided strong support for each of the individual receptor family groupings, but weak support between groups, making evolutionary inferences difficult. A more critical analysis of the PACAP/VIP receptor family confirmed previous hypotheses that the vasoactive intestinal peptide receptor (VPAC(1)R) gene is the ancestral form of the receptor.

Animals↗

Structural and evolutionary comparisons of four alleles of the mouse immunoglobulin kappa chain gene, Igk-VSer.

The mouse Igk-VSer gene encodes an immunoglobulin kappa light chain variable region which gives rise to two phenotypic polymorphisms of mouse kappa chains. The nucleotide sequences of coding and flanking regions of the Igk-VSerc and Igk-VSerd alleles found in recently inbred strains of wild mice are compared with those of the Igk-VSera and Igk-VSerb alleles described previously. Results suggest that the gene is evolving randomly and that framework 2 and complementarity determining region 2 are preserved, presumably for overall light chain structure. Results indicate that all four alleles have an octamer motif upstream of the gene which should be functional and allow prediction of whether or not the product of the germ line gene will be detectable as either the IB-peptide or Ef1a phenotypic polymorphism. Southern hybridization of genomic DNA using as probe a 1-kb Xba I-Xba I fragment located approximately 4 kb upstream of the BALB/c Igk-VSerb coding region demonstrated the presence of homologous DNA in mice bearing the Igk-VSera allele and absence from mice bearing the Igk-VSerc and Igk-VSerd alleles. Nucleotide sequence comparison of BALB/c and SK/CamRk (Igk-VSerd) DNA in this region demonstrated that BALB/c contained an insertion 2.4 kb in length which was absent from SK/CamRk. Both strains contain DNA homologous to the reverse complement of the mouse Bam5 repetitive element at the point of the insertion, with BALB/c containing approximately 70 nucleotides more of the element than SK/CamRk. Surprisingly, the strains containing DNA related to the Xba I-Xba I probe are not those determined to be the most similar by nucleotide sequence comparisons and by the Phylogenetic Analysis Using Parsimony program. The evolutionary relationship of the alleles and a possible basis for the inconsistency presented by the Xba I-Xba I fragment-related DNA are discussed.

Alleles↗

The hotspot conversion paradox and the evolution of meiotic recombination.

Studies of meiotic recombination have revealed an evolutionary paradox. Molecular and genetic analysis has shown that crossing over initiates at specific sites called hotspots, by a recombinational-repair mechanism in which the initiating hotspot is replaced by a copy of its homolog. We have used computer simulations of large populations to show that this mechanism causes active hotspot alleles to be rapidly replaced by inactive alleles, which arise by rare mutation and increase by recombination-associated conversion. Additional simulations solidified the paradox by showing that the known benefits of recombination appear inadequate to maintain its mechanism. Neither the benefits of accurate segregation nor those of recombining flanking genes were sufficient to preserve active alleles in the face of conversion. A partial resolution to this paradox was obtained by introducing into the model an additional, nonmeiotic function for the sites that initiate recombination, consistent with the observed association of hotspots with functional sites in chromatin. Provided selection for this function was sufficiently strong, active hotspots were able to persist in spite of frequent conversion to inactive alleles. However, this explanation is unsatisfactory for two reasons. First, it is unlikely to apply to obligately sexual species, because observed crossover frequencies imply maintenance of many hotspots per genome, and the viability selection needed to preserve these would drive the species to extinction. Second, it fails to explain why such a genetically costly mechanism of recombination has been maintained over evolutionary time. Thus the paradox persists and is likely to be resolved only by significant changes to the commonly accepted mechanism of crossing over.

Computer Simulation↗

Genomic organization and expression of the human mono-ADP-ribosyltransferase ART3 gene.

Here we describe an RT-PCR analysis of mono-ADP-ribosyltransferase 3 (ART3) mRNA expression in macrophages, testis, semen, tonsil, heart and skeletal muscle and the complete gene structure as obtained by sequence alignment of PCR products with a human genomic clone (GenBank accession no. AC112719). Twelve exons (ex1-12) were found to make up the coding region of the gene (one more than previously published). Two prominent classes of ART3 splice variants could be distinguished by the presence or absence of ex2 which encodes most of ART3 protein. Among the ex2-containing mRNA species, the most frequently amplified variant did not include exons 9 to 11, except in skeletal muscle, in which the major splice variant lacked ex10 only. Two different, previously not reported 5' non-translated regions (5' UTRs) were identified, demonstrating the presence of two alternative promoters that we termed palpha and pbeta. Whereas the 5'UTR originating from palpha, was split up into three exons, a single exon represented the 5' UTR of pbeta transcripts. Strikingly, in heart, skeletal muscle and tonsils the upstream promoter palpha was totally inactive and ART3 transcription appears to be driven solely by pbeta. In all other cell types tested, transcription started mainly (if not exclusively) at palpha. Thus, ART3 expression in human cells appears to be governed by a combination of differential splicing and tissue-preferential use of two alternative promoters. This specific use is evolutionary conserved as shown by analysis of the 5' UTR of the mouse ART3 mRNA.

5' Untranslated Regions↗

The crocodilian mitochondrial control region: general structure, conserved sequences, and evolutionary implications.

We present the first comprehensive analysis of the crocodilian control region. We have analyzed sequences from all three families of Crocodylia (Crocodylidae, Gavialidae, Alligatoridae), incorporating all genera except Paleosuchus and Melanosuchus. Within the control region of other vertebrates, several sequence motifs and their order appear to be conserved. Herein, we compare aligned crocodilian D-loop sequences to homologous sequences from other vertebrates ranging from fish to birds. Among other findings, we have discovered that while domain I tends to be shorter than the same region in mammals and birds, it contains sequences similar in structure to both the goose-hairpin and termination associated sequences (TAS). Domain II is highly conservative with regard to size among the taxa examined and contains several of the conserved sequence boxes characterized in other vertebrates. Domain III contains several interesting sequence motifs including tandemly repeated sequences, a long poly-A region in the Crocodylidae, and possible bidirection promoter sequences.

Alligators and Crocodiles↗

Current approaches to whole genome phylogenetic analysis.

It has long been known that evolutionary trees (phylogenies) can be estimated by comparing the DNA or protein sequences of homologous genes across different organisms. More recently, attempts have been made to estimate phylogenies by comparing entire genomes. These attempts have focused largely on comparisons of gene content and gene order. Many different methods have been proposed for making these comparisons. These include primarily maximum parsimony and distance methods, although more recently maximum likelihood and Bayesian methods are being developed. This paper discusses each of these approaches in turn, including their merits and limitations, and any software which is available to make use of them.

Evolution, Molecular↗

The RNA polymerase III-dependent family of genes in hemiascomycetes: comparative RNomics, decoding strategies, transcription and evolutionary implications.

We present the first comprehensive analysis of RNA polymerase III (Pol III) transcribed genes in ten yeast genomes. This set includes all tRNA genes (tDNA) and genes coding for SNR6 (U6), SNR52, SCR1 and RPR1 RNA in the nine hemiascomycetes Saccharomyces cerevisiae, Saccharomyces castellii, Candida glabrata, Kluyveromyces waltii, Kluyveromyces lactis, Eremothecium gossypii, Debaryomyces hansenii, Candida albicans, Yarrowia lipolytica and the archiascomycete Schizosaccharomyces pombe. We systematically analysed sequence specificities of tRNA genes, polymorphism, variability of introns, gene redundancy and gene clustering. Analysis of decoding strategies showed that yeasts close to S.cerevisiae use bacterial decoding rules to read the Leu CUN and Arg CGN codons, in contrast to all other known Eukaryotes. In D.hansenii and C.albicans, we identified a novel tDNA-Leu (AAG), reading the Leu CUU/CUC/CUA codons with an unusual G at position 32. A systematic 'p-distance tree' using the 60 variable positions of the tRNA molecule revealed that most tDNAs cluster into amino acid-specific sub-trees, suggesting that, within hemiascomycetes, orthologous tDNAs are more closely related than paralogs. We finally determined the bipartite A- and B-box sequences recognized by TFIIIC. These minimal sequences are nearly conserved throughout hemiascomycetes and were satisfactorily retrieved at appropriate locations in other Pol III genes.

Ascomycota↗