Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Rice LSD1-like Genes: Genome-Wide Characterization and Evidence Linking OsLSD3 to Plant Height.

LSD1-like zinc-finger proteins participate in programmed cell death, redox homeostasis, and stress responses in plants, but their functional diversification and contributions to agronomic variation in rice remain poorly defined. This study aimed to characterize the rice LSD1-like gene family and evaluate the potential agronomic roles of selected members, with particular emphasis on OsLSD3. Genome-wide analyses were integrated with OsLSD3 natural variation and haplotype analyses in 4666 rice accessions, CRISPR/Cas9 mutant phenotyping in the ZH11 background, and subcellular localization assays. Seven LSD1-like genes were identified and showed substantial divergence in protein architecture, gene organization, promoter cis-element profiles, and tissue- and stress-responsive expression. OsLSD3 formed six population-structured haplotypes, and two common Japonica haplotypes differed significantly in plant height. Consistently, two independent oslsd3 mutant lines were taller than the wild type, whereas additional changes in grain-related traits were line-specific. OsLSD2 and OsLSD3 localized mainly to the nucleus, while OsLSD4 was predominantly nuclear with weak cytoplasmic localization. These results identify OsLSD3 as the strongest candidate among the examined members for further investigation of plant height- and grain-related traits, while OsLSD2 and OsLSD4 represent additional candidates for grain-trait regulation. Further validation using additional alleles and environments is required.

LSD1-like↗

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning↗

Characterization of genomic organization of the adenosine A2A receptor gene by molecular and bioinformatics analyses.

The adenosine A(2A) receptor (A(2A)R) is abundantly expressed in brain and emerging as an important therapeutic target for Parkinson's disease and potentially other neuropsychiatric disorders. To understand the molecular mechanisms of A(2A)R gene expression, we have characterized the genomic organization of the mouse and human A(2A)R genes by molecular and bioinformatic analyses. Three new exons (m1A, m1B and m1C) encoding the 5' untranslated regions (5'-UTRs) of mouse A(2A)R mRNA were identified by rapid amplification of 5' cDNA end (5' RACE), RT-PCR analysis and genome sequence analyses. Similar bioinformatics analysis also suggested six variants of the non-coding "exon 1" (h1A, h1B, h1C, h1D, h1E and h1F) in the human A(2A)R gene, which were confirmed by RT-PCR analysis, while three of the human exon 1 variants (h1D, h1E and h1F) were likewise verified by 5' oligonucleotide capping analysis suggesting multiple transcription start sites. Importantly, RT-PCR and quantitative PCR analysis demonstrated that the A(2A)R transcripts with different exon 1 variants displayed tissue-specific expression patterns. For instance, the mouse exon m1A mRNA was detected only in brain (specifically striatum) and the human exon h1D mRNA in lymphoreticular system. Furthermore, the determination of the three new transcription start sites of human A(2A)R gene by 5' oligonucleotide capping and bioinformatics analyses led to the identification of three corresponding promoter regions which contain several important cis elements, providing additional target for further molecular dissection of A(2A)R gene expression. Finally, our analysis indicates that A(2A)R mRNA and a novel transcript partially overlapping with the 3' exon h3, but in opposite orientation to the A(2A)R gene, could conceivably form duplexes to mutually regulate transcript expression. Thus, combined molecular and bioinformatics analyses revealed a new A(2A)R genomic structure, with conserved coding exons 2 and 3 and divergent, tissue-specific exon 1 variants encoding for 5'-UTR. This raises the possibility of generating multiple tissue-specific A(2A)R mRNA species by alternative promoters with varying regulatory susceptibility.

Animals↗

The complexity of the mammalian transcriptome.

A comprehensive understanding of protein and regulatory networks is strictly dependent on the complete description of the transcriptome of cells. After the determination of the genome sequence of several mammalian species, gene identification is based on in silico predictions followed by evidence of transcription. Conservative estimates suggest that there are about 20,000 protein-encoding genes in the mammalian genome. In the last few years the combination of full-length cDNA cloning, cap-analysis gene expression (CAGE) tag sequencing and tiling arrays experiments have unveiled unexpected additional complexities in the transcriptome. Here we describe the current view of the mammalian transcriptome focusing on transcripts diversity, the growing non-coding RNA world, the organization of transcriptional units in the genome and promoter structures. In-depth analysis of the brain transcriptome has been challenging due to the cellular complexity of this organ. Here we present a computational analysis of CAGE data from different regions of the central nervous system, suggesting distinctive mechanisms of brain-specific transcription.

Alternative Splicing↗

A novel approach for nontargeted data analysis for metabolomics. Large-scale profiling of tomato fruit volatiles.

To take full advantage of the power of functional genomics technologies and in particular those for metabolomics, both the analytical approach and the strategy chosen for data analysis need to be as unbiased and comprehensive as possible. Existing approaches to analyze metabolomic data still do not allow a fast and unbiased comparative analysis of the metabolic composition of the hundreds of genotypes that are often the target of modern investigations. We have now developed a novel strategy to analyze such metabolomic data. This approach consists of (1) full mass spectral alignment of gas chromatography (GC)-mass spectrometry (MS) metabolic profiles using the MetAlign software package, (2) followed by multivariate comparative analysis of metabolic phenotypes at the level of individual molecular fragments, and (3) multivariate mass spectral reconstruction, a method allowing metabolite discrimination, recognition, and identification. This approach has allowed a fast and unbiased comparative multivariate analysis of the volatile metabolite composition of ripe fruits of 94 tomato (Lycopersicon esculentum Mill.) genotypes, based on intensity patterns of >20,000 individual molecular fragments throughout 198 GC-MS datasets. Variation in metabolite composition, both between- and within-fruit types, was found and the discriminative metabolites were revealed. In the entire genotype set, a total of 322 different compounds could be distinguished using multivariate mass spectral reconstruction. A hierarchical cluster analysis of these metabolites resulted in clustering of structurally related metabolites derived from the same biochemical precursors. The approach chosen will further enhance the comprehensiveness of GC-MS-based metabolomics approaches and will therefore prove a useful addition to nontargeted functional genomics research.

Automation↗

A human DNA segment encompassing leucine and methionine tRNA pseudogenes localized on chromosome 6.

A human genomic clone, designated LHtlm8, that strongly hybridized to a mammalian leucine tRNA(IAG) probe, was found to encompass a pair of tRNA pseudogenes that are transcribed in a homologous cell extract. A leucine tRNA(AAG) pseudogene (TRLP1) is 2.1-kb upstream and of opposite polarity to a methionine elongator tRNA(CAU) pseudogene (TRMEP1). TRLP1 has three nucleotide variations (97% identity) from its cognate leucine tRNA(IAG), while TRMEP1 has a 78% identity with its cognate tRNA. Similar to a number of other eukaryotic tRNA pseudogenes, presumptive precursor tRNA transcripts are generated from the two pseudogenes in vitro, but possibly due to their aberrant and unstable secondary and tertiary structures, no detectable mature tRNA products are observed. The two tRNA pseudogenes are encompassed within a 9.6-kb EcoRI fragment that has been assigned to the chromosomal locus, 6pter-q13, by Southern blot hybridization of human-rodent somatic cell hybrid DNAs with probes derived from the cloned tRNA pseudogenes and flanking sequences. A 4.4-kb EcoRI fragment also harbored in clone LHtlm8 was mapped to human chromosome 11, suggesting that the two EcoRI fragments were inadvertantly ligated together during construction of the genomic library.

Animals↗

Cloning, characterization, and genetics of the juvenile hormone esterase gene from Heliothis virescens.

The gene for juvenile hormone esterase (JHE) was cloned from Heliothis virescens (Lepidoptera: Noctuidae). A genomic library was constructed from embryonic DNA and screened with a homologous N-terminal probe from the JHE cDNA. Five genomic clones were isolated and analyzed by dot blot hybridization using regions of the JHE cDNA as probes. Clone C hybridized to both 5' and 3' probes from the JHE cDNA, suggesting that clone C contains both ends of JHE gene. This was verified by sequencing the ends of the JHE gene from clone C using primers from both the 5' and 3' ends of the JHE cDNA. Additional sequencing and restriction mapping were used to characterize the gene. The gene is c. 8 kb long and contains four introns with consensus intron-exon junctions. One of the introns is relatively large (4 kb) and is situated near the extreme 5' end of the gene. Genetic analysis of RFLP variation in interspecific and intraspecific crosses shows that the JHE locus is single-copy with no closely related paralogs and is autosomally encoded in Heliothis. Therefore the developmental pattern of expression of this gene and the previously documented sequence variation in cDNA clones is not explainable by reference to a JHE gene family with distinct structural loci for the different forms.

Animals↗

Gene by environment QTL mapping through multiple trait analyses in blood pressure salt-sensitivity: identification of a novel QTL in rat chromosome 5.

BACKGROUND: The genetic mechanisms underlying interindividual blood pressure variation reflect the complex interplay of both genetic and environmental variables. The current standard statistical methods for detecting genes involved in the regulation mechanisms of complex traits are based on univariate analysis. Few studies have focused on the search for and understanding of quantitative trait loci responsible for gene x environmental interactions or multiple trait analysis. Composite interval mapping has been extended to multiple traits and may be an interesting approach to such a problem. METHODS: We used multiple-trait analysis for quantitative trait locus mapping of loci having different effects on systolic blood pressure with NaCl exposure. Animals studied were 188 rats, the progenies of an F2 rat intercross between the hypertensive and normotensive strain, genotyped in 179 polymorphic markers across the rat genome. To accommodate the correlational structure from measurements taken in the same animals, we applied univariate and multivariate strategies for analyzing the data. RESULTS: We detected a new quantitative train locus on a region close to marker R589 in chromosome 5 of the rat genome, not previously identified through serial analysis of individual traits. In addition, we were able to justify analytically the parametric restrictions in terms of regression coefficients responsible for the gain in precision with the adopted analytical approach. CONCLUSION: Future work should focus on fine mapping and the identification of the causative variant responsible for this quantitative trait locus signal. The multivariable strategy might be valuable in the study of genetic determinants of interindividual variation of antihypertensive drug effectiveness.

Animals↗

Global transcriptome analysis of the heat shock response of Shewanella oneidensis.

Shewanella oneidensis is an important model organism for bioremediation studies because of its diverse respiratory capabilities. However, the genetic basis and regulatory mechanisms underlying the ability of S. oneidensis to survive and adapt to various environmentally relevant stresses is poorly understood. To define this organism's molecular response to elevated growth temperatures, temporal gene expression profiles were examined in cells subjected to heat stress by using whole-genome DNA microarrays for S. oneidensis. Approximately 15% (n = 711) of the total predicted S. oneidensis genes (n = 4,648) represented on the microarray were significantly up- or downregulated (P < 0.05) over a 25-min period after shift to the heat shock temperature. As expected, the majority of the genes that showed homology to known chaperones and heat shock proteins in other organisms were highly induced. In addition, a number of predicted genes, including those encoding enzymes in glycolysis and the pentose cycle, serine proteases, transcriptional regulators (MerR, LysR, and TetR families), histidine kinases, and hypothetical proteins were induced. Genes encoding membrane proteins were differentially expressed, suggesting that cells possibly alter their membrane composition or structure in response to variations in growth temperature. A substantial number of the genes encoding ribosomal proteins displayed downregulated coexpression patterns in response to heat stress, as did genes encoding prophage and flagellar proteins. Finally, a putative regulatory site with high conservation to the Escherichia coli sigma32-binding consensus sequence was identified upstream of a number of heat-inducible genes.

Bacterial Proteins↗

DNA sequence variation and molecular genotyping of natural killer leukocyte immunoglobulin-like receptor, LILRA3.

Leukocyte immunoglobulin-like receptors (LILRs) resemble killer cell immunoglobulin-like receptors (KIR) in structure and function and the KIR and LILR gene families form the major part of the leukocyte receptor cluster (LRC) of human chromosome 19q13.4. Unlike KIR, the LILR gene clusters do not vary in gene number. However, some individuals lack expression of LILRA3. This null allele has a 6.7-kb deletion, which encompasses the first six translated exons. This haplotype enabled unambiguous direct sequencing of LILRA3 alleles using genomic DNA from individuals heterozygous for the deletion. We have performed nucleotide sequencing of a 2.5-kb region within LILRA3 and identified eight bi-allelic substitutions, four of which were non-synonymous. Two from four previously identified LILRA3 cDNA sequences were confirmed and a further six alleles characterised, of which four will encode unique peptides. At least one of the polymorphic positions identified (encoding residue 84 of the first Ig domain) is likely to directly influence ligand binding. A PCR-SSP molecular genotyping system was developed and used to describe a panel of 172 Caucasoid individuals from South-East England. Six alleles were present in this group but they were unevenly distributed, as three alleles accounted for 88% of the studied chromosomes.

Antigens, CD↗

Molecular analysis of human glycophorin MiIX gene shows a silent segment transfer and untemplated mutation resulting from gene conversion via sequence repeats.

The human glycophorin (HGp) loci that define the red blood cell surface antigens of the MNSs blood group system exhibit considerable allelic variation. Previous studies have identified gene conversion events involving HGpA(alpha) and HGpB(delta) that produced delta-alpha-delta hybrid genes which differ in the location of breakpoints. This report presents the molecular analysis of HGpMilX, the first example of a reverse alpha-delta-alpha hybrid gene that specifies a newly described phenotype of the Miltenberger complex. A novel restriction fragment unique to the HGpMilX gene was detected by Southern blot hybridization. The structure of the genomic region encoding the entire extracellular domain of the MilX protein was determined. Nucleotide sequencing of amplified genomic DNA showed that a silent segment of the HGpB(delta) gene had been transposed to replace the internal part of exon III in the HGpA(alpha) gene, thereby resulting in the formation of the MilX allele with an alpha-delta-alpha configuration. The proximal alpha-delta breakpoint was found to be flanked by a direct repeat of the acceptor splice site, whereas the distal delta-alpha breakpoint was localized to a palindromic region. This DNA rearrangement, with a minimal transfer of 16 templated nucleotides and a single mutation of untemplated adenyl nucleotide, not only created two intraexon hybrid junctions but transactivated the expression of a new stretch of amino acid residues in the MilX protein. Such a segment replacement may have occurred through the directional transfer from one duplex to the other via the mechanism of gene conversion. The occurrence of HGpMilX as another hybrid derived from parts of parent genes underlines the role of the recombinational "hotspot" in the generation of allelic diversity in the glycophorin family.

Amino Acid Sequence↗

Mouse hepatitis virus strain A59 RNA polymerase gene ORF 1a: heterogeneity among MHV strains.

Gene 1, the putative RNA replicase gene of coronaviruses, is expressed via two large overlapping open reading frames (ORF 1a and ORF 1b). We have determined the nucleotide sequence of ORF 1a, encoded within the first 13.7 kb of gene 1, for the coronavirus mouse hepatitis virus strain A59 (MHV-A59). Putative papain-like protease domains, a picornavirus 3C-like protease domain, two hydrophobic domains, and a domain "X" of unknown function, previously identified in other coronaviruses (1-3), are also present in ORF 1a of MHV-A59. Comparison between the ORF 1a sequence of MHV-A59 and the published sequence of the JHM strain of MHV (2) showed a high degree of similarity with the exception of several short regions. We sequenced one region of MHV-JHM that contained an 18 amino acid insertion relative to A59 and four other regions in which the sequences of the two strains differed. The MHV-2 and MHV-3 strains were also sequenced in some of these regions. Our analysis confirmed the presence of only one heterogeneous region in ORF 1a of MHV-A59 and MHV-JHM which is also present in MHV-2. Our findings indicate the need to modify the published sequence of MHV-JHM.

Amino Acid Sequence↗

The b-32 protein from maize endosperm: characterization of genomic sequences encoding two alternative central domains.

As derived from a cDNA clone, the structure of the b-32 protein of Zea mays, a putative regulatory factor of zein expression, has a central acidic region separated by two domains covered by secondary structure motifs. In this work, three b-32 genomic clones were selected from two genomic libraries obtained from the maize inbred lines W64A and A69Y. The nucleotide sequences of the complete coding region of each b-32 gene, as well as long stretches of their 5' and 3' flanking regions, were determined. Introns are not present in the b-32 genomic sequences. Minor variations among the three genes and an earlier reported b-32 cDNA indicates that they constitute a gene family showing a characteristic polymorphism. Such a polymorphism is highly evident in large segments of the upstream regulatory sequences. Interestingly, when compared with cDNA (W64A) or with gene b-32.120 (W64A), the genes b-32.129 (W64A) and b-32.152 (A69Y) show three jumps of the reading frame in the central part of the coding region, resulting in a completely different sequence of the b-32 protein central domain. In all cases, variations in the N- and C-terminal domains account only for microheterogeneity.

Amino Acid Sequence↗

Variation on an Src-like theme.

The modularity of protein architecture and the diversity of protein domains hint at a vast combinatorial richness. But evolution appears to have been relatively conservative about selecting new combinations. When a particular grouping of domains within a polypeptide chain can perform a concerted function, that combination tends to reappear in multiple genomic contexts. In other words, once a molecular solution to a functional problem has emerged, it is reused rather than reinvented.

Evolution, Molecular↗

Geomagnetics and society interact in weekly and broader multiseptans underlying health and environmental integrity.

Evidence for the ubiquity and partial endogenicity of about-weekly (circaseptan) components and multiples and/or submultiples thereof (the multiseptans) accumulates as longer and denser records become available. Often attributed to a mere response to the social schedule, circaseptan components now have been documented to characterize environmental variables related to primarily non-photic solar effects. Plausibly, like circadians, circaseptans are anchored in genomes, from bacteria to humans, via both an internal and external evolution. If so, circaseptans, like circadians, may be found in the absence of a 7-day schedule, whereas the social schedule may play a synchronizing role and be responsible for the detection of prominent weekly variations in population statistics. The wobbliness of multiseptans and other components of some environmental time structures (chronomes) may correspond to the wobbliness of multiseptans found in cardiovascular morbidity statistics. Here, the latter stem primarily, but not exclusively, from an extensive database on the incidence of daily calls for an ambulance in Moscow, Russia from 1979-1981. A modulation of multiseptans and other chronome components of both environmental and biological variables by the about 11-year solar activity cycle (and of other low-frequency signals reviewed elsewhere) may account for prior controversies and scepticism about a variety of non-photic effects on biota. This is notably the case when relatively short series are analyzed without consideration of effects of unassessed long-term variations; this is the task of the new field of chronomics. In the spectral element of the chronomes of geophysical and biospherical variability, there are natural near weeks,apart from any precise 7-day periodicity.

Analysis of Variance↗

Evolution of protein function, from a structural perspective.

The recent growth in structural data, and ensuing analyses, have revealed the structural and functional versatility of protein families. With respect to enzymes, local active-site mutations, variations in surface loops and recruitment of additional domains accommodate the diverse substrate specificities and catalytic activities observed within several superfamilies. Conversely, some functions have more than one structural solution, having evolved independently several times during evolution. Combined with the existence of multi-functional genes, which have arisen by gene recruitment, these phenomena must be considered in the process of genome annotation.

Animals↗

Typing methods to approach Pneumocystis carinii genetic heterogeneity.

The study of the genetic heterogeneity of P. carinii is complicated by the lack of an in vitro culture system, as well as by the likely occurrence of co-infections with several special forms or types in a single host. Karyotyping and multilocus enzyme electrophoresis are useful for studies at the evolutionary level. However, these methods require a large number of cells, which prevents their use for the special form infecting humans. DNA sequence analysis of genomic regions is useful to study P. carinii diversity, both at the evolutionary and epidemiological levels. To type the special form specific to humans, several methods are currently used to detect polymorphism in PCR products of polymorphic regions of the genome: DNA sequencing, type-specific hybridisations, and single-strand conformation polymorphism. All these methods still need evaluation. The frequency of potential co-infections in humans determined by these various methods is different. The differences could be due to methodological problems or to real variations between patient populations, geographical locations and/or prophylaxis regimens. In the future, elucidating the population structure of P. carinii and the frequency of potential co-infections is going to be crucial for a better understanding of its epidemiology, and thus for a better prevention of P. carinii pneumonia in humans.

DNA, Fungal↗

Polymorphism in the procyclic acidic repetitive protein gene family of Trypanosoma brucei.

The expression of procyclic acidic repetitive protein (PARP) by Trypanosoma brucei is strongly induced during the transition of bloodstream form to cultured procyclic trypomastigotes in vitro. The membrane-associated protein is distinguished by a central domain consisting of tandemly repeated glutamate-proline dipeptides. The trypanosome genome contains eight PARP genes, at least four of which are expressed. A minimum of four distinct PARP mRNA species comprises two classes of PARP mRNA, based upon divergent 3' untranslated region sequences, and these mRNAs encode polypeptides that exhibited an inverse relation between molecular weight and isoelectric point. Comparative analysis of PARP gene structure indicated that these polypeptides differ by variation in size of the dipeptide repeat domain. Comparison of PARP genes and polypeptides of three independent T. brucei isolates suggested that PARP is not a homogeneous species but instead represents a family of polymorphic proteins.

Animals↗