Search PubMed⌕ Search

Biomedical subjects

Marco Marra

Publications and source records attributed to Marco Marra.

14 recordsLinked to original sources

Cross-Platform Methylation-Based Site of Origin Classification for Squamous Cell Carcinomas.

Squamous cell carcinomas (SCCs) are one of the most common cancer types and can arise at nearly any anatomic site. Because SCCs are one of the most common metastases, do not have reliable site-specific morphologic or genomic features, and have considerable morphologic and immunohistochemical overlap with urothelial carcinomas, distinguishing between primary and metastatic squamous-appearing tumors can be challenging. This distinction can be critical to clinical management. We present Squamous cell carcinoma Methylation for Origin Site (SquaMOS), a methylation-based classifier to predict site of origin of squamous-appearing carcinomas. Trained on publicly available array-based methylation data from 1062 primary SCCs (from lung, head and neck, cervix, and esophagus) and urothelial carcinomas, SquaMOS predicted site of origin in primary tumors with 96.1% accuracy in an internal test set (n = 458) and 97.4% accuracy in an external test set from 3 institutions (n = 78). On metastatic tumors (n = 51), SquaMOS predictions were 96.1% accurate. SquaMOS was directly applicable to shallow Nanopore sequencing data (CpG probe site coverage, 0.25-2.88×) with an accuracy of 91.7% (n = 36; 100% accurate for high-confidence predictions). When tested on SCCs outside the training set types (n = 15, including 3 metastases to lung), no cases were misclassified as of lung origin, supporting accuracy of lung vs nonlung origin classification for diverse SCC types. Overall, we demonstrate highly accurate performance of the SquaMOS classifier on primary and metastatic tumors from multiple data sources, robust to suboptimal tumor purity. We illustrate transferability of our array-based classifier to low-depth Nanopore sequencing data, a potentially rapid means of site of origin determination in a clinical setting.

Humans↗

Identification of novel lung genes in bronchial epithelium by serial analysis of gene expression.

A description of the transcriptome of human bronchial epithelium should provide a basis for studying lung diseases, including cancer. We have deduced global gene expression profiles of bronchial epithelium and lung parenchyma, based on a vast dataset of nearly two million sequence tags from 21 serial analysis of gene expression (SAGE) libraries from individuals with a history of smoking. Our analysis suggests that the transcriptome of the bronchial epithelium is distinct from that of lung parenchyma and other tissue types. Moreover, our analysis has identified novel bronchial-enriched genes such as MS4A8B, and has demonstrated the use of SAGE for the discovery of novel transcript variants. Significantly, gene expression associated with ciliogenesis is evident in bronchial epithelium, and includes the expression of transcripts specifying axonemal proteins DNAI2, SPAG6, ASP, and FOXJ1 transcription factor. Moreover, expression of potential regulators of ciliogenesis such as MDAC1, NYD-SP29, ARMC3, and ARMC4 were also identified. This study represents a comprehensive delineation of the bronchial and parenchyma transcriptomes, identifying more than 20,000 known and hypothetical genes expressed in the human lung, and constitutes one of the largest human SAGE studies reported to date.

Aged↗

Mating factor linkage and genome evolution in basidiomycetous pathogens of cereals.

Sex in basidiomycete fungi is controlled by tetrapolar mating systems in which two unlinked gene complexes determine up to thousands of mating specificities, or by bipolar systems in which a single locus (MAT) specifies different sexes. The genus Ustilago contains bipolar (Ustilago hordei) and tetrapolar (Ustilago maydis) species and sexual development is associated with infection of cereal hosts. The U. hordei MAT-1 locus is unusually large (approximately 500 kb) and recombination is suppressed in this region. We mapped the genome of U. hordei and sequenced the MAT-1 region to allow a comparison with mating-type regions in U. maydis. Additionally the rDNA cluster in the U. hordei genome was identified and characterized. At MAT-1, we found 47 genes along with a striking accumulation of retrotransposons and repetitive DNA; the latter features were notably absent from the corresponding U. maydis regions. The tetrapolar mating system may be ancestral and differences in pathogenic life style and potential for inbreeding may have contributed to genome evolution.

Edible Grain↗

Generation, annotation, analysis and database integration of 16,500 white spruce EST clusters.

BACKGROUND: The sequencing and analysis of ESTs is for now the only practical approach for large-scale gene discovery and annotation in conifers because their very large genomes are unlikely to be sequenced in the near future. Our objective was to produce extensive collections of ESTs and cDNA clones to support manufacture of cDNA microarrays and gene discovery in white spruce (Picea glauca [Moench] Voss). RESULTS: We produced 16 cDNA libraries from different tissues and a variety of treatments, and partially sequenced 50,000 cDNA clones. High quality 3' and 5' reads were assembled into 16,578 consensus sequences, 45% of which represented full length inserts. Consensus sequences derived from 5' and 3' reads of the same cDNA clone were linked to define 14,471 transcripts. A large proportion (84%) of the spruce sequences matched a pine sequence, but only 68% of the spruce transcripts had homologs in Arabidopsis or rice. Nearly all the sequences that matched the Populus trichocarpa genome (the only sequenced tree genome) also matched rice or Arabidopsis genomes. We used several sequence similarity search approaches for assignment of putative functions, including blast searches against general and specialized databases (transcription factors, cell wall related proteins), Gene Ontology term assignation and Hidden Markov Model searches against PFAM protein families and domains. In total, 70% of the spruce transcripts displayed matches to proteins of known or unknown function in the Uniref100 database (blastx e-value < 1e-10). We identified multigenic families that appeared larger in spruce than in the Arabidopsis or rice genomes. Detailed analysis of translationally controlled tumour proteins and S-adenosylmethionine synthetase families confirmed a twofold size difference. Sequences and annotations were organized in a dedicated database, SpruceDB. Several search tools were developed to mine the data either based on their occurrence in the cDNA libraries or on functional annotations. CONCLUSION: This report illustrates specific approaches for large-scale gene discovery and annotation in an organism that is very distantly related to any of the fully sequenced genomes. The ArboreaSet sequences and cDNA clones represent a valuable resource for investigations ranging from plant comparative genomics to applied conifer genetics.

Arabidopsis↗

The genome of the kinetoplastid parasite, Leishmania major.

Leishmania species cause a spectrum of human diseases in tropical and subtropical regions of the world. We have sequenced the 36 chromosomes of the 32.8-megabase haploid genome of Leishmania major (Friedlin strain) and predict 911 RNA genes, 39 pseudogenes, and 8272 protein-coding genes, of which 36% can be ascribed a putative function. These include genes involved in host-pathogen interactions, such as proteolytic enzymes, and extensive machinery for synthesis of complex surface glycoconjugates. The organization of protein-coding genes into long, strand-specific, polycistronic clusters and lack of general transcription factors in the L. major, Trypanosoma brucei, and Trypanosoma cruzi (Tritryp) genomes suggest that the mechanisms regulating RNA polymerase II-directed transcription are distinct from those operating in other eukaryotes, although the trypanosomatids appear capable of chromatin remodeling. Abundant RNA-binding proteins are encoded in the Tritryp genomes, consistent with active posttranscriptional regulation of gene expression.

Animals↗

Serial analysis of gene expression reveals conserved links between protein kinase A, ribosome biogenesis, and phosphate metabolism in Ustilago maydis.

The switch from budding to filamentous growth is a key aspect of invasive growth and virulence for the fungal phytopathogen Ustilago maydis. The cyclic AMP (cAMP) signaling pathway regulates dimorphism in U. maydis, as demonstrated by the phenotypes of mutants with defects in protein kinase A (PKA). Specifically, a mutant lacking the regulatory subunit of PKA encoded by the ubc1 gene displays a multiple-budded phenotype and fails to incite disease symptoms, although proliferation does occur in the plant host. A mutant with a defect in a catalytic subunit of PKA, encoded by adr1, has a constitutively filamentous phenotype and is nonpathogenic. We employed serial analysis of gene expression to examine the transcriptomes of a wild-type strain and the ubc1 and adr1 mutants to further define the role of PKA in U. maydis. The mutants displayed changes in the transcript levels for genes encoding ribosomal proteins, genes regulated by the b mating-type proteins, and genes for metabolic functions. Importantly, the ubc1 mutant displayed elevated transcript levels for genes involved in phosphate acquisition and storage, thus revealing a connection between cAMP and phosphate metabolism. Further experimentation indicated a phosphate storage defect and elevated acid phosphatase activity for the ubc1 mutant. Elevated phosphate levels in culture media also enhanced the filamentous growth of wild-type cells in response to lipids, a finding consistent with PKA regulation of morphogenesis in U. maydis. Overall, these findings extend our understanding of cAMP signaling in U. maydis and reveal a link between phosphate metabolism and morphogenesis.

Acid Phosphatase↗

Genome sequence of the Brown Norway rat yields insights into mammalian evolution.

The laboratory rat (Rattus norvegicus) is an indispensable tool in experimental medicine and drug development, having made inestimable contributions to human health. We report here the genome sequence of the Brown Norway (BN) rat strain. The sequence represents a high-quality 'draft' covering over 90% of the genome. The BN rat sequence is the third complete mammalian genome to be deciphered, and three-way comparisons with the human and mouse genomes resolve details of mammalian evolution. This first comprehensive analysis includes genes and proteins and their relation to human disease, repeated sequences, comparative genome-wide studies of mammalian orthologous chromosomal regions and rearrangement breakpoints, reconstruction of ancestral karyotypes and the events leading to existing species, rates of variation, and lineage-specific and lineage-independent evolutionary events such as expansion of gene families, orthology relations and protein evolution.

Animals↗

Integrated and sequence-ordered BAC- and YAC-based physical maps for the rat genome.

As part of the effort to sequence the genome of Rattus norvegicus, we constructed a physical map comprised of fingerprinted bacterial artificial chromosome (BAC) clones from the CHORI-230 BAC library. These BAC clones provide approximately 13-fold redundant coverage of the genome and have been assembled into 376 fingerprint contigs. A yeast artificial chromosome (YAC) map was also constructed and aligned with the BAC map via fingerprinted BAC and P1 artificial chromosome clones (PACs) sharing interspersed repetitive sequence markers with the YAC-based physical map. We have annotated 95% of the fingerprint map clones in contigs with coordinates on the version 3.1 rat genome sequence assembly, using BAC-end sequences and in silico mapping methods. These coordinates have allowed anchoring 358 of the 376 fingerprint map contigs onto the sequence assembly. Of these, 324 contigs are anchored to rat genome sequences localized to chromosomes, and 34 contigs are anchored to unlocalized portions of the rat sequence assembly. The remaining 18 contigs, containing 54 clones, still require placement. The fingerprint map is a high-resolution integrative data resource that provides genome-ordered associations among BAC, YAC, and PAC clones and the assembled sequence of the rat genome.

Animals↗

A physical map of the mouse genome.

A physical map of a genome is an essential guide for navigation, allowing the location of any gene or other landmark in the chromosomal DNA. We have constructed a physical map of the mouse genome that contains 296 contigs of overlapping bacterial clones and 16,992 unique markers. The mouse contigs were aligned to the human genome sequence on the basis of 51,486 homology matches, thus enabling use of the conserved synteny (correspondence between chromosome blocks) of the two genomes to accelerate construction of the mouse map. The map provides a framework for assembly of whole-genome shotgun sequence data, and a tile path of clones for generation of the reference sequence. Definition of the human-mouse alignment at this level of resolution enables identification of a mouse clone that corresponds to almost any position in the human genome. The human sequence may be used to facilitate construction of other mammalian genome maps using the same strategy.

Animals↗

Transferrin receptor 2 (TfR2) and HFE mutational analysis in non-C282Y iron overload: identification of a novel TfR2 mutation.

Hereditary hemochromatosis (HH) is classically associated with a Cys282Tyr (C282Y) mutation of the HFE gene. Non-C282Y HH is a heterogeneous group accounting for 15% of HH in Northern Europe. Pathogenic mutations of the transferrin receptor 2 (TfR2) gene have been identified in 4 Italian pedigrees with the latter syndrome. The goal of this study was to perform a mutational analysis of the TfR2 and HFE genes in a cohort of non-C282Y iron overload patients of mixed ethnic backgrounds. Several sequence variants were identified within the TfR2 gene, including a homozygous missense change in exon 17, c2069 A-->C, which changes a glutamine to a proline residue at position 690. This putative mutation was found in a severely affected Portuguese man and 2 family members with the same genotype. In summary, pathologic TfR2 mutations are present outside of Italy, accounting for a small proportion of non-C282Y HH.

Canada↗

Temperature-regulated transcription in the pathogenic fungus Cryptococcus neoformans.

The basidiomycete fungus Cryptococcus neoformans is an opportunistic pathogen of worldwide importance that causes meningitis, leading to death in immunocompromised individuals. Unlike many basidiomycete fungi, C. neoformans is thermotolerant, and its ability to grow at 37 degrees C is considered to be a virulence factor. We used serial analysis of gene expression (SAGE) to characterize the transcriptomes of C. neoformans strains that represent two varieties with different polysaccharide capsule serotypes. These include a serotype D strain of the C. neoformans variety neoformans and a serotype A strain of variety grubii. In this report, we describe the construction and characterization of SAGE libraries from each strain grown at 25 degrees C and 37 degrees C. The SAGE data reveal transcriptome differences between the two strains, even at this early stage of analysis, and identify sets of genes with higher transcript levels at 25 degrees C or 37 degrees C. Notably, growth at the lower temperature increased transcript levels for histone genes, indicating a general influence of temperature on chromatin structure. At 37 degrees C, we noted elevated transcript levels for several genes encoding heat shock proteins and translation machinery. Some of these genes may play a role in temperature-regulated phenotypes in C. neoformans, such as the adaptation of the fungus to growth in the host and the dimorphic transition between budding and filamentous growth. Overall, this work provides the most comprehensive gene expression data available for C. neoformans; this information will be a critical resource both for gene discovery and genome annotation in this pathogen.

Blotting, Northern↗