Search PubMed⌕ Search

Biomedical subjects

Anuj Kumar

Publications and source records attributed to Anuj Kumar.

At least 19 recordsLinked to original sources

The emergence of putative epistatic mutations and iSNVs in SARS-CoV-2 XBB.1.16 variants linked with alteration in immunogenic determinants.

The SARS-CoV-2 XBB variants have been proposed to evolve towards immune evasion against vaccination or natural infection, which may contribute to higher transmissibility. The XBB.1.16 independently emerged due to accumulation of two important substitutions, E180V and T478R in the spike protein. Its pseudoviral infectivity and evasion of humoral immunity were similar to XBB.1 and XBB.1.5. In March 2023, XBB.1.16 had outcompeted other dominant XBB variants in India, which indicate a potential growth advantage. Here, intra-host single nucleotide variations (iSNV) and mutations were screened in SARS-CoV-2 genomes in closely related individuals at two time points: at symptoms onset, and during recovery. The prominence of putative epistatic iSNVs (E180V, G184V, G252V, D253G, and P521S/T) in XBB.1.16 variants were detected during the recovery phase. E180V exhibits mutational constellations with the G252V and P521T in a subset of samples, and this pattern was also detected in contemporary SARS-CoV-2 genomes. Higher order protein structural predictions suggested that the putative epistatic interactions among E180V, G184V, and G252V, D253G may be associated with S protein folding and structural stability. This study involving genomics and computational analyses highlights the potential role of these putative epistatic interactions in immune evasion, which may have contributed to dominance of XBB variants.

Humans↗

Deciphering the etiology of the 2024 outbreak of undiagnosed febrile illness in Panzi, Democratic Republic of the Congo.

In late 2024, an outbreak of over 400 cases of undiagnosed febrile illness, predominantly presenting as fever and cough, was reported in Panzi Health Zone, southwestern Democratic Republic of the Congo. Here we conducted an epidemiological and laboratory investigation to determine the etiology of the outbreak. Clinical data and specimens were prospectively collected from 108 individuals, of whom 59/108 (54.6%) were female. Children aged <5&#x2009;years were the most affected (47/108, 43.5%); 14/32 (43.7%) were malnourished. Oro/nasopharyngeal swabs from 96/108 individuals were PCR tested; 26 blood samples were sequenced. Plasmodium falciparum was detected in 56/108 (51.8%) individuals. Co-infections were also detected, with influenza A(H1N1)pdm09 virus in 16/56 (28.6%) and severe acute respiratory syndrome coronavirus 2 in 10/56 (17.9%) individuals. No novel pathogens were detected via metagenomics. Our findings suggest that the outbreak was primarily associated with a surge in malaria cases, with concurrent viral respiratory infections. Increasing decentralized laboratory capacity and strengthening broader health systems remain crucial for faster outbreak detection and investigation.

Disease Outbreaks↗

Volatile organic compounds: sampling methods and their worldwide profile in ambient air.

The atmosphere is a particularly difficult analytical system because of the very low levels of substances to be analysed, sharp variations in pollutant levels with time and location, differences in wind, temperature and humidity. This makes the selection of an efficient sampling technique for air analysis a key step to reliable results. Generally, methods for volatile organic compounds sampling include collection of the whole air or preconcentration of samples on adsorbents. All the methods vary from each other according to the sampling technique, type of sorbent, method of extraction and identification technique. In this review paper we discuss various important aspects for sampling of volatile organic compounds by the widely used and advanced sampling methods. Characteristics of various adsorbents used for VOCs sampling are also described. Furthermore, this paper makes an effort to comprehensively review the concentration levels of volatile organic compounds along with the methodology used for analysis, in major cities of the world.

Air Pollutants↗

Organelle DB: an updated resource of eukaryotic protein localization and function.

Organelle DB (http://organelledb.lsi.umich.edu) is a web-accessible relational database presenting a supplemented catalog of organelle-localized proteins and major protein complexes. Since its release in 2004, Organelle DB has grown by 20% to encompass over 30,000 proteins from 138 eukaryotic organisms. Each protein in Organelle DB is presented with its subcellular localization, primary sequence and a detailed description of its function, as available. All records in Organelle DB have been annotated using controlled vocabulary from the Gene Ontology consortium. Protein localization data are inherently visual, and Organelle DB is a significant repository of biological images, housing 1500 micrographs of yeast cells carrying stained proteins. Furthermore, we report here the development of Organelle View, an extension of Organelle DB for the interactive visualization of organelles and subcellular structures in the budding yeast Saccharomyces cerevisiae. Organelle View offers a dimensional representation of a yeast cell; users can search Organelle View for proteins of interest, and the organelles housing these proteins will be highlighted in the cell image. Among other applications, Organelle View may serve as an educational aid engaging introductory biology students through a visually 'fun' interface. Organelle View can be accessed from the Organelle DB home page or directly at http://organelleview.lsi.umich.edu.

Animals↗

Genomic analysis of insertion behavior and target specificity of mini-Tn7 and Tn3 transposons in Saccharomyces cerevisiae.

Transposons are widely employed as tools for gene disruption. Ideally, they should display unbiased insertion behavior, and incorporate readily into any genomic DNA to which they are exposed. However, many transposons preferentially insert at specific nucleotide sequences. It is unclear to what extent such bias affects their usefulness as mutagenesis tools. Here, we examine insertion site specificity and global insertion behavior of two mini-transposons previously used for large-scale gene disruption in Saccharomyces cerevisiae: Tn3 and Tn7. Using an expanded set of insertion data, we confirm that Tn3 displays marked preference for the AT-rich 5 bp consensus site TA[A/T]TA, whereas Tn7 displays negligible target site preference. On a genome level, both transposons display marked non-uniform insertion behavior: certain sites are targeted far more often than expected, and both distributions depart drastically from Poisson. Thus, to compare their insertion behavior on a genome level, we developed a windowed Kolmogorov-Smirnov (K-S) test to analyze transposon insertion distributions in sequence windows of various sizes. We find that when scored in large windows (>300 bp), both Tn3 and Tn7 distributions appear uniform, whereas in smaller windows, Tn7 appears uniform while Tn3 does not. Thus, both transposons are effective tools for gene disruption, but Tn7 does so with less duplication and a more uniform distribution, better approximating the behavior of the ideal transposon.

Base Sequence↗

A systems biology approach to learning autophagy.

With its relevance to our understanding of eukaryotic cell function in the normal and disease state, autophagy is an important topic in modern cell biology; yet, few textbooks discuss autophagy beyond a two- or three-sentence summary. Here, we report an undergraduate/graduate class lesson for the in-depth presentation of autophagy using an active learning approach. By our method, students will work in small groups to solve problems and interpret an actual data set describing genes involved in autophagy. The problem-solving exercises and data set analysis will instill within the students a much greater understanding of the autophagy pathway than can be achieved by simple rote memorization of lecture materials; furthermore, the students will gain a general appreciation of the process by which data are interpreted and eventually formed into an understanding of a given pathway. As the data sets used in these class lessons are largely genomic and complementary in content, students will also understand first-hand the advantage of an integrative or systems biology study: No single data set can be used to define the pathway in full-the information from multiple complementary studies must be integrated in order to recapitulate our present understanding of the pathways mediating autophagy. In total, our teaching methodology offers an effective presentation of autophagy as well as a general template for the discussion of nearly any signaling pathway within the eukaryotic kingdom.

Autophagy↗

Impact of Ni(II), Zn(II) and Cd(II) on biogassification of potato waste.

A study was conducted on anaerobic digestion of potato waste and cattle manure mixture, inoculated with 12% inoculum and diluted to 1:1 substrate water ratio at 37 +/- 1 degrees C. Initially pH of substrate was found to be 4.5 to 5.0. Lime and sodium bicarbonate solutions were employed to adjust the pH to 7.5. Biogas production continued up to 10 and 7 days, when lime and sodium bicarbonate solutions were used to adjust the pH, respectively. Biogassification potential was studied in response to different ratio of waste and cattle manure. Biogas production rate was higher when potato waste and cattle manure were used in 50:50 ratio. Effect of two different concentrations (2.5 and 5.0 ppm) of three heavy metals viz. (Ni (II), Zn (II) and Cd (II)) on anaerobic digestion of substrate (potato waste--cattle manure, 50:50) was studied. At 2.5 ppm, all the three heavy metals increased biogas production rate over the control value. The percentage increase in biogas production over the control was highest by Cd, followed by Ni and Zn. In all the treatments, methane content of biogas increased with increase in time after feeding. Various physico-chemical parameters viz. total solids, total volatile solids, total organic carbon and chemical oxygen demand considerably declined after 7 days of digestion and decline was greater in presence of heavy metals as compared to control. The physico-chemical parameters revealed maximum decrease in the presence of 2.5-ppm concentrations of heavy metals with the substrate. Among all the three heavy metals employed in the study, Cd++ at 2.5 ppm was found to produce maximum biogas production rate. The use of three heavy metals to enhance biogas production from potato and other horticultural waste is discussed.

Animals↗

Organelle DB: a cross-species database of protein localization and function.

To efficiently utilize the growing body of available protein localization data, we have developed Organelle DB, a web-accessible database cataloging more than 25,000 proteins from nearly 60 organelles, subcellular structures and protein complexes in 154 organisms spanning the eukaryotic kingdom. Organelle DB is the first on-line resource devoted to the identification and presentation of eukaryotic proteins localized to organelles and subcellular structures. As such, Organelle DB is a strong resource of data from the human proteome as well as from the major model organisms Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster, Caenorhabditis elegans and Mus musculus. In particular, Organelle DB is a central repository of yeast data, incorporating results--and actual fluorescent imagesfrom ongoing large-scale studies of protein localization in S.cerevisiae. Each protein in Organelle DB is presented with its sequence and, as available, a detailed description of its function; functions were extracted from relevant model organism databases, and links to these databases are provided within Organelle DB. To facilitate data interoperability, we have annotated all protein localizations using vocabulary from the Gene Ontology consortium. We also welcome new data for inclusion in Organelle DB, which may be freely accessed at http://organelledb.lsi.umich.edu.

Animals↗

Teaching systems biology: an active-learning approach.

With genomics well established in modern molecular biology, recent studies have sought to further the discipline by integrating complementary methodologies into a holistic depiction of the molecular mechanisms underpinning cell function. This genomic subdiscipline, loosely termed "systems biology," presents the biology educator with both opportunities and obstacles: The benefit of exposing students to this cutting-edge scientific methodology is manifest, yet how does one convey the breadth and advantage of systems biology while still engaging the student? Here, I describe an active-learning approach to the presentation of systems biology. In graduate classes at the University of Michigan, Ann Arbor, I divided students into small groups and asked each group to interpret a sample data set (e.g., microarray data, two-hybrid data, homology-search results) describing a hypothetical signaling pathway. Mimicking realistic experimental results, each data set revealed a portion of this pathway; however, students were only able to reconstruct the full pathway by integrating all data sets, thereby exemplifying the utility in a systems biology approach. Student response to this cooperative exercise was extremely positive. In total, this approach provides an effective introduction to systems biology appropriate for students at both the undergraduate and graduate levels.

Data Collection↗

Large-scale mutagenesis of the yeast genome using a Tn7-derived multipurpose transposon.

We present here an unbiased and extremely versatile insertional library of yeast genomic DNA generated by in vitro mutagenesis with a multipurpose element derived from the bacterial transposon Tn7. This mini-Tn7 element has been engineered such that a single insertion can be used to generate a lacZ fusion, gene disruption, and epitope-tagged gene product. Using this transposon, we generated a plasmid-based library of approximately 300,000 mutant alleles; by high-throughput screening in yeast, we identified and sequenced 9032 insertions affecting 2613 genes (45% of the genome). From analysis of 7176 insertions, we found little bias in Tn7 target-site selection in vitro. In contrast, we also sequenced 10,174 Tn3 insertions and found a markedly stronger preference for an AT-rich 5-base pair target sequence. We further screened 1327 insertion alleles in yeast for hypersensitivity to the chemotherapeutic cisplatin. Fifty-one genes were identified, including four functionally uncharacterized genes and 25 genes involved in DNA repair, replication, transcription, and chromatin structure. In total, the collection reported here constitutes the largest plasmid-based set of sequenced yeast mutant alleles to date and, as such, should be singularly useful for gene and genome-wide functional analysis.

Amino Acid Sequence↗

A novel mitochondrial protein, Tar1p, is encoded on the antisense strand of the nuclear 25S rDNA.

In eukaryotes, it is widely assumed that genes coding for proteins and structural RNAs do not overlap. Using a transposon-tagging strategy to globally analyze the Saccharomyces cerevisiae genome for expressed genes, we identified multiple insertions in an open reading frame that is contained fully within and transcribed antisense to the 25S rRNA gene in the nuclear rDNA repeat region on Chromosome XII. Expression of this gene, TAR1 (Transcript Antisense to Ribosomal RNA), can be detected at the RNA and protein levels, and the primary sequence of the corresponding 124-amino-acid protein is conserved in several yeast species. Tar1p was found to localize to mitochondria, and overexpression of the protein suppresses the respiration-deficient petite phenotype of a point mutation in mitochondrial RNA polymerase that affects mitochondrial gene expression and mtDNA stability. These findings indicate that coding information for protein and structural RNAs can overlap, raising issues regarding the coevolution of such complex genes, and also suggest that rDNA transcription and mitochondrial function are coordinately regulated in eukaryotic cells.

Amino Acid Sequence↗

Subcellular localization of the yeast proteome.

Protein localization data are a valuable information resource helpful in elucidating eukaryotic protein function. Here, we report the first proteome-scale analysis of protein localization within any eukaryote. Using directed topoisomerase I-mediated cloning strategies and genome-wide transposon mutagenesis, we have epitope-tagged 60% of the Saccharomyces cerevisiae proteome. By high-throughput immunolocalization of tagged gene products, we have determined the subcellular localization of 2744 yeast proteins. Extrapolating these data through a computational algorithm employing Bayesian formalism, we define the yeast localizome (the subcellular distribution of all 6100 yeast proteins). We estimate the yeast proteome to encompass approximately 5100 soluble proteins and >1000 transmembrane proteins. Our results indicate that 47% of yeast proteins are cytoplasmic, 13% mitochondrial, 13% exocytic (including proteins of the endoplasmic reticulum and secretory vesicles), and 27% nuclear/nucleolar. A subset of nuclear proteins was further analyzed by immunolocalization using surface-spread preparations of meiotic chromosomes. Of these proteins, 38% were found associated with chromosomal DNA. As determined from phenotypic analyses of nuclear proteins, 34% are essential for spore viability--a percentage nearly twice as great as that observed for the proteome as a whole. In total, this study presents experimentally derived localization data for 955 proteins of previously unknown function: nearly half of all functionally uncharacterized proteins in yeast. To facilitate access to these data, we provide a searchable database featuring 2900 fluorescent micrographs at http://ygac.med.yale.edu.

Algorithms↗

A question of size: the eukaryotic proteome and the problems in defining it.

We discuss the problems in defining the extent of the proteomes for completely sequenced eukaryotic organisms (i.e. the total number of protein-coding sequences), focusing on yeast, worm, fly and human. (i) Six years after completion of its genome sequence, the true size of the yeast proteome is still not defined. New small genes are still being discovered, and a large number of existing annotations are being called into question, with these questionable ORFs (qORFs) comprising up to one-fifth of the 'current' proteome. We discuss these in the context of an ideal genome-annotation strategy that considers the proteome as a rigorously defined subset of all possible coding sequences ('the orfome'). (ii) Despite the greater apparent complexity of the fly (more cells, more complex physiology, longer lifespan), the nematode worm appears to have more genes. To explain this, we compare the annotated proteomes of worm and fly, relating to both genome-annotation and genome evolution issues. (iii) The unexpectedly small size of the gene complement estimated for the complete human genome provoked much public debate about the nature of biological complexity. However, in the first instance, for the human genome, the relationship between gene number and proteome size is far from simple. We survey the current estimates for the numbers of human genes and, from this, we estimate a range for the size of the human proteome. The determination of this is substantially hampered by the unknown extent of the cohort of pseudogenes ('dead' genes), in combination with the prevalence of alternative splicing. (Further information relating to yeast is available at http://genecensus.org/yeast/orfome)

Animals↗

A small reservoir of disabled ORFs in the yeast genome and its implications for the dynamics of proteome evolution.

We surveyed the sequenced Saccharomyces cerevisiae genome (strain S288C) comprehensively for open reading frames (ORFs) that could encode full-length proteins but contain obvious mid-sequence disablements (frameshifts or premature stop codons). These pseudogenic features are termed disabled ORFs (dORFs). Using homology to annotated yeast ORFs and non-yeast proteins plus a simple region extension procedure, we have found 183 dORFs. Combined with the 38 existing annotations for potential dORFs, we have a total pool of up to 221 dORFs, corresponding to less than approximately 3% of the proteome. Additionally, we found 20 pairs of annotated ORFs for yeast that could be merged into a single ORF (termed a mORF) by read-through of the intervening stop codon, and may comprise a complete ORF in other yeast strains. Focussing on a core pool of 98 dORFs with a verifying protein homology, we find that most dORFs are substantially decayed, with approximately 90% having two or more disablements, and approximately 60% having four or more. dORFs are much more yeast-proteome specific than live yeast genes (having about half the chance that they are related to a non-yeast protein). They show a dramatically increased density at the telomeres of chromosomes, relative to genes. A microarray study shows that some dORFs are expressed even though they carry multiple disablements, and thus may be more resistant to nonsense-mediated decay. Many of the dORFs may be involved in responding to environmental stresses, as the largest functional groups include growth inhibition, flocculation, and the SRP/TIP1 family. Our results have important implications for proteome evolution. The characteristics of the dORF population suggest the sorts of genes that are likely to fall in and out of usage (and vary in copy number) in a strain-specific way and highlight the role of subtelomeric regions in engendering this diversity. Our results also have important implications for the effects of the [PSI+] prion. The dORFs disabled by only a single stop and the mORFs (together totalling 35) provide an estimate for the extent of the sequence population that can be resurrected readily through the demonstrated ability of the [PSI+] prion to cause nonsense-codon read-through. Also, the dORFs and mORFs that we find have properties (e.g. growth inhibition, flocculation, vanadate resistance, stress response) that are potentially related to the ability of [PSI+] to engender substantial phenotypic variation in yeast strains under different environmental conditions. (See genecensus.org/pseudogene for further information.)

Chromosomes, Fungal↗

The TRIPLES database: a community resource for yeast molecular biology.

TRIPLES is a web-accessible database of TRansposon-Insertion Phenotypes, Localization and Expression in Saccharomyces cerevisiae-a relational database housing nearly half a million data points generated from an ongoing study using large-scale transposon mutagenesis to characterize gene function in yeast. At present, TRIPLES contains three principal data sets (i.e. phenotypic data, protein localization data and expression data) for over 3500 annotated yeast genes as well as several hundred non-annotated open reading frames. In addition, the TRIPLES web site provides online order forms linked to each data set so that users may request any strain or reagent generated from this project free of charge. In response to user requests, the TRIPLES web site has undergone several recent modifications. Our localization data have been supplemented with approximately 500 fluorescent micrographs depicting actual staining patterns observed upon indirect immunofluorescence analysis of indicated epitope-tagged proteins. These localization data, as well as all other data sets within TRIPLES, are now available in full as tab-delimited text. To accommodate increased reagent requests, all orders are now cataloged in a separate database, and users are notified immediately of order receipt and shipment. Also, TRIPLES is one of five sites incorporated into the new functional analysis tool Function Junction provided by the Saccharomyces Genome Database. TRIPLES may be accessed from the Yale Genome Analysis Center (YGAC) homepage at http://ygac.med.yale.edu.

Computer Graphics↗