Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Multilocus sequence typing of serotype III group B streptococcus and correlation with pathogenic potential.

Serotype III group B streptococcus (GBS) causes more invasive disease in infants than do other serotypes in North America. We used multilocus sequence typing to identify clones within 28 invasive serotype III GBS isolates identified from a population-based study and 55 serotype III GBS colonizing isolates from a cohort of women from the same population. Ten allelic sequence types (STs) were identified and primarily involved 2 profiles: ST-19 (57.1% of invasive isolates and 58.2% of colonizing isolates) and ST-17 (32.1% of invasive isolates and 29.1% of colonizing isolates). On concatenation, the 10 allelic profiles converged into 3 groups. Group 1 consisted of ST-19 complex, ST-36, and ST-1, and was closely related to reference genome 2603V/R (serotype V). Group 2 consisted of ST-17 complex. Group 3 consisted of ST-23 complex and was closely related to the serotype III genome strain NEM 316. Neither of the major sequence types or groups was more commonly associated with invasion (P=.61) or with lower levels of maternal capsular polysaccharide-specific IgG (0.89 microg/mL and 0.39 microg/mL, respectively) for ST-19 and ST-17 (P=.86). The close association of genomic strain 2603V/R (serotype V) with ST-19 suggests that the phenomenon of capsule switching may have occurred.

Alleles↗

Genomewide conserved epitope profiles of HIV-1 predicted by biophysical properties of MHC binding peptides.

We propose a new method for predicting MHC binding of peptides using biophysical parameters of the constituent amino acids. Unlike conventional matrix-based methods, our method does not assume independent binding of the individual side chains and uses a model that simultaneously represents all the residues. The model discovers the quantified 9-mer "property model" within the longer peptides that are most common among binders. Prediction for a new peptide is based on its statistical "distance" from the extracted peptide property model. MHC-specific peptide property models were constructed from compiled binder/nonbinder data using this method. We report the results of cross-validation of the prediction method and comparison with other methods. The comparison suggests that our method performs substantially better for some MHC class II molecules and equally well for other MHC types. To demonstrate large-scale utility, 30 HIV-1 reference genomes covering diverse subtypes were analyzed. Regions that are likely to bind MHC (A2, DR1, or DR4) and that are conserved across the HIV-1 subtypes were identified. These "epitope profiles" of the diverse HIV-1 strains can also be visually presented to facilitate discovery of conserved patterns naturally occurring in the viral genomes. As an essential step in designing vaccines, the revealed patterns may provide valuable information in identifying the immunologically important regions.

Computer Simulation↗

Characterization of the expression and immunogenicity of poliovirus replicons that encode simian immunodeficiency virus SIVmac239 Gag or envelope SU proteins.

The effectiveness of the poliovirus vaccines to induce both systemic and mucosal immunity has prompted the development of this virus as a vector in which to express foreign proteins. Our laboratory has previously reported on the construction and characterization of poliovirus genomes that encode HIV-1 proteins (Porter DC, et al.: J Virol 1996;70:2643-2649). To develop this system further, we have constructed poliovirus genomes, referred to as replicons, which encode the SIVmac239 Gag or Env SU in place of the poliovirus capsid gene (P1). Since the replicons do not encode capsid proteins, they are encapsidated into poliovirus by passage with a recombinant vaccinia virus, VVP1, which provides the poliovirus capsid proteins in trans. Using this system, we have derived stocks of the encapsidated replicons which encode the SIVmac239 or Env SU protein. Infection of cells with the replicon that encodes SIVmac239 Gag resulted in the expression of a 55-kDa protein that was released from the infected cells. Analysis of the sedimentation of the released proteins by sucrose density gradient centrifugation revealed that the protein was released from the cell in the form of a virus-like particle. Infection of cells with the replicons encoding the SIVmac239 Env SU resulted in the expression of a 63-kDa protein, corresponding to the molecular mass predicted for the nonglycosylated SIVmac239 SU protein. A second protein with a molecular mass greater than 160 kDa was also immunoprecipitated. After enzymatic deglycosylation, this protein migrated at a molecular mass consistent with that for an Env SU dimer. Analysis of the medium from cells infected with the replicon encoding SIVmac239 Env SU revealed the presence of a protein of molecular mass 85-90 kDa, possibly representing a fragment of the SIVmac239 or Env SU protein. To determine the immunogenicity of the replicons encoding SIVmac239 Gag or Env SU, transgenic mice that express the human receptor for poliovirus, and are thus susceptible to poliovirus, were immunized via the intramuscular route. A serum antibody response to SIV envelope was detected following booster immunization, establishing that the encapsidated replicon was immunogenic. Finally, we demonstrate that the replicons have the capacity to infect peripheral blood mononuclear monocytes/macrophages, suggesting that this cell is a possible target for in vivo infection. The results of our studies, then, lend further support for the development and application of recombinant poliovirus replicons in a vaccine strategy.

Animals↗

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software↗

GCfix: a fast and accurate fragment length-specific method for correcting GC bias in cell-free DNA.

MOTIVATION: Cell-free DNA (cfDNA) analysis has wide-ranging clinical applications due to its noninvasive nature. However, cfDNA fragmentomics and copy number analysis can be complicated by GC bias. There is a lack of GC correction software based on rigorous cfDNA GC bias analysis. Furthermore, there is no standardized metric for comparing GC bias correction methods across large sample sets, nor a rigorous experiment setup to demonstrate their effectiveness on cfDNA data at various coverage levels. RESULTS: We present GCfix, a method for robust GC bias correction in cfDNA data across diverse coverages. Developed following an in-depth analysis of cfDNA GC bias at the region and fragment length levels, GCfix is both fast and accurate. It works on all reference genomes and generates correction factors, tagged BAM files, and corrected coverage tracks. We also introduce two orthogonal performance metrics for (i) comparing the fragment count density distribution of GC content between expected and corrected samples, and (ii) evaluating coverage profile improvement post-correction. GCfix outperforms existing cfDNA GC bias correction methods on these metrics. AVAILABILITY AND IMPLEMENTATION: GCfix software and code for reproducing the figures are publicly accessible on GitHub: https://github.com/Rafeed-bot/GCfix_Software.

Software↗

AdDeam: a fast and scalable tool for estimating and clustering reference-level damage profiles.

MOTIVATION: DNA damage patterns, such as increased frequencies of C→T and G→A substitutions at fragment ends, are widely used in ancient DNA studies to assess authenticity and detect contamination. In metagenomic studies, fragments can be mapped against multiple references or de novo assembled contigs to identify those likely to be ancient. Generating and comparing damage profiles, however, can be both tedious and time-consuming. Although tools exist for estimating damage in single reference genomes and metagenomic datasets, none efficiently cluster damage patterns. RESULTS: To address this methodological gap, we developed AdDeam, a tool that combines rapid damage estimation with clustering for streamlined analyses and easy identification of potential contaminants or outliers. Our tool takes aligned ancient DNA (aDNA) fragments from various samples or contigs as input, computes damage patterns, clusters them, and outputs representative damage profiles per cluster, a probability of each sample pertaining to a cluster, as well as a Principal Component Analysis of the damage patterns for each sample for fast visualisation. We evaluated AdDeam on both simulated and empirical datasets. AdDeam effectively distinguishes different damage levels, such as uracil-DNA glycosylase-treated samples, sample-specific damages from specimens of different time periods, and can also distinguish between contigs containing modern or ancient fragments, providing a clear framework for aDNA authentication and facilitating large-scale analyses. AVAILABILITY AND IMPLEMENTATION: AdDeam is publicly available at https://github.com/LouisPwr/AdDeam and can also be installed via Bioconda. It is implemented in Python and C++. All analysis scripts and datasets are available at https://github.com/LouisPwr/AdDeamAnalysis and on Zenodo under: 10.5281/zenodo.15052427.

Software↗

WinPCA: a package for windowed principal component analysis.

SUMMARY: With chromosomal reference genomes and population-scale whole genome-sequencing becoming increasingly accessible, contemporary studies often include characterizations of the genomic landscape as it varies along chromosomes, commonly termed genome scans. While traditional summary statistics like FST and dXY between pre-assigned populations remain integral to characterizing the genomic divergence profile, PCA differs by providing single-sample resolution, thereby supporting the identification of polymorphic inversions, introgression and other types of divergent sequence that may not be fully aligned with global population structure. Here, we introduce WinPCA, a user-friendly package to compute, polarize and visualize genetic principal components in windows along the genome. To accommodate low-coverage whole genome-sequencing datasets, WinPCA can optionally make use of PCAngsd methods to compute principal components in a genotype likelihood framework. WinPCA accepts variant data in either VCF or BEAGLE format and can generate rich plots for interactive data exploration and downstream presentation. AVAILABILITY AND IMPLEMENTATION: WinPCA is implemented in Python and freely available at https://github.com/MoritzBlumer/winpca and https://doi.org/10.5281/zenodo.15614979.

Software↗

Benchmarking methods for measuring biosynthetic gene cluster similarity and determination of gene cluster families.

MOTIVATION: Natural products are often produced by a set of biosynthetic enzymes that are encoded by genes clustered together in the producer's genome, referred to as a biosynthetic gene cluster (BGC). The ability to compare and cluster BGCs is essential for several applications, including predicting which bacteria will make a known product and assessing the potential diversity of natural products produced by a set of bacteria. There are multiple methods for comparing and clustering BGCs based on their similarity, but there has been a lack of investigation into how strongly BGC similarity relates to product structural similarity and how these methods perform relative to each other. RESULTS: Using publicly available databases, we developed a benchmark dataset to assess how well different BGC similarity metrics correlate with the structural similarity of their products and how well these methods cluster BGCs. We found that all methods showed moderate correlation between BGC and structural similarity, with correlations improving for more similar BGCs and varying significantly by BGC biosynthetic class. Analysis of outliers revealed some outliers were due to mistakes or omissions in public datasets, while others represented deviation between BGC similarity and product structural similarity. All methods generally performed better on clustering metrics, with BiG-SCAPE performing the best after errors in the public datasets had been corrected. AVAILABILITY AND IMPLEMENTATION: Scripts and data required to reproduce the results are available at https://github.com/aswalker-lab/BGC-clustering-benchmark and processed similarity, clusters, and scaffolds are also available at https://huggingface.co/datasets/allie-walker/BGC-clustering-benchmark. Code is also available at Zenodo: 10.5281/zenodo.17373546.

Multigene Family↗

Columba: fast approximate pattern matching with optimized search schemes.

MOTIVATION: Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. RESULTS: This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first, based on the bidirectional FM-index, prioritizes speed. The second, Columba RLC, uses run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Benchmarks on the human genome, as well as bacterial and human pan-genome datasets, demonstrate that Columba is much faster than existing lossless aligners and even competitive with lossy tools. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses. AVAILABILITY AND IMPLEMENTATION: The source code of Columba is available at https://github.com/biointec/columba under AGPL license. Scripts to reproduce the benchmarks and analyses are available at https://doi.org/10.5281/zenodo.15849246.

Software↗

RAmpSim: a thermodynamic simulator for hybridization capture in metagenomic sequencing.

MOTIVATION: Simulators that generate synthetic datasets help address the lack of ground truth for developing and benchmarking computational tools. Many read simulators assume uniform sampling across reference genomes; however, for newer capture-based sequencing technologies (e.g. TELSeq), this assumption is intentionally broken to oversample regions of interest. Along with systematic biases arising from probe multiplicity, sequence composition, and species abundances inherent to capture-based sequencing, this mismatch between modeling assumptions and the characteristics of real data necessitates the design of a new capture-based sequencing-specific simulator. RESULTS: We present RAmpSim, a fast simulator that models bait-target hybridization and fragment capture using a thermodynamic nearest-neighbor energy model and Boltzmann-weighted sampling of binding sites. Fragments are generated through multinomial sampling parameterized by bait concentration, binding energy, and genomic abundance before being passed to existing models of platform-specific errors. Implemented in Rust, RAmpSim reproduces empirical within-genome coverage and cross-species enrichment patterns observed in capture-based metagenomic datasets. RAmpSim generally outperforms a uniform baseline with respect to position-based earth mover's distance when compared against the empirical coverage distribution. Classification analysis also shows high recall in recovering empirical high-coverage regions while outperforming a uniform baseline. AVAILABILITY: Code, example scripts, and data sources are available at https://github.com/az002/RAmpSim.git.

Metagenomics↗

GiantHost: a domain-adaptive and uncertainty-aware framework for giant virus host prediction.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty-a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent. RESULTS: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts-revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone. AVAILABILITY: The source code of GiantHost is available via: https://github.com/FuchuanQu/GiantHost.

Giant Viruses↗

Quantitative DNA methylation analysis based on four-dye trace data from direct sequencing of PCR amplificates.

MOTIVATION: Methylation of cytosines in DNA plays an important role in the regulation of gene expression, and the analysis of methylation patterns is fundamental for the understanding of cell differentiation, aging processes, diseases and cancer development. Such analysis has been limited, because technologies for detailed and efficient high-throughput studies have not been available. We have developed a novel quantitative methylation analysis algorithm and workflow based on direct DNA sequencing of PCR products from bisulfite-treated DNA with high-throughput sequencing machines. This technology is a prerequisite for success of the Human Epigenome Project, the first large genome-wide sequencing study for DNA methylation in many different tissues. Methylation in tissue samples which are compositions of different cells is a quantitative information represented by cytosine/thymine proportions after bisulfite conversion of unmethylated cytosines to uracil and PCR. Calculation of quantitative methylation information from base proportions represented by different dye signals in four-dye sequencing trace files needs a specific algorithm handling imbalanced and overscaled signals, incomplete conversion, quality problems and basecaller artifacts. RESULTS: The algorithm we developed has several key properties: it analyzes trace files from PCR products of bisulfite-treated DNA sequenced directly on ABI machines; it yields quantitative methylation measurements for individual cytosine positions after alignment with genomic reference sequences, signal normalization and estimation of effectiveness of bisulfite treatment; it works in a fully automated pipeline including data quality monitoring; it is efficient and avoids the usual cost of multiple sequencing runs on subclones to estimate DNA methylation. The power of our new algorithm is demonstrated with data from two test systems based on mixtures with known base compositions and defined methylation. In addition, the applicability is proven by identifying CpGs that are differentially methylated in real tissue samples.

Algorithms↗

The Puf3 protein is a transcript-specific regulator of mRNA degradation in yeast.

Eukaryotic post-transcriptional regulation is often specified by control elements within mRNA 3'- untranslated regions (3'-UTRs). In order to identify proteins that regulate specific mRNA decay rates in Saccharomyces cerevisae, we analyzed the role of five members of the Puf family present in the yeast genome (referred to as JSN1/PUF1, PUF2, PUF3, PUF4 and MPT5/PUF5). Yeast strains lacking all five Puf proteins showed differential expression of numerous yeast mRNAs. Examination of COX17 mRNA indicates that Puf3p specifically promotes decay of this mRNA by enhancing the rate of deadenylation and subsequent turnover. Puf3p also binds to the COX17 mRNA 3'-UTR in vitro. This indicates that the function of Puf proteins as specific regulators of mRNA deadenylation has been conserved throughout eukaryotes. In contrast to the case in Caenorhabditis elegans and Drosophila, yeast Puf3p does not affect translation of COX17 mRNA. These observations indicate that Puf proteins are likely to play a role in the control of transcript-specific rates of degradation in yeast by interacting directly with the mRNA turnover machinery.

3' Untranslated Regions↗

Pervasive positive selection on X-linked ampliconic genes in primates.

Mammalian sex chromosomes harbour ampliconic gene families, which are multi-copy genes with ≥97% sequence identity, predominantly expressed in testis tissue and essential for male fertility. The amplification of testis-specific genes is conserved across mammals, yet the specific gene families that expand show striking lineage-specific variation. Previous studies suggest a dynamic turnover with adaptive evolution for several of these families, but their analysis has been limited by the quality of reference genomes of repetitive regions. To characterise the molecular evolutionary processes of ampliconic gene families on both sex chromosomes, we analysed telomere-to-telomere genome assemblies from eight primate species spanning 25 million years of evolution. We identified 53 X-linked and 19 Y-linked ampliconic gene families with dynamic copy number variation. Gene conversion through palindromic pairing and tandem arrays maintained high sequence similarity despite accumulating mutations. X-linked families maintained conserved chromosomal positions despite copy number changes, whereas Y-linked families showed frequent positional turnover. Strikingly, multiple X-linked families (GAGE, SSX, CSAG, and VCX) showed pervasive positive selection across the primate phylogeny and multiple (MAGEB, CT45, HSFX) showed lineage specific positive selection. Y-linked families predominantly evolve under purifying selection. Examining intraspecific copy number variation of the X-linked ampliconic families in chimpanzees, humans, and gorillas, we found variation among individuals but clear differences between species, with the largest families varying the most. These patterns could suggest that sperm competition, meiotic drive, or dosage-dependent selection drive the rapid, lineage-specific evolution of testis-expressed ampliconic genes in primates.

Journal Article↗

Hypoxia Response Is Associated with Reduced HPV Activity and Tumor Microenvironment Remodeling in Cervical Cancer.

Human papillomavirus (HPV) significantly influences cervical cancer progression and treatment, yet its interactions with the tumor microenvironment remain incompletely understood. We performed single-cell and spatial transcriptomic sequencing on cervical cancer samples to explore these interactions. By aligning sequencing reads to a merged HPV16-human reference genome, we characterized HPV16 heterogeneity and its association with host states at the single-cell and single-gene levels. E5 transcriptional activity was negatively associated with the host interferon response, indicating a role in immune evasion. A hypoxic environment was correlated with the downregulation of E5 activity and elevated MHC-I expression, which may contribute to stronger interactions between hypoxic cancer cells and cytotoxic CD8⁺ T cells. Additionally, HPV16 integration in host cells was associated with increased fatty acid metabolism. These findings suggest that combining anti-angiogenic drugs and fatty acid metabolism inhibitors has the potential to improve cervical cancer treatment.

Cervical cancer↗

A palindrome-mediated mechanism distinguishes translocations involving LCR-B of chromosome 22q11.2.

Two known recurrent constitutional translocations, t(11;22) and t(17;22), as well as a non-recurrent t(4;22), display derivative chromosomes that have joined to a common site within the low copy repeat B (LCR-B) region of 22q11.2. This breakpoint is located between two AT-rich inverted repeats that form a nearly perfect palindrome. Breakpoints within the 11q23, 17q11 and 4q35 partner chromosomes also fall near the center of palindromic sequences. In the present work the breakpoints of a fourth translocation involving LCR-B, a balanced ependymoma-associated t(1;22), were characterized not only to localize this junction relative to known genes, but also to further understand the mechanism underlying these rearrangements. FISH mapping was used to localize the 22q11.2 breakpoint to LCR-B and the 1p21 breakpoint to single BAC clones. STS mapping narrowed the 1p21.2 breakpoint to a 1990 bp AT-rich region, and junction fragments were amplified by nested PCR. Junction fragment-derived sequence indicates that the 1p21.2 breakpoint splits a 278 nt palindrome capable of forming stem-loop secondary structure. In contrast, the 1p21.2 reference genomic sequence from clones in the database does not exhibit this configuration, suggesting a predisposition for regional genomic instability perhaps etiologic for this rearrangement. Given its similarity to known chromosomal fragile site (FRA) sequences, this polymorphic 1p21.2 sequence may represent one of the FRA1 loci. Comparative analysis of the secondary structure of sequences surrounding translocation breakpoints that involve LCR-B with those not involving this region indicate a unique ability of the former to form stem-loop structures. The relative likelihood of forming these configurations appears to be related to the rate of translocation occurrence. Further analysis suggests that constitutional translocations in general occur between sequences of similar melting temperature and propensity for secondary structure.

Base Sequence↗

Identification of a repeated sequence in the genome of the sea urchin which is transcribed by RNA polymerase III and contains the features of a retroposon.

A repeated sequence element which is located about 200 nucleotides upstream from the protein-coding portion of the muscle actin gene (probably within a large 5' intron) in the genome of the sea urchin, Strongylocentrotus purpuratus has been characterized, and shown to contain the sequence features which indicate that it has been transposed by means of an RNA intermediate. This retroposon-like sequence, SURF1-1, is a member of a family which is dispersed and repeated about 800 times in the genome, referred to as SURF1 (sea urchin retroposon family 1). In vitro transcription of this sequence by RNA polymerase III defines a 300 nucleotide transcription unit which is bounded by a short direct repeated sequence. The 3' end of this unit contains a simple 21 nucleotide A+T-rich sequence characteristic of retroponons, and a consensus B box portion of an internal RNA polymerase III promotor is located 60 to 80 nucleotides downstream from the two sites of transcription initiation. This sequence also contains a 40 nucleotide region that is related to several tRNA sequences (containing the B box), and a 79 nucleotide sequence which is homologous to a repeated sequence previously shown to be present within the 3' untranslated portions of the Spec1 and Spec2 mRNAs of this species (1). Analysis of transcripts of this sequence family in RNA from several embryonic stages indicates that its expression is highest at 11 hours postfertilization (about 128 cells) and drops as development proceeds. Furthermore, most or all, transcription of this sequence family in nuclei isolated from 11 hour embryos is by RNA polymerase III, and is from the same strand which is transcribed in vitro.

Actins↗

HUMHOT: a database of human meiotic recombination hot spots.

Meiotic recombination occurs preferentially at certain regions in the genome referred to as hot spots. The number of hot spots known in humans has increased manifold in recent years. The identification of these hot spots in humans is of great interest to population and medical geneticists since they influence the structure of Linkage Disequilibrium and Haplotype blocks in human populations, whose patterns have applications in mapping disease genes. HUMHOT is a web-based database of Human Meiotic Recombination Hot Spots. The database comprises DNA sequences corresponding to the hot spot regions from the literature that have been mapped to a high resolution (<4 kb) in humans. It also provides flanking sequence information for the hot spot region along with references describing the hot spot. The database can be queried based on hot spot identity, chromosome position or by homology to user-defined sequences. It is also updated with new hot spot sequences as they are discovered and provides hyperlinks to commonly used tools for estimating recombination rates, performing genetic analysis and new advances in our understanding of meiotic hot spots. Public access to the HUMHOT database is available at http://www.jncasr.ac.in/humhot.

Base Sequence↗