Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Genome assembly of Astatotilapia latifasciata uncovers B chromosome-linked chromatin reorganization.

B chromosomes (Bs) are supernumerary genomic elements found in many eukaryotes, yet their full sequence composition, functional potential, and regulatory impact on the host genome remain unclear. Here, we present a chromosome-level genome assembly of the cichlid fish Astatotilapia latifasciata, integrating PacBio long reads, Illumina short reads, and Hi-C chromatin contact maps to resolve both A and B chromosomes. The 0.93 Gb assembly (N50 = 36.2 Mb) includes a 34 Mb B chromosome containing 789 predicted protein-coding genes and a markedly higher density of transposable elements (TEs), especially long terminal repeats (LTR) retrotransposons. Transcriptome profiling revealed that B-linked genes are predominantly transcriptionally repressed relative to their A chromosome paralogs. Hi-C-based chromatin modeling uncovered distinct 3D structural configurations associated with the B chromosome, including fewer topologically associating domains (TADs), reduced loop formation, and altered compartmentalization. These changes are linked to long-range chromatin interactions and genomic rearrangements, suggesting that the B chromosome reshapes the nuclear architecture of the host genome. Our study proposes a potential regulatory role of Bs in genome and provides a genomic resource for investigating chromosome evolution in cichlids.

Animals↗

Chromosome-level genome assembly of Ceroplastes pseudoceriferus Green, 1935 (Hemiptera: Coccidae).

Soft scales (Hemiptera: Coccidae) are significant polyphagous pests and majority of which are invasive species. The 364.14 Mb chromosome-level genome of Ceroplastes pseudoceriferus was assembled in this work, with a contig N50 length of 6.16 Mb and scafold N50 length of 21.24 Mb. Approximately 99.89% of assembled sequences were anchored into 18 chromosomes with the assistance of Hi-C reads. Furthermore, approximately 53.98% of the genome was composed of repetitive elements. In total, 10,475 protein-coding genes were predicted, of which 9503 (90.72%) genes were functionally annotated. The BUSCO analysis demonstrated the completeness of the genome annotation is 92.54%. This genome represents first high-quality chromosome level assembly of Coccidae, thereby advancing our knowledge of Coccidae insects and developing effective management strategies that protect crops, forests, and natural ecosystems.

Animals↗

A chromosome-level genome assembly and annotation of Cercis chuniana (Fabaceae).

The genus Cercis L., at the base of the subfamily Cercidoideae of Fabaceae, is known for its ecological adaptability and significant medicinal, ornamental, and economic value. However, the lack of a high-quality genome hinders the understanding of the evolution of Cercis and Fabaceae. In this study, we present a chromosome-level genome of Cercis chuniana by combining Illumina short reads, PacBio HiFi long reads, and Hi-C data. The final genome size is 355.53 Mb, consisting of 12 contigs with a N50 of 42.34 Mb. Notably, 344.24 Mb, corresponding to 96.82% of the genome, was anchored to seven chromosomes. The assembly comprises 24.83% repetitive sequences, including 19.32% long terminal repeats. Additionally, a total of 33,837 protein-coding genes were predicted in the genome, with 32,709 (96.67%) genes successfully annotated. The high-quality genome assembly of C. chuniana not only bridges the existing gap in genomic data and offers important resources for molecular studies of this species, but also provides essential insights for future studies on speciation, functional and comparative genomics within the Fabaceae family.

Genome, Plant↗

A telomere-to-telomere reference genome assembly of the red silk cotton tree (Bombax ceiba).

Bombax ceiba, an important ornamental tree and potential fiber resource in the textile industry, is widely distributed in tropical and subtropical regions. In this study, we assembled a nearly gap-free telomere-to-telomere (T2T) genome of B. ceiba using Illumina, PacBio High-fidelity (HiFi), ONT ultra-long, and Hi-C sequencing technologies. The genome spanned approximately 807.89 Mb, with a scaffold N50 of 16.58 Mb, and 754.68 Mb (93.41%) of genomic sequences were anchored onto 48 pseudo-chromosomes. Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis revealed a completeness of 99.40%, identifying 1,378 single-copy and 213 duplicated genes out of 1,614. The genome contained 67.72% (547.11 Mb) repeat regions, with 39,708 predicted protein-coding genes. Collectively, our study provides valuable genomic data for investigating the evolutionary history of the Malvaceae family.

Genome, Plant↗

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38 Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55 Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87 Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals↗

Chromosome-level genome assembly of a cosmopolitan marine harmful algal bloom diatom species Chaetoceros socialis (Chaetocerotaceae).

Chaetoceros socialis is a cosmopolitan diatom species that is crucial for maintaining marine ecosystem structure and driving elemental cycles. C. socialis can form harmful algal blooms (HABs) that may cause a negative impact on the marine ecosystems. Whole-genome information for C. socialis is still unavailable, which may hinder more targeted studies on its ecological adaptive responses and evolutionary drivers. To address this gap, we employed cutting-edge genomic technologies including PacBio single-molecule real-time (SMRT) sequencing and high-throughput chromatin conformation capture (Hi-C) to achieve the first chromosome-level genome assembly of C. socialis. The assembled genome is 60.22 Mb in size with a scaffold N50 of 7.81 Mb and has been anchored to eight pseudochromosomes. A total of 13,378 protein-coding genes were predicted, of which 12,069 (90.22%) were functionally annotated. This high-quality genomic resource provides a fundamental data platform for systematically elucidating the ecological adaptation mechanisms of C. socialis.

Chromosomes↗

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant↗

A high-quality chromosome-level genome assembly of apple of Peru (Nicandra physalodes).

Nicandra physalodes, a member of the Solanaceae family, is known for its medicinal potential and strong natural insect-repellent properties, which are mainly attributed to its bioactive withanolides and alkaloids. Despite its ecological and pharmacological significance, genomic information for this species has remained limited. Here, we generated a chromosome-level reference genome for N. physalodes based on PacBio high-fidelity (HiFi) long-read sequencing and Hi-C scaffolding. The assembled genome is 933.97 Mb in size, with a contig N50 of 87.37 Mb, and 99.95% (933.54 Mb) of the sequences anchored to 10 pseudochromosomes. Repetitive elements account for 73.06% of the genome, and 27,925 protein-coding genes were predicted, 97.81% of which were functionally annotated. This genomic resource provides a valuable foundation for investigating the genetic basis of specialized metabolite biosynthesis, insect resistance, and environmental adaptation in N. physalodes, as well as for comparative studies within the Solanaceae family.

Genome, Plant↗

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29 Mb; scaffold N50: 7.23 Mb) with 31 chromosomes and a genome size of 200.86 Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow↗

The first near telomere-to-telomere genome assembly of Panulirus homarus homarus.

The scalloped spiny lobster (Panulirus homarus homarus) is an economically important decapod crustacean with high aquaculture potential. Several chromosome-level genomes of this species have been reported. But the lobster or even entire shrimps did not have the telomere-to-telomere assembly until now. Therefore, we present the first near telomere-to-telomere genome assembly of P. h. homarus generated by using pure Oxford Nanopore Technologies ultra-long (ONT) reads and Hi-C sequencing. The final assembly anchored to 73 chromosomes with a contig N50 of 41.1 Mb. 73 chromosomes contain entire 146 telomeres, of which 51 chromosomes have no gaps. A total of 38,396 protein-coding genes were predicted. BUSCO analysis showed a completeness score of 99.8%, indicating a high degree of assembly completeness. This high-quality genomic dataset provides a valuable resource for comparative genomics, evolutionary studies, and genome-assisted breeding of spiny lobsters.

Animals↗

Chromosome-level genome assembly of the horned turban snail Turbo cornutus.

The horned turban snail (Turbo cornutus) is an ecologically and economically important herbivorous gastropod inhabiting nearshore rocky reef habitats. T. cornutus represents a valuable coastal fishery resource in East Asia. Here, we present a chromosome-level genome assembly for T. cornutus generated using a combination of PacBio HiFi long-read and Illumina short-read sequencing and Hi-C scaffolding. The assembled genome spanned 1.93 Gb and was organized into 18 pseudo-chromosomes, representing 99.50% of the total assembly. The contig and scaffold N50 lengths were 41.02 Mb and 104.01 Mb, respectively, with repeat sequences constituting 59.07% of the genome. A total of 28,920 protein-coding genes were predicted, and genome completeness was assessed at 99.3% using the BUSCO mollusca_odb12 dataset. This chromosome-level genome assembly provides a reference for future studies on the biology of T. cornutus, the organization of the gastropod genome, and comparative genomics.

Animals↗

Whole-genome sequencing and analysis of the endophytic fungus Alternaria alternata Y-2 from Leymus chinensis.

To explore the genetic basis and functional potential of beneficial symbiosis between the endophytic fungus Alternaria alternata Y-2 and its host Leymus chinensis, we performed Illumina-based draft whole-genome sequencing and systematic bioinformatic analysis. Although this assembly does not reach telomere-to-telomere completeness, it provides high-quality gene-level information for gene prediction, functional annotation, carbohydrate-active enzyme (CAZyme) identification, and secondary metabolite biosynthetic gene cluster analysis. The final genome size of A. alternata Y-2 was 34,383,676 bp with a GC content of 51.0%, containing 12,724 predicted protein-coding genes, 90 tRNAs, and 12 rRNAs. BUSCO assessment showed 98.9% completeness, supporting the high quality of this draft genome. A total of 12,627 genes were successfully annotated in the NCBI NR database, and 17,183 genes were functionally categorized using GO terms. In total, 448 CAZyme genes and 21 secondary metabolite biosynthetic gene clusters were identified, which are potentially involved in lignocellulose degradation, cellular redox homeostasis and biosynthesis of bioactive metabolites. Based on ITS sequence alignment, NR annotation, and phylogenetic analysis of single-copy orthologous genes, the strain was confidently identified as A. alternata. This study firstly reports the draft genome of an endophytic A. alternata strain derived from L. chinensis and provides valuable genetic resources for exploring the endophytic lifestyle, stress tolerance, and bioactive metabolite potential of this fungus.

Alternaria↗

Separation and preconcentration phenomena in internally heated poly(dimethylsilicone) capillaries: preliminary modelling and demonstration studies.

The concept of achieving low-resolution separations in internally heated capillary membranes is discussed in terms of controlling the diffusion coefficients of volatile organic compounds in poly(dimethylsilicone) membranes in space and time. The behaviour of 1,1,1-trichloroethane in polydimethylsilicone was used in conjunction with a mixed-physics finite element model, incorporating second order partial differential equations, to describe time and spatial variations of mass-flux, membrane temperature and diffusion coefficients. The model, coded with Femlab, predicted highly non-linear diffusion coefficient profiles resulting from temperature programming a 500 [micro sign]m thick membrane, with an increase in the diffusion coefficient of approximately 30% in the last 30% of the membrane thickness. Simulations of sampling hypothetical analytes, with disparate temperature dependent diffusion coefficient relationships, predicted distinct thermal desorption profiles with selectivities that reflected the extent of diffusion through the membrane. The predicted desorption profiles of these analytes also indicated that low resolution separations were possible. An internally heated poly(dimethylsilicone) capillary membrane was constructed from a 10 cm long, 1.5 mm od capillary with 0.5 mm thick walls. Thirteen aqueous standards of volatile organic compounds of environmental significance were studied, and low-resolution separations were indicated, with temperature programming of the membrane enabling desorption profiles to be differentiated. Further, analytically useful relationships in the [micro sign]g cm(-3) concentration range were demonstrated with correlation coefficients >0.96 observed for linear regressions of desorption profile intensities to analyte concentrations.

Journal Article↗

Human carboxypeptidase E. Isolation and characterization of the cDNA, sequence conservation, expression and processing in vitro.

Carboxypeptidase E (CPE), which cleaves C-terminal amino acid residues and is involved in neuropeptide processing, is itself subject to intracellular processing. Human CPE cDNA was isolated and sequence comparisons were made with those of a previously isolated brain cDNA (M1622) encoding rat CPE and of other human carboxypeptidases (M and N). Human (2.5 kb) and rat (2.1 kb) CPE cDNAs approximated to the size of their respective mRNAs; additional sequences were located in putative 5' and 3' untranslated regions of human CPE mRNA. There is 79% sequence similarity between human and rat CPE cDNAs, with greater similarity (89%) over the coding region and short sections of the non-coding sequence. The predicted 476-amino acid-residue sequences of human and rat preproCPEs are highly conserved (96% identity), with lower degree of similarity of the N-terminal signal peptide (76%). Human CPE showed 51% and 43% sequence similarity to human CPN and CPM respectively, with discrete regions of divergence dispersed between the highly conserved mechanistically implicated regions. Antiserum generated from a fusion protein, synthesized in Escherichia coli from constructs of the human cDNA, recognized an approx. 50 kDa membrane protein and a smaller soluble protein in rat and human brain preparations, corresponding to the two forms of native CPE. Human CPE mRNA transcripts directed the synthesis in reticulocyte lysate of a 54 kDa translation product, which in the presence of dog pancreas microsomal membranes was co-translationally processed with cleavage, insertion into membranes and glycosylation. Three processed forms were generated, the largest (56 kDa) and smallest (52 kDa) being equally glycosylated. The membrane association of the processed translation products and of native brain membrane CPE, detected immunologically, was resistant to moderate alkali but not pH 11.5 extraction. These results are consistent with secondary-structure predictions that CPE is a peripheral membrane protein. The dissimilar regions of human carboxypeptidases may provide information on sequences responsible for their different cellular disposition.

Amino Acid Sequence↗

L-Mandelate dehydrogenase from Rhodotorula graminis: cloning, sequencing and kinetic characterization of the recombinant enzyme and its independently expressed flavin domain.

The l-mandelate dehydrogenase (L-MDH) from the yeast Rhodotorula graminis is a mitochondrial flavocytochrome b2 which catalyses the oxidation of mandelate to phenylglyoxylate coupled with the reduction of cytochrome c. We have used the N-terminal sequence of the enzyme to isolate the gene encoding this enzyme using the PCR. Comparison of the genomic sequence with the sequence of cDNA prepared by reverse transcription PCR revealed the presence of 11 introns in the coding region. The predicted amino acid sequence indicates a close relationship with the flavocytochromes b2 from Saccharomyces cerevisiae and Hansenula anomala, with about 40% identity to each. The sequence shows that a key residue for substrate specificity in S. cerevisiae flavocytochrome b2, Leu-230, is replaced by Gly in L-MDH. This substitution is likely to play an important part in determining the different substrate specificities of the two enzymes. We have developed an expression system and purification protocol for recombinant L-MDH. In addition, we have expressed and purified the flavin-containing domain of L-MDH independently of its cytochrome domain. Detailed steady-state and pre-steady-state kinetic investigations of both L-MDH and its independently expressed flavin domain have been carried out. These indicate that L-MDH is efficient with both physiological (cytochrome c, kcat=225 s-1 at 25 degrees C) and artificial (ferricyanide, kcat=550 s-1 at 25 degrees C) electron acceptors. Kinetic isotope effects with [2-2H]mandelate indicate that H-C-2 bond cleavage contributes somewhat to rate-limitation. However, the value of the isotope effect erodes significantly as the catalytic cycle proceeds. Reduction potentials at 25 degrees C were measured as -120 mV for the 2-electron reduction of the flavin and -10 mV for the 1-electron reduction of the haem. The general trends seen in the kinetic studies show marked similarities to those observed previously with the flavocytochrome b2 (L-lactate dehydrogenase) from S. cerevisiae.

Alcohol Oxidoreductases↗

Functional expression and characterization of the cytoplasmic aminopeptidase P of Caenorhabditis elegans.

Aminopeptidase P (AP-P; X-Pro aminopeptidase; EC 3.4.11.9) cleaves the N-terminal X-Pro bond of peptides and occurs in mammals as both cytosolic and plasma membrane forms, encoded by separate genes. In mammals, the plasma membrane AP-P can function as a kininase, but little is known about the physiological role of the cytosolic enzyme. The C. elegans genome contains a single gene encoding AP-P (W03G9.4), analysis of which predicts regions displaying high levels of amino-acid sequence homology between the predicted gene product and mammalian cytoplasmic AP-P, with the absolute conservation of key catalytic residues. The sequence of an EST (yk91g4), comprising the open reading frame of W03G9.4, confirmed the predicted genomic structure of the gene and the prediction that W03G9.4 codes for a nonsecreted protein with a molecular mass of 68 kDa. Nematodes transformed with a promoter reporter construct, W03G9.4:GFP, showed high levels of fluorescence in the intestine of larvae and adult hermaphrodites, indicating that the intestine is a major site of W03G9.4 expression. yk91g4 tagged with a hexahistidine and DLYDDDDK peptide epitope was expressed in Escherichia coli to yield, after affinity purification, a recombinant protein with a molecular mass of 71 kDa. The recombinant W03G9.4 removed the N-terminal amino acid from bradykinin (RPPGFSPFR), a Caenorhabditis elegans neuropeptide (KPSFVRFamide) and Lem Trp 1 (APSGFLGVRamide), but did not display activity towards angiotensin I (NRVYIHPFHL), des-Arg bradykinin and AF1 (KNEFIRFamide). The activity towards bradykinin was inhibited by EDTA and 1, 10 phenanthroline, as expected for a metalloenzyme, and also by apstatin (IC50, 1 microM), a selective inhibitor of mammalian AP-P. A Km of 45 microM and an optimum pH of 7-8 was observed with bradykinin as the substrate. The activity of the nematode AP-P, like its mammalian counterparts, was strongly influenced by metal ions, with Co2+, Mn2+ and Zn2+ all inhibiting the hydrolysis of bradykinin. We conclude that W03G9.4 codes for a cytoplasmic AP-P with very similar enzymatic properties to those of mammalian AP-P, and we suggest that the enzyme has a physiological role in the intracellular hydrolysis of proline-containing peptides absorbed from the lumen of the intestine.

Amino Acid Sequence↗

Rga5p is a specific Rho1p GTPase-activating protein that regulates cell integrity in Schizosaccharomyces pombe.

Schizosaccharomyces pombe Rho1p regulates (1,3)beta-d-glucan synthesis and is required for cell integrity maintenance and actin cytoskeleton organization, but nothing is known about the regulation of this protein. At least nine different S. pombe genes code for proteins predicted to act as Rho GTPase-activating proteins (GAPs). The results shown in this paper demonstrate that the protein encoded by the gene named rga5+ is a GAP specific for Rho1p. rga5+ overexpression is lethal and causes morphological alterations similar to those reported for Rho1p inactivation. rga5+ deletion is not lethal and causes a mild general increase in cell wall biosynthesis and morphological alterations when cells are grown at 37 degrees C. Upon mild overexpression, Rga5p localizes to growth areas and possesses both in vivo and in vitro GAP activity specific for Rho1p. Overexpression of rho1+ in rga5Delta cells is lethal, with a morphological phenotype resembling that of the overexpression of the constitutively active allele rho1G15V. In addition (1,3)beta-d-glucan synthase activity, regulated by Rho1p, is increased in rga5Delta cells and decreased in rga5-overexpressing cells. Moreover, the increase in (1,3)beta-d-glucan synthase activity caused by rho1+ overexpression is considerably higher in rga5Delta than in wild-type cells. Genetic interactions suggest that Rga5p is also important for the regulation of the other known Rho1p effectors, Pck1p and Pck2p.

Cell Wall↗

Analysis of the chromosome sequence of the legume symbiont Sinorhizobium meliloti strain 1021.

Sinorhizobium meliloti is an alpha-proteobacterium that forms agronomically important N(2)-fixing root nodules in legumes. We report here the complete sequence of the largest constituent of its genome, a 62.7% GC-rich 3,654,135-bp circular chromosome. Annotation allowed assignment of a function to 59% of the 3,341 predicted protein-coding ORFs, the rest exhibiting partial, weak, or no similarity with any known sequence. Unexpectedly, the level of reiteration within this replicon is low, with only two genes duplicated with more than 90% nucleotide sequence identity, transposon elements accounting for 2.2% of the sequence, and a few hundred short repeated palindromic motifs (RIME1, RIME2, and C) widespread over the chromosome. Three regions with a significantly lower GC content are most likely of external origin. Detailed annotation revealed that this replicon contains all housekeeping genes except two essential genes that are located on pSymB. Amino acid/peptide transport and degradation and sugar metabolism appear as two major features of the S. meliloti chromosome. The presence in this replicon of a large number of nucleotide cyclases with a peculiar structure, as well as of genes homologous to virulence determinants of animal and plant pathogens, opens perspectives in the study of this bacterium both as a free-living soil microorganism and as a plant symbiont.

Bacterial Proteins↗