Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “multiple clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Patterns of gene divergence and VL promoter activity in immunoglobulin light chain clusters of the channel catfish.

The structure and genomic organization of V, J, and C segments in multiple gene clusters of the G class of catfish light (L) chain were determined to study evolutionary patterns of cluster divergence and regulatory function. The results showed that the organizational pattern is conserved; two VL segments reside in opposite transcriptional orientation to a pseudogene JL segment (psiJ), a functional JL segment, and a CL segment. Structures within the central V-psiJ-J-C regions indicate that cluster duplication occurred after the V-V-psiJ-J-C organization became fixed. VL divergence in gene clusters subsequently occurred by mechanisms that principally targeted complementarity-determining regions. The sequence of the VL-flanking regions, which contained regulatory octamer, TATA, and Pax-5 (BSAP) binding site motifs, was conserved during G cluster duplication and within VL-flanking regions in divergent lineages of bony fish. Reporter assays of catfish B cells transfected with a 742-bp VL-flanking fragment showed promoter activity in the absence of enhancer elements. The promoter's activity doubled when coupled with the catfish IgH enhancer. A 136-bp fragment containing the motifs conserved in bony fish phylogeny and located between the leader initiation codon and the initiation sites of sterile transcripts served as a minimal promoter and provided the highest B-cell activity. These constructs, however, did not act as promoters in catfish non-lymphoid cells or mammalian BJA-B B cells or fibroblasts. These results indicate that the structure and function of VL promoter regions in the regulation and tissue specificity of L-chain gene expression evolved early in phylogeny.

Amino Acid Sequence↗

Enhancing genome recovery across metagenomic samples using MAGmax.

SUMMARY: The number of metagenome-assembled genomes (MAGs) is rapidly increasing with the growing scale of metagenomic studies, driving fast progress in microbiome research. Sample-wise assembly has become the standard due to its computational efficiency and strain-level resolution. It requires dereplication, the removal of near-identical genomes assembled in different metagenomic samples. We present MAGmax, an efficient dereplication tool that enhances both the quantity and quality of MAGs through a strategy of bin merging and reassembly. Unlike dRep, which selects a single representative bin per genome cluster, MAGmax merges multiple bins within a cluster and reassembles them to increase coverage. MAGmax produces more dereplicated, higher-quality MAGs than dRep at 1.6× its speed and using three times less memory. AVAILABILITY AND IMPLEMENTATION: The MAGmax open source software, implemented in Rust, is available under the GPLv3 license at https://github.com/soedinglab/MAGmax.

Metagenomics↗

Childhood infections and risk of multiple sclerosis.

Multiple sclerosis has been hypothesized to be the result from an aberrant immune response possibly triggered by delayed exposure to a common childhood infection. Because the vast majority of previous studies testing this hypothesis have been based on a history of childhood infections recalled years to decades later in adulthood, we investigated whether age at six common childhood infections was associated with risk of multiple sclerosis, using information recalled in the childhood of a historical cohort of school children in Denmark. Cases included all individuals with multiple sclerosis in the country born between 1940 and 1975, who had attended school in the capital, Copenhagen. Controls were age- and sex-matched peers. School health records were obtained for all subjects. The records contained information on measles, pertussis, scarlet fever, birth order, sibship size, social class of the father, school years, and name of school and attended school classes for children born since 1940 (n(cases) = 455, n(controls) = 1801). For children born since 1950, the records also contained information on rubella, varicella and mumps (n(cases) = 182, n(controls) = 690). Neither age at infection with measles, rubella, varicella, mumps, pertussis and scarlet fever (upper age limit, 14 years) nor the cumulative number of these infections between the ages of 10 and 14 years was associated with the risk of multiple sclerosis. In addition, the risk of multiple sclerosis was not associated with birth order or social class. No clustering of multiple sclerosis in school classes was observed. Our findings suggest that measles, rubella, mumps, varicella, pertussis and scarlet fever, even if acquired late in childhood, are not associated with increased risk of multiple sclerosis later in life.

Adolescent↗

OrthoMCL: identification of ortholog groups for eukaryotic genomes.

The identification of orthologous groups is useful for genome annotation, studies on gene/protein evolution, comparative genomics, and the identification of taxonomically restricted sequences. Methods successfully exploited for prokaryotic genome analysis have proved difficult to apply to eukaryotes, however, as larger genomes may contain multiple paralogous genes, and sequence information is often incomplete. OrthoMCL provides a scalable method for constructing orthologous groups across multiple eukaryotic taxa, using a Markov Cluster algorithm to group (putative) orthologs and paralogs. This method performs similarly to the INPARANOID algorithm when applied to two genomes, but can be extended to cluster orthologs from multiple species. OrthoMCL clusters are coherent with groups identified by EGO, but improved recognition of "recent" paralogs permits overlapping EGO groups representing the same gene to be merged. Comparison with previously assigned EC annotations suggests a high degree of reliability, implying utility for automated eukaryotic genome annotation. OrthoMCL has been applied to the proteome data set from seven publicly available genomes (human, fly, worm, yeast, Arabidopsis, the malaria parasite Plasmodium falciparum, and Escherichia coli). A Web interface allows queries based on individual genes or user-defined phylogenetic patterns (http://www.cbil.upenn.edu/gene-family). Analysis of clusters incorporating P. falciparum genes identifies numerous enzymes that were incompletely annotated in first-pass annotation of the parasite genome.

Animals↗

Genome cluster database. A sequence family analysis platform for Arabidopsis and rice.

The genome-wide protein sequences from Arabidopsis (Arabidopsis thaliana) and rice (Oryza sativa) spp. japonica were clustered into families using sequence similarity and domain-based clustering. The two fundamentally different methods resulted in separate cluster sets with complementary properties to compensate the limitations for accurate family analysis. Functional names for the identified families were assigned with an efficient computational approach that uses the description of the most common molecular function gene ontology node within each cluster. Subsequently, multiple alignments and phylogenetic trees were calculated for the assembled families. All clustering results and their underlying sequences were organized in the Web-accessible Genome Cluster Database (http://bioinfo.ucr.edu/projects/GCD) with rich interactive and user-friendly sequence family mining tools to facilitate the analysis of any given family of interest for the plant science community. An automated clustering pipeline ensures current information for future updates in the annotations of the two genomes and clustering improvements. The analysis allowed the first systematic identification of family and singlet proteins present in both organisms as well as those restricted to one of them. In addition, the established Web resources for mining these data provide a road map for future studies of the composition and structure of protein families between the two species.

Algorithms↗

Proteins of Escherichia coli come in sizes that are multiples of 14 kDa: domain concepts and evolutionary implications.

Initial attempts to correlate the distribution of gene density (number of gene loci per unit length on the linkage map) with the distribution of lengths of coding sequences have led to the observation that 46% of approximately 1000 sampled proteins in Escherichia coli have molecular masses of n X 14,000 +/- 2500 daltons (n = 1, 2, ...). This clustering around multiples of 14,000 contrasts with the 36% one would expect in these ranges if the sizes were uniformly distributed. The entire distribution is well fit by a sum of normal or lognormal distributions located at multiples of 14,000, which suggests that the percentage of E. coli proteins governed by the underlying sizing mechanism is much greater than 50%. Clustering of protein molecular sizes around multiples of a unit size also is suggested by the distribution of well-characterized HeLa cell proteins. The distribution of gene lengths for E. coli suggests regular clustering, which implies that the clustering of protein molecular masses is not an artifact of the molecular mass measurement by gel electrophoresis. These observations suggest the existence of a fundamental structural unit. The rather uniform size of this structural unit (without any apparent sequence homology) suggests that a general principle such as geometrical or physical optimization at the DNA or protein level is responsible. This suggestion is discussed in relation to experimental evidence for the domain structure of proteins and to existing hypotheses that attempt to account for these domains. Microevolution would appear to be accommodated by incremental changes within this fundamental unit, whereas macroevolution would appear to involve "quantum" changes to the next stable size of protein.

Bacterial Proteins↗

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans↗

Psittacine beak and feather disease in three captive sulphur-crested cockatoos (Cacatua galerita) in Thailand.

Three sulphur-crested cockatoos (Cacatua galerita) were diagnosed as psittacine beak and feather disease (PBFD). Histopathology of the feather pulp and follicles showed intracytoplasmic botryoid clusters or granular inclusion bodies in epithelial cells and macrophages. Electron microscopy revealed multiple cytoplasmic clusters of electron dense viral particles corresponding to the inclusions. PBFD virus (circovirus) DNA-specific product was detected from formalin-fixed paraffin-embedded feathers by nested polymerase chain reaction (PCR) method.

Animals↗

Factor analysis of ethnic variation in the multiple metabolic (insulin resistance) syndrome in three Canadian populations.

This study describes and compares the pattern of risk factor clustering in multiple metabolic (insulin resistance) syndrome (MMS) in three Canadian ethnic groups (Indians, Inuit, non-Aboriginal Canadians). Three cross-sectional, population-based sample surveys in three contiguous regions of Canada were conducted during the late 1980s and early 1990s (Ontario, Manitoba, Northwest Territories). The combined dataset consists of 873 Cree-Ojibwa Indians from northern Ontario and Manitoba, 387 Inuit from the Northwest Territories, and 2,670 non-Aboriginal Canadians (predominantly of European origin) in the province of Manitoba. The samples are representative of the noninstitutionalized, adult population of their respective catchment areas. Factor analysis transformed 10 anthropometric and metabolic variables into three uncorrelated factors. Three factors, which together account for 64.3% of the variance, can be identified: an "obesity factor" (factor loadings for weight, height, waist and hip girth, and HDL-cholesterol); a "blood pressure factor" (factor loadings for mean systolic and diastolic BP and total cholesterol); and a "lipid/glucose factor" (factor loadings for triglycerides, total cholesterol, HDL, and fasting plasma glucose). Fasting insulin is available for only a subset of the data and separate analysis shows that it groups with glucose. Factor scores generated by the factor analysis differ according to ethnic group, diabetes status, and sex on multivariate analysis of variance. Indians have the highest scores for all three factors. Inuit have the lowest obesity scores and are not significantly different from non-Aboriginal people with regard to the other two factors. MMS is prevalent in diverse ethnic groups but varies in the pattern of phenotypic expression, with some components more prominent in some groups.

Adult↗

Genus-specific protein binding to the large clusters of DNA repeats (short regularly spaced repeats) present in Sulfolobus genomes.

Short regularly spaced repeats (SRSRs) occur in multiple large clusters in archaeal chromosomes and as smaller clusters in some archaeal conjugative plasmids and bacterial chromosomes. The sequence, size, and spacing of the repeats are generally constant within a cluster but vary between clusters. For the crenarchaeon Sulfolobus solfataricus P2, the repeats in the genome fall mainly into two closely related sequence families that are arranged in seven clusters containing a total of 441 repeats which constitute ca. 1% of the genome. The Sulfolobus conjugative plasmid pNOB8 contains a small cluster of six repeats that are identical in sequence to one of the repeat variants in the S. solfataricus chromosome. Repeats from the pNOB8 cluster were amplified and tested for protein binding with cell extracts from S. solfataricus. A 17.5-kDa SRSR-binding protein was purified from the cell extracts and sequenced. The protein is N terminally modified and corresponds to SSO454, an open reading frame of previously unassigned function. It binds specifically to DNA fragments carrying double and single repeat sequences, binding on one side of the repeat structure, and producing an opening of the opposite side of the DNA structure. It also recognizes both main families of repeat sequences in S. solfataricus. The recombinant protein, expressed in Escherichia coli, showed the same binding properties to the SRSR repeat as the native one. The SSO454 protein exhibits a tripartite internal repeat structure which yields a good sequence match with a helix-turn-helix DNA-binding motif. Although this putative motif is shared by other archaeal proteins, orthologs of SSO454 were only detected in species within the Sulfolobus genus and in the closely related Acidianus genus. We infer that the genus-specific protein induces an opening of the structure at the center of each DNA repeat and thereby produces a binding site for another protein, possibly a more conserved one, in a process that may be essential for higher-order stucturing of the SRSR clusters.

Amino Acid Sequence↗

Sarcoplasmic reticulum Ca2+ release flux underlying Ca2+ sparks in cardiac muscle.

Discrete events of Ca2+ release from the sarcoplasmic reticulum (SR) have been described in cardiac, skeletal, and smooth muscle. In skeletal muscle these release events originate at individual channels. In cardiac muscle, however, it remains a question of debate whether localized Ca2+ release transients, termed Ca2+ sparks, originate from single release channels or multiple channels clustered in close vicinity. Generalizing methods used earlier to describe cell-averaged Ca2+ release, we derived, as a function of space and time, the flux of Ca2+ release that underlies Ca2+ sparks. Using the method to analyze spontaneous sparks recorded with confocal microscopy in dissociated cat atrial cells, we obtained in most cases single sparks of Ca2+ release that appear to originate from approximately 1-microm-wide regions. In many cases, doublets, triplets, and greater groups of release sparks were observed. This multiplicity, the estimated release flux magnitude, and existing data on the structure of junctions between SR and plasmalemma suggest that individual release sparks result from the opening of multiple Ca2+ release channels clustered within discrete SR junctional regions.

Animals↗

Multiple comparisons model-based clustering and ternary pattern tree numerical display of gene response to treatment: procedure and application to the preclinical evaluation of chemopreventive agents.

Microarray technology has greatly aided the identification of genes that are expressed differentially. Statistical analysis of such data by multiple comparisons procedures has been slow to develop, in part, because methods to cluster the results of such comparisons in biologically meaningful ways have not been available. We isolated and analyzed, by Northern blot and GeneChip, replicate liver RNA samples (n = 4/group) from rats fed with control diet or diet containing one of three chemopreventive compounds, selected because their pharmacological activities, including RNA expression response, are relatively well understood. We report on a classification tree, based on the results of nonparametric multiple comparisons, which results in the bipolar hierarchical clustering of genes in relation to their response to treatment. In addition to identifying treatment-responsive genes, application of this procedure to our test study identified the known pharmacological relationships among the treatment groups without supervision. Also, small treatment-specific subsets of genes were identified that may be indicative of additional pharmacophores present in the test compounds.

Animals↗

The TANG cluster comprising ten nitrate transporter genes controls fruit sweetness and size in tomato.

Sucrose is a major transport form of photoassimilated carbon in tomato, Arabidopsis, and many other plant species, and plays a critical regulatory role in plant growth, development, and fruit quality. Plant vacuoles function as storage organelles, accumulating substantial quantities of metabolically inactive nitrates as a nitrogen reserve and soluble sugars as a carbon reserve. Consequently, the balance between nitrate and sucrose accumulation determines plant growth dynamics and fruit taste. In this study, we identified a gene cluster designated TANG (Total soluble solidsAccumulation viaNitrate transporterGene cluster), comprising ten nitrate transporter genes that are significantly associated with sucrose accumulation in tomato. This gene cluster mediates the transport of nitrate between the cytoplasm and vacuole, thereby influencing its storage. Functional disruption of TANG8, a member of the gene cluster, results in either enhanced sugar accumulation or increased fruit size. Selective disruption of multiple TANG cluster members yields fruits with elevated sweetness and increased fruit size in S. pimpinellifolium. The interaction between the TANG members and a tonoplast localized Sucrose Transporter 4 provides insight into the competitive accumulation of nitrate and sugar. The multiplex editing of a gene cluster provides a successful example of engineering crops with high quality and yield.

Gene cluster↗

Three methods of developing MMPI taxonomies of sexual offenders.

In a sample of 261 state hospital sexual offenders, Minnesota Multiphasic Personality Inventory (MMPI) profiles did not differ for offenders with adult victims versus offenders with child victims when offender age was controlled. MMPI 2-point analyses for the whole sample revealed five common codes that were independent of victim maturity. The sample was randomly divided in half and subjected to a cluster-analytic procedure which revealed two MMPI clusters. The first cluster was unelevated, with Scale 4 as its high point. The second cluster had multiple elevations, with Scales 8, 4, 2, and 7 as the highest scales. These clusters were replicated in a cluster analysis of the second half of the sample. However, when the sample was recombined, the two clusters were not externally validated basis on demographic and criminological variables. The results suggest that common psychological variables among sexual offenders may have more discriminative value than victim maturity in developing sexual offender taxonomies.

Adult↗

Assessment of clonal relationships in ipsilateral and bilateral multiple breast carcinomas by comparative genomic hybridisation and hierarchical clustering analysis.

The issue of whether multiple, ipsilateral or bilateral, breast carcinomas represent multiple primary tumours or dissemination of a single carcinomatous process has been difficult to resolve, especially for individual patients. We have addressed the problem by comparative genomic hybridisation analysis of 26 tumours from 12 breast cancer patients with multiple ipsilateral and/or bilateral carcinoma lesions. Genomic imbalances were detected in 25 of the 26 (96%) tumours. Using the genomic imbalances detected in these 26 lesions as well as those previously found by us in an independent series of 35 unifocal breast carcinomas, we compared a probabilistic model for likelihood of independence with unsupervised hierarchical clustering methodologies to determine the clonal relatedness of multiple tumours in breast cancer patients. We conclude that CGH analysis of multiple breast carcinomas followed by unsupervised hierarchical clustering of the genomic imbalances is more reliable than previous criteria to determine the tumours' clonal relationship in individual patients, that most ipsilateral breast carcinomas arise through intramammary spreading of a single breast cancer, and that most patients with bilateral breast carcinomas have two different diseases.

Adult↗

Plasmid-mediated dissemination of blaKPC-3 and multidrug resistance genes among different species of Klebsiella.

Carbapenem resistance is a serious threat to public health because carbapenems are used as last-resort antibiotics. Carbapenem resistance gene KPC (Klebsiella pneumoniae carbapenemase) inactivates a broad range of β-lactam substrates. In this manuscript, we examined intra-host transmission of blaKPC-3 via interspecies gene transfer. Two carbapenem-resistant Klebsiella pneumoniae isolates and one Klebsiella michiganensis isolate were identified from two patients. Genetic relations of these isolates were investigated with whole-genome sequencing (WGS). Hybrid assembly of bacterial genomes showed the three isolates carried plasmids that harbor common antimicrobial resistance (AMR) gene clusters that confer multidrug-class resistance, including carbapenems. Our results suggest that AMR gene clusters are disseminated across the species as fragments rather than as complete, intact plasmids.IMPORTANCEAn antimicrobial resistance gene cluster encompassing multiple drug classes on plasmids could lead a drug-susceptible pathogen to gain multidrug resistance. Interspecies gene transfer enables K. michiganensis to become multidrug-resistant through the acquisition of clustered, plasmid-encoded resistance genes spanning multiple antibiotic classes.

Plasmids↗

Gene proximity analysis across whole genomes via PQ trees.

Permutations on strings representing gene clusters on genomes have been studied earlier by Uno and Yagiura (2000), Heber and Stoye (2001), Bergeron et al. (2002), Eres et al. (2003), and Schmidt and Stoye (2004) and the idea of a maximal permutation pattern was introduced by Eres et al. (2003). In this paper, we present a new tool for representation and detection of gene clusters in multiple genomes, using PQ trees (Booth and Leuker, 1976): this describes the inner structure and the relations between clusters succinctly, aids in filtering meaningful from apparently meaningless clusters, and also gives a natural and meaningful way of visualizing complex clusters. We identify a minimal consensus PQ tree and prove that it is equivalent to a maximal pi pattern (Eres et al., 2003) and each subgraph of the PQ tree corresponds to a nonmaximal permutation pattern. We present a general scheme to handle multiplicity in permutations and also give a linear time algorithm to construct the minimal consensus PQ tree. Further, we demonstrate the results on whole genome datasets. In our analysis of the whole genomes of human and rat, we found about 1.5 million common gene clusters but only about 500 minimal consensus PQ trees, with E. Coli K-12 and B. Subtilis genomes, we found only about 450 minimal consensus PQ trees out of about 15,000 gene clusters, and when comparing eight different Chloroplast genomes, we found only 77 minimal consensus PQ trees out of about 6,700 gene clusters. Further, we show specific instances of functionally related genes in two of the cases.

Algorithms↗