Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “multiple clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

TFBScluster: a resource for the characterization of transcriptional regulatory networks.

SUMMARY: One major challenge of the post-sequencing era of the human genome project will be the functional annotation of the non-coding portion of the genome, in particular gene regulatory sequences. We have developed a new web-based tool, TFBScluster, which performs genome-wide identification of transcription factor binding site clusters that are conserved in multiple mammalian genomes. Clusters representing candidate gene regulatory elements can be filtered further, based on the presence or absence of additional user-defined DNA sequence motifs or by constraining the orientation or order of binding sites. Comprehensive results files, returned by email, are designed to facilitate experimental validation of computationally identified candidate gene regulatory sequences. TFBScluster, therefore, has the potential to contribute to deciphering transcriptional networks that regulate a wide range of mammalian developmental processes.

Algorithms↗

Identification of a novel HIV-1 complex circulating recombinant form (CRF18_cpx) of Central African origin in Cuba.

BACKGROUND: Analysis of partial pol and env sequences have indicated a high diversity of HIV-1 genetic forms in Cuba, including two potential novel circulating recombinant forms (CRF): U/H and D/A. OBJECTIVES: To determine whether U/H recombinant viruses from Cuba, detected in 7% of samples, represent a novel HIV-1 CRF, and to identify non-Cuban viruses related to this recombinant form. METHODS: Near full-length genome amplification was carried out by nested polymerase chain reaction in four overlapping DNA segments of two epidemiologically unlinked viruses in uncultured peripheral blood mononuclear cells. The sequences were analysed phylogenetically. Recombinant structures and phylogenetic relationships were analysed by bootscanning and by maximum likelihood. Searches for related viruses in databases were initially based on sequence homology and sharing of signature nucleotides. RESULTS: Both Cuban viruses clustered uniformly in bootscans all along the genome with each other and with a virus from Cameroon, CM53379, indicating that all three represent the same recombinant form. Their genome comprised multiple segments clustering with subtypes A1, F, G, H and K, as well as segments failing to cluster with recognized subtypes. The newly defined CRF, designated CRF18_cpx, was phylogenetically related in partial segments to CRF13_cpx, CRF04_cpx and 36 additional viruses, most of them from Central Africa. One of the viruses from Cameroon, sequenced in the near full-length genome, was a CRF18_cpx/subtype G secondary recombinant. CONCLUSIONS: A novel HIV-1 complex circulating recombinant form (CRF18_cpx) has been identified that is circulating in Cuba and Central Africa.

Africa, Central↗

The homeodomain of PDX-1 mediates multiple protein-protein interactions in the formation of a transcriptional activation complex on the insulin promoter.

Activation of insulin gene transcription specifically in the pancreatic beta cells depends on multiple nuclear proteins that interact with each other and with sequences on the insulin gene promoter to build a transcriptional activation complex. The homeodomain protein PDX-1 exemplifies such interactions by binding to the A3/4 region of the rat insulin I promoter and activating insulin gene transcription by cooperating with the basic-helix-loop-helix (bHLH) protein E47/Pan1, which binds to the adjacent E2 site. The present study provides evidence that the homeodomain of PDX-1 acts as a protein-protein interaction domain to recruit multiple proteins, including E47/Pan1, BETA2/NeuroD1, and high-mobility group protein I(Y), to an activation complex on the E2A3/4 minienhancer. The transcriptional activity of this complex results from the clustering of multiple activation domains capable of interacting with coactivators and the basal transcriptional machinery. These interactions are not common to all homeodomain proteins: the LIM homeodomain protein Lmx1.1 can also activate the E2A3/4 minienhancer in cooperation with E47/Pan1 but does so through different interactions. Cooperation between Lmx1.1 and E47/Pan1 results not only in the aggregation of multiple activation domains but also in the unmasking of a potent activation domain on E47/Pan1 that is normally silent in non-beta cells. While more than one activation complex may be capable of activating insulin gene transcription through the E2A3/4 minienhancer, each is dependent on multiple specific interactions among a unique set of nuclear proteins.

Animals↗

Structural organization of the immunoglobulin heavy chain locus in the channel catfish: the IgH locus represents a composite of two gene clusters.

Two structurally-related genomic clusters of catfish immunoglobulin heavy chain gene segments are known. The first gene cluster contains DH and JH segments, as well as the C region exons encoding the functional Cmu. The second gene cluster contains multiple VH gene segments representing different VH families, a germline-joined VDJ, a single JH segment, and at least two pseudogene Cmu exons. It was not known whether these gene clusters were linked, nor was the organization or the location of VH segments associated within the first gene cluster known. Pulsed-field gel electrophoresis studies have been used to determine the structural organization of these gene clusters. Restriction mapping studies show that the two gene clusters are closely linked; the second gene cluster is located upstream from the first with the Cmu regions within the clusters separated by about 725kb. The clusters are in the same relative transcriptional orientation, and the results indicate that the complete IgH locus spans no more than 1000kb and may be as small as 750-800kb. VH gene segments are located both upstream and downstream of the pseudo-Cmu exons; however, no VH gene segments that hybridized with the VH specific probes were detected downstream of the functional Cmu. These studies coupled with earlier sequence analyses indicate that the catfish IgH locus arose from a massive internal duplication event. Subsequent gene rearrangement within the duplicated cluster likely resulted in the presence of the germline VDJ and the deletion of intervening V, D and J segments. Transposition by a member of the Tc1/mariner family of transposable elements appears to have led to the disruption of the duplicated Cmu.

Animals↗

Clustering of RNA secondary structures with application to messenger RNAs.

There is growing evidence of translational gene regulation at the mRNA level, and of the important roles of RNA secondary structure in these regulatory processes. Because mRNAs likely exist in a population of structures, the popular free energy minimization approach may not be well suited to prediction of mRNA structures in studies of post-transcriptional regulation. Here, we describe an alternative procedure for the characterization of mRNA structures, in which structures sampled from the Boltzmann-weighted ensemble of RNA secondary structures are clustered. Based on a random sample of full-length human mRNAs, we find that the minimum free energy (MFE) structure often poorly represents the Boltzmann ensemble, that the ensemble often contains multiple structural clusters, and that the centroids of a small number of structural clusters more effectively characterize the ensemble. We show that cluster-level characteristics and statistics are statistically reproducible. In a comparison between mRNAs and structural RNAs, similarity is observed for the number of clusters and the energy gap between the MFE structure and the sampled ensemble. However, for structural RNAs, there are more high-frequency base-pairs in both the Boltzmann ensemble and the clusters, and the clusters are more compact. The clustering features have been incorporated into the Sfold software package for nucleic acid folding and design.

Base Pairing↗

Phosphate availability regulates biosynthesis of two antibiotics, prodigiosin and carbapenem, in Serratia via both quorum-sensing-dependent and -independent pathways.

Serratia sp. ATCC 39006 produces two secondary metabolite antibiotics, 1-carbapen-2-em-3-carboxylic acid (Car) and the red pigment, prodigiosin (Pig). We have previously reported that production of Pig and Car is controlled by N-acyl homoserine lactone (N-AHL) quorum sensing, with synthesis of N-AHLs directed by the LuxI homologue SmaI, and is also regulated by Rap, a member of the SlyA family. We now describe further characterization of the SmaI quorum-sensing system and its connection with other regulatory mechanisms. We show that the genes responsible for biosynthesis of Pig, pigA-O, are transcribed as a single polycistronic message in an N-AHL-dependent manner. The smaR gene, transcribed convergently with smaI and predicted to encode the LuxR homologue partner of SmaI, was shown to possess a negative regulatory function, which is uncommon among the LuxR-type transcriptional regulators. SmaR represses transcription of both the pig and car gene clusters in the absence of N-AHLs. Specifically, we show that SmaIR exerts its effect on car gene expression via transcriptional control of carR, encoding a pheromone-independent LuxR homologue. Transcriptional activation of the pig and car gene clusters also requires a functional Rap protein, but Rap dependency can be bypassed by secondary mutations. Transduction of these suppressor mutations into wild-type backgrounds confers a hyper-Pig phenotype. Multiple mutations cluster in a region upstream of the pigA gene, suggesting this region may represent a repressor target site. Two mutations mapped to genes encoding pstS and pstA homologues, which are parts of a high-affinity phosphate transport system (Pst) in Escherichia coli. Disruption of pstS mimicked phosphate limitation and caused concomitant hyper-production of Pig and Car, which was mediated, in part, through increased transcription of the smaI gene. The Pst and SmaIR systems define distinct, yet overlapping, regulatory circuits which form part of a complex regulatory network controlling the production of secondary metabolites in Serratia ATCC 39006.

4-Butyrolactone↗

RUMINA: high-throughput deduplication of unique molecular identifiers for amplicon and whole-genome sequencing with enhanced error correction.

MOTIVATION: Unique molecular identifiers (UMIs) are widely used in next-generation sequencing to enable accurate molecular counting and error correction. However, challenges remain in accurately collapsing UMI clusters, especially when read counts are low or sparse read clusters arise from barcode sequencing errors. RESULTS: We present RUMINA, a Rust-based pipeline for UMI-aware deduplication and error correction, optimized for both amplicon and shotgun sequencing. RUMINA supports multiple UMI cluster strategies, alongside majority-rule read selection independent of mapping quality, as well as discrete handling of 1-2 read clusters, paired-end merging, and read-length stratification. Benchmarking using simulated HIV population sequencing data and real-world iCLIP and TCR datasets showed that RUMINA improves ultra-low frequency SNV detection (0.01%-1%), reduces false positives, enhances reproducibility, and processes sequencing data up to 10-fold faster than existing tools. By integrating UMI- and sequence-level correction in a high-performance framework, RUMINA offers a fast, scalable, and robust solution for UMI-enabled sequencing workflows. AVAILABILITY AND IMPLEMENTATION: RUMINA is implemented in Rust and distributed as open-source code and precompiled binaries. Source code and installation instructions are available at https://github.com/greninger-lab/rumina. Documentation associated with this manuscript is available at https://github.com/greninger-lab/rumina_paper.

High-Throughput Nucleotide Sequencing↗

Contents and relationship of elements in human hair for a non-industrialised population in Poland.

The concentrations of 11 elements: Pb, Mn, Fe, Cd, Cu, Ni, Cr, Zn, Na, K and Ca in hair were determined by AAS. Hair samples (n = 266) were collected between 1990 and 1994 from inhabitants of the Silesian Beskid in the south of Poland (non-industrialised region). The effects of age (1-30 years old, 31-80 years old), sex (male, female) and colour of hair on the heavy metal levels were determined. Using statistical methods of cluster analysis, multiple regression analysis and factor analysis we obtained information concerning relations among metals in the hair. The strongest relations between metals in the hair are as follows: Fe-Mn, Cr-K and Cd-Pb in the first cluster and Zn-Ni in the second cluster. For our population (n = 266, non-industrialised region in Poland) we obtained a factor loading > 0.7. Factor 1 was contributed by Na and K; Factor 2 by Pb, Cd, Mn and Fe; Factor 3 by Ca (a negative correlation); Factor 4 by Ni; Factor 5 by Cu; Factor 6 by Co, Cd and K; and Factor 7 by Cr, Pb, Mn, Fe and K. These seven factors explain 77.7% variance. We obtained linear multiple dependence (P < 0.05) as follows: Mn = f (Cd, Fe, Ca); Na = f (Zn, K, -Ca); K = f (-Zn, Cr, Na); Pb = f (Cd, Zn, Cr, Fe); Zn = f (Ni, Na, -Cr, -K); Fe = f (Pb, Mn); Cr = f (Pb, K, -Zn); Co = f (Cd, Cr); and Ca = f (Mn, -Zn, -Na). These relations can be useful to explain relationships among the metals in man.

Adolescent↗

Clusters of internally primed transcripts reveal novel long noncoding RNAs.

Non-protein-coding RNAs (ncRNAs) are increasingly being recognized as having important regulatory roles. Although much recent attention has focused on tiny 22- to 25-nucleotide microRNAs, several functional ncRNAs are orders of magnitude larger in size. Examples of such macro ncRNAs include Xist and Air, which in mouse are 18 and 108 kilobases (Kb), respectively. We surveyed the 102,801 FANTOM3 mouse cDNA clones and found that Air and Xist were present not as single, full-length transcripts but as a cluster of multiple, shorter cDNAs, which were unspliced, had little coding potential, and were most likely primed from internal adenine-rich regions within longer parental transcripts. We therefore conducted a genome-wide search for regional clusters of such cDNAs to find novel macro ncRNA candidates. Sixty-six regions were identified, each of which mapped outside known protein-coding loci and which had a mean length of 92 Kb. We detected several known long ncRNAs within these regions, supporting the basic rationale of our approach. In silico analysis showed that many regions had evidence of imprinting and/or antisense transcription. These regions were significantly associated with microRNAs and transcripts from the central nervous system. We selected eight novel regions for experimental validation by northern blot and RT-PCR and found that the majority represent previously unrecognized noncoding transcripts that are at least 10 Kb in size and predominantly localized in the nucleus. Taken together, the data not only identify multiple new ncRNAs but also suggest the existence of many more macro ncRNAs like Xist and Air.

Animals↗

p53 gene mutations in Bowen's disease in Koreans: clustering in exon 5 and multiple mutations.

We analyzed the p53 protein expression and gene mutations to evaluate the role of ultraviolet radiation or other carcinogens, and possible racial differences in 17 samples from 12 Korean patients with Bowen's disease. A simple microdissection technique was used to collect the tumor cells selectively. p53 protein expression was found in eight of 17 (47%) samples. Abnormalities in polymerase chain reaction (PCR)-single-strand conformation polymorphism (SSCP) analysis were observed in 16 (94%) samples. A total of 14 missense mutations were detected in eight (47%) samples; 11 were clustered in exon 5 and the remaining three were located in exon 8. UV-like mutations were seen in five of 14 (36%) mutations, but no CC to TT transitions, UV-fingerprint mutations were observed. Multiple mutations were present in two cases and double mutation in a single case. Each lesion in multiple Bowen's disease showed different mutations and was suggested to be of different clonal origins. TP53-loss of heterozygosity (LOH) was detected in four out of 15 (27%) informative samples. Clustering of mutations in exon 5 suggests the role of another carcinogen in Koreans or Asians other than the UVR. Microdissection would increase the detection rate of the p53 gene mutations and LOH not only in skin cancer but also in precancerous lesions.

Adult↗

Differences in the subgingival microbiota of Swedish and USA subjects who were periodontally healthy or exhibited minimal periodontal disease.

BACKGROUND: Previous studies have shown differences in the mean proportions of subgingival species in samples from periodontitis subjects in different countries, which may relate to differences in diet, genetics, disease susceptibility and manifestation. The purpose of the present investigation was to determine whether there were differences in the subgingival microbiotas of Swedish and American subjects who exhibited periodontal health or minimal periodontal disease. METHOD: One hundred and fifty eight periodontally healthy or minimally diseased subjects (N Sweden=79; USA=79) were recruited. Subjects were measured at baseline for plaque, gingivitis, BOP, suppuration, pocket depth and attachment level at 6 sites per tooth. Subgingival plaque samples taken from the mesial aspect of each tooth at baseline were individually analyzed, in one laboratory, for their content of 40 bacterial species using checkerboard DNA-DNA hybridization (total samples=4345). % DNA probe counts comprised by each species was determined for each site and averaged across sites in each subject. Significance of differences in proportions of each species between countries was determined using ancova adjusting for age, mean pocket depth, gender and smoking status. p values were adjusted for multiple comparisons. Cluster analysis was performed to group subjects based on their subgingival microbial profiles using a chord coefficient and an average unweighted linkage sort. RESULTS: On average, all species were detected in samples from subjects in both countries. After adjusting for multiple comparisons, 5 species were in significantly higher adjusted mean percentages in Swedish than American subjects: Actinomyces naeslundii genospecies 1 (9.7, 3.3); Streptococcus sanguis (2.5, 1.2); Eikenella corrodens (1.7, 1.0); Tannerella forsythensis (3.5, 2.3) and Prevotella melaninogenica (6.3, 1.8). Leptotrichia buccalis was in significantly higher adjusted mean percentages in American (5.5) than Swedish subjects (3.0). Cluster analysis grouped 121 subjects into 8 microbial profiles. Twenty four of the 40 test species examined differed significantly among cluster groups. Five clusters were dominated by American subjects and 2 clusters by Swedish subjects. Fifty eight of 79 (73%) of the Swedish subjects fell into 1 cluster group dominated by high proportions of A. naeslundii genospecies 1, Prevotella nigrescens, T. forsythensis and P. melaninogenica. Other clusters were characterized by high proportions of Actinomyces gerencseriae, Veillonella parvula, Capnocytophaga gingivalis, Prevotella intermedia, Eubacterium saburreum, L. buccalis and Neisseria mucosa. CONCLUSIONS: The microbial profiles of subgingival plaque samples from Swedish and American subjects who exhibited periodontal health or minimal disease differed. The heterogeneity in subgingival microbial profiles was more pronounced in the American subjects, possibly because of greater genetic and microbiologic diversity in the American subjects sampled.

Adult↗

Metabolic syndrome amplifies the age-associated increases in vascular thickness and stiffness.

OBJECTIVES: We sought to evaluate whether the clustering of multiple components of the metabolic syndrome (MS) has a greater impact on these vascular parameters than individual components of MS. BACKGROUND: Intima-media thickness (IMT) and vascular stiffness have been shown to be independent predictors of adverse cardiovascular events. The MS is defined as the clustering of three or more of the cardiovascular risk factors of dysglycemia, hypertension, dyslipidemia, and obesity. METHODS: Carotid IMT and stiffness were derived via B-mode ultrasonography in 471 participants from the Baltimore Longitudinal Study on Aging, who were without clinical cardiovascular disease and not receiving antihypertensive therapy. RESULTS: The MS conferred a disproportionate increase in carotid IMT (+16%, p < 0.0001) and stiffness (+32%, p < 0.0001), compared with control subjects. Multiple regression models, which included age, gender, smoking, low-density lipoprotein, as well as each individual component of MS as continuous variables, showed that MS was an independent determinant of both IMT (p = 0.002) and stiffness (p = 0.012). The MS was associated with a greater prevalence of subjects whose values were in the highest quartiles of IMT, stiffness, or both. CONCLUSIONS: Even after taking into account each individual component of MS, the clustering of at least three of these components is independently associated with increased IMT and stiffness. This suggests that the components of MS interact to synergistically impact vascular thickness and stiffness. Future studies should examine whether the excess cardiovascular risk associated with MS is partly mediated through the amplified alterations in these vascular properties.

Adult↗

Species specificity in rodent pheromone receptor repertoires.

The mouse V1R putative pheromone receptor gene family consists of at least 137 intact genes clustered at multiple chromosomal locations in the genome. Species-specific pheromone receptor repertoires may partly explain species-specific social behavior. We conducted a genomic analysis of an orthologous pair of mouse and rat V1R gene clusters to test for species specificity in rodent pheromone systems. Mouse and rat have lineage-specific V1R repertoires in each of three major subfamilies at these loci as a result of postspeciation duplications, gene loss, and gene conversions. The onset of this diversification roughly coincides with a wave of Line1 (L1) retrotranspositions into the two loci. We propose that L1 activity has facilitated postspeciation V1R duplications and gene conversions. In addition, we find extensive homology among putative V1R promoter regions in both species. We propose a regulatory model in which promoter homogenization could ensure that V1R genes are equally competitive for a limiting transcriptional structure to account for mutually exclusive V1R expression in vomeronasal neurons.

Animals↗

Characterizing magnetoencephalographic spike sources in children with tuberous sclerosis complex.

PURPOSE: Tuberous sclerosis complex (TSC) often causes medically intractable seizures. Magnetoencephalography (MEG) localizes epileptiform discharges. To evaluate the use of MEG spike sources (MEGSSs) for localizing epileptic zones in TSC patients, we characterized MEGSSs and correlated them to EEG and magnetic resonance imaging (MRI) results. METHODS: We analyzed data from seven children who underwent prolonged video-EEG, MEG, and MRI. We classified MEGSSs as clusters (six or more spike sources, 1 cm between sources regardless of number of sources). RESULTS: A single, unilateral cluster with additional scatters occurred in two patients; these predominantly lateralized dipoles correlated to prominent tubers on MRI and ictal/interictal EEG zones. Bilateral clusters with scatters existed in two patients; cluster locations partly overlapped multiple prominent tubers. These patients also had bilateral or diffuse interictal discharges, bilateral or generalized seizures, and changing seizure types and EEG findings. Only bilateral scatters occurred in three patients; scatters partly overlapped EEG interictal/ictal-onset regions; one patient had coexisting generalized seizures. In one patient with equally bilateral scatters, scatters overlapped a prominent tuber and interictal/ictal-onset zones in the right frontal region. CONCLUSIONS: MEG contributes to information from EEG and MRI for localizing epileptogenic zones in children with TSC. A single cluster with scatters in a unilateral hemisphere predicts a primary epileptogenic zone or hemisphere; bilateral or multiple clusters indicate bilateral primary or potential epileptogenic zones; and bilateral scatters without clusters may indicate epileptogenic zones that are hidden within extensive areas of scattered MEGSSs.

Brain Mapping↗

Annotating nonspecific SAGE tags with microarray data.

SAGE (serial analysis of gene expression) detects transcripts by extracting short tags from the transcripts. Because of the limited length, many SAGE tags are shared by transcripts from different genes. Relying on sequence information in the general gene expression database has limited power to solve this problem due to the highly heterogeneous nature of the deposited sequences. Considering that the complexity of gene expression at a single tissue level should be much simpler than that in the general expression database, we reasoned that by restricting gene expression to tissue level, the accuracy of gene annotation for the nonspecific SAGE tags should be significantly improved. To test the idea, we developed a tissue-specific SAGE annotation database based on microarray data (). This database contains microarray expression information represented as UniGene clusters for 73 normal human tissues and 18 cancer tissues and cell lines. The nonspecific SAGE tag is first matched to the database by the same tissue type used by both SAGE and microarray analysis; then the multiple UniGene clusters assigned to the nonspecific SAGE tag are searched in the database under the matched tissue type. The UniGene cluster presented solely or at higher expression levels in the database is annotated to represent the specific gene for the nonspecific SAGE tags. The accuracy of gene annotation by this database was largely confirmed by experimental data. Our study shows that microarray data provide a useful source for annotating the nonspecific SAGE tags.

Cell Line↗

Cluster analysis of consensus water sites in thrombin and trypsin shows conservation between serine proteases and contributions to ligand specificity.

Cluster analysis is presented as a technique for analyzing the conservation and chemistry of water sites from independent protein structures, and applied to thrombin, trypsin, and bovine pancreatic trypsin inhibitor (BPTI) to locate shared water sites, as well as those contributing to specificity. When several protein structures are superimposed, complete linkage cluster analysis provides an objective technique for resolving the continuum of overlaps between water sites into a set of maximally dense microclusters of overlapping water molecules, and also avoids reliance on any one structure as a reference. Water sites were clustered for ten superimposed thrombin structures, three trypsin structures, and four BPTI structures. For thrombin, 19% of the 708 microclusters, representing unique water sites, contained water molecules from at least half of the structures, and 4% contained waters from all 10. For trypsin, 77% of the 106 microclusters contained water sites from at least half of the structures, and 57% contained waters from all three. Water site conservation correlated with several environmental features: highly conserved microclusters generally had more protein atom neighbors, were in a more hydrophilic environment, made more hydrogen bonds to the protein, and were less mobile. There were significant overlaps between thrombin and trypsin conserved water sites, which did not localize to their similar active sites, but were concentrated in buried regions including the solvent channel surrounding the Na+ site in thrombin, which is associated with ligand selectivity. Cluster analysis also identified water sites conserved in thrombin but not trypsin, and vice versa, providing a list of water sites that may contribute to ligand discrimination. Thus, in addition to facilitating the analysis of water sites from multiple structures, cluster analysis provides a useful tool for distinguishing between conserved features within a protein family and those conferring specificity.

Animals↗

Clustering of risk behaviors with cigarette consumption: A population-based survey.

OBJECTIVE: This study assessed clustering of multiple risk behaviors (i.e., low leisure-time physical activity, low fruits/vegetables intake, and high alcohol consumption) with level of cigarette consumption. METHODS: Data from the 2002 Swiss Health Survey, a population-based cross-sectional telephone survey assessing health and self-reported risk behaviors, were used. 18,005 subjects (8052 men and 9953 women) aged 25 years old or more participated. RESULTS: Smokers more frequently had low leisure time physical activity, low fruits/vegetables intake, and high alcohol consumption than non- and ex-smokers. Frequency of each risk behavior increased steadily with cigarette consumption. Clustering of risk behaviors increased with cigarette consumption in both men and women. For men, the odds ratios of multiple (> or =2) risk behaviors other than smoking, adjusted for age, nationality, and educational level, were 1.14 (95% confidence interval: 0.97, 1.33) for ex-smokers, 1.24 (0.93, 1.64) for light smokers (1-9 cigarettes/day), 1.72 (1.36, 2.17) for moderate smokers (10-19 cigarettes/day), and 3.07 (2.59, 3.64) for heavy smokers (> or =20 cigarettes/day) versus non-smokers. Similar odds ratios were found for women for corresponding groups, i.e., 1.01 (0.86, 1.19), 1.26 (1.00, 1.58), 1.62 (1.33, 1.98), and 2.75 (2.30, 3.29). CONCLUSIONS: Counseling and intervention with smokers should take into account the strong clustering of risk behaviors with level of cigarette consumption.

Adult↗

Diplotype trend regression analysis of the ADH gene cluster and the ALDH2 gene: multiple significant associations with alcohol dependence.

The set of alcohol-metabolizing enzymes has considerable genetic and functional complexity. The relationships between some alcohol dehydrogenase (ADH) and aldehyde dehydrogenase (ALDH) genes and alcohol dependence (AD) have long been studied in many populations, but not comprehensively. In the present study, we genotyped 16 markers within the ADH gene cluster (including the ADH1A, ADH1B, ADH1C, ADH5, ADH6, and ADH7 genes), 4 markers within the ALDH2 gene, and 38 unlinked ancestry-informative markers in a case-control sample of 801 individuals. Associations between markers and disease were analyzed by a Hardy-Weinberg equilibrium (HWE) test, a conventional case-control comparison, a structured association analysis, and a novel diplotype trend regression (DTR) analysis. Finally, the disease alleles were fine mapped by a Hardy-Weinberg disequilibrium (HWD) measure (J). All markers were found to be in HWE in controls, but some markers showed HWD in cases. Genotypes of many markers were associated with AD. DTR analysis showed that ADH5 genotypes and diplotypes of ADH1A, ADH1B, ADH7, and ALDH2 were associated with AD in European Americans and/or African Americans. The risk-influencing alleles were fine mapped from among the markers studied and were found to coincide with some well-known functional variants. We demonstrated that DTR was more powerful than many other conventional association methods. We also found that several ADH genes and the ALDH2 gene were susceptibility loci for AD, and the associations were best explained by several independent risk genes.

Adult↗