Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multigene families”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

On the evolution of multigene families.

Multigene families are classified into three groups: small families as exemplified by hemoglobin genes of mammals; middlesize multigene families, by genes of mammalian histocompatibility antigens; and large multigene families, by variable region genes of immunoglobulins. Facts and theories on these evolving multigene families are reviewed, with special reference to the population genetics of their concerted evolution. It is shown that multigene families are evolving under continued occurrence of unequal (but homologous) crossing-over and gene conversion, and that mechanisms for maintaining genetic variability are totally different from the conventional models of population genetics. Thus, in view of widespread occurrence of multigene families in genomes of higher organisms, the evolutionary theory based mainly on change of gene frequency at each locus would appear to need considerable revision.

Biological Evolution↗

Antibody genes and other multigene families.

The multigene family is a unit of chromosomal organization. Its gene members are closely linked, homologous in sequence, and have overlapping functions. Multigene families can be divided into three categories - simple-sequence, multiplicational, and information - by a variety of structural and functional criteria. Multigene families exhibit two novel evolutionary features - coincidental evolution and rapid change in family size - that suggest that they all share one or more evolutionary mechanisms. Natural selection cannot act directly on individual genes in a family because of their identical or overlapping functions; hence selection must operate on the family as a whole or on blocks of genes within the family. The mechanism(s) for coincidental evolution expand out variant genes within a family so they can be acted on by natural selection, and accordingly, permit multigene families to evolve adaptively. The close linkage of the genes in a family appears to be a consequence of the fact that their ocntrol and evolutionary mechanisms may only operate on tandemly linked genes. New multigene families may evolve from a single gene or from other multigene families. In addition to evolving new functions, the latter mode of evolution generate a new multigene family whose members are preadapted to interact with those of the old family. These family interactions can lead to the evolution of more sophisticated molecular machines or to the regulation of one family by the second. Multigene families may be large or small. The three categories of multigene families allow potential multigene families to be identified and they suggest specific experimental approaches for the study of new families. Some of the most interesting genetic systems under investigation today are known or potential informational multigene families. This is not fortuitous in that many of the most interesting aspects of phenotype are complex ones with correspondingly complex genetic, evolutionary, and regulatory requirements. One of the frontiers in modern genetics is the identification characterization, and understanding of informational multigene families.

Alleles↗

The organization, expression, and evolution of antibody genes and other multigene families.

The multigene family is a unit of chromosomal organization. Its gene members are closely linked, homologous in sequence, and have overlapping functions. Multigene families can be divided into three catagories: simple-sequence, multiplicational, and informational-by a variety of structural and functional criteria. Multigene families exhibit two novel evolutionary features-coincidental evolution and rapid change in family size-that suggest that they all share one or more evolutionary mechanisms. Natural selection cannot act directly upon individual genes in a family because of their identical or overlapping functions; hence selection must operate upon the family as a whole or upon blocks of genes within the family. The mechanism(s) for coincidental evolution expands out variant genes within a family so they can be acted upon by natural selection and, accordingly, permits multigene families to evolve adaptively. The control mechanisms in multiplicational families appear to promote the rapid expression of many gene copies. In contrast, the regulatory mechanisms of informational families promote the selection, expression, and amplification of appropriate units of information. The close linkage of the genes in a family appears to be a consequence of the fact that their control and evolutionary mechanisms may only operate on tandemly linked genes. New multigene families may evolve from a single gene or from other multigene families. In addition to evolving new functions, the latter mode of evolution generates a new multigene family whose members are preadapted to interact with those of the old family. These family interactions can lead to the evolution of more sophisticated molecular machines or to the regulation of one family by a second. Multigene families may be large or small. The three catagories of multigene families allow potential multigene families to be identified, and they suggest specific experimental approaches for the study of new families. Some of the most interesting genetic systems under the investigation today are known or potential informational multigene families. This is not fortuitous in that many of the most interesting aspects of phenotype are complex ones with correspondingly complex genetic, evolutionary, and regulatory requirements. One of the frontiers in modern genetics is the identification, characterization, and understanding of informational multigene families.

Alleles↗

Equinatoxins, pore-forming proteins from the sea anemone Actinia equina, belong to a multigene family.

The multigene family of equinatoxins, pore-forming proteins from sea anemone Actinia equina, has been studied at the protein and gene levels. We report the cDNA sequence of a new, sphingomyelin inhibited equinatoxin, EqtIV. The N-terminal sequences of natural Eqt I and III were also determined, confirming two isoforms of EqtI, differing at position 13. The number of Eqt genes determined by Southern blot hybridization was found to be more than five, indicating that Eqts belong to a multigene family.

Amino Acid Sequence↗

Human non-histone chromosomal protein HMG-17: identification, characterization, chromosome localization and RFLPs of a functional gene from the large multigene family.

The multigene family of chromosomal protein HMG-17 is the largest known human retropseudogene family. A functional gene was identified and isolated by screening cDNA-selected genomic clones with a set of 5 oligonucleotides whose sequence corresponded to regions in which the sequence of the retropseudogenes differed from that of the cDNA and which did not span previously identified exon/intron junctions. A 7195 bp genomic fragment containing 6 exons, ranging in size from 30 to 817 bp, two of which encode the entire DNA binding domain of the protein, was sequenced. The gene has features which are typical to "housekeeping" genes and is characterized by a very high content of G + C residues in a 1.4 kb fragment starting 500 bp from the cap site and by an "HTF" island in the 5' region. Transcriptional regulatory signals, exon/intraon boundaries and features characteristic of "housekeeping" genes are evolutionary conserved between the human and chicken genes. The HMG-17 gene was localized to human chromosome 1p12-34. RFLP's useful for further mapping were detected. The experimental evidence presented leads to the assumption that the gene characterized is the only functional human HMG-17 gene.

Amino Acid Sequence↗

Expression of members of the Saccharomyces cerevisiae hsp70 multigene family.

The hsp70 multigene family of Saccharomyces cerevisiae is a complex multigene family, composed of members exhibiting complex patterns of regulation. Expression of some members is induced after a heat shock, whereas expression of others is repressed. Some members of the family are expressed during exponential growth. One gene, SSA3, shows an unusual pattern of expression during approach to stationary phase. While most RNAs decrease in abundance, SSA3 RNA levels dramatically increase. The constitutive expression of SSA3 in cells lacking adenylate cyclase activity suggests that cAMP modulates SSA3 expression.

Fungal Proteins↗

Complexity and expression of the glutamine synthetase multigene family in the amphidiploid crop Brassica napus.

In the amphidiploid genome of oilseed rape (Brassica napus) the diploid ancestral genomes of B. campestris and B. oleracea have been merged. As a result of this crossing event, all gene loci, gene families, or multigene families of the A and C genome types encoding a certain protein are now combined in one plant genome. In the case of the multigene family for glutamine synthetase, the key enzyme of nitrogen assimilation, six different cDNA sequences were isolated from leaf and root specific libraries. One sequence pair (BnGSL1/BnGSL2) was characterized by the presence of amino-terminal transit peptides, a typical feature of all nuclear encoded chloroplast proteins. Two other cDNA pairs (BnGSR1-1/BnGSR1-2 and BnGSR2-1/BnGSR2-2) with very high homology between each other were found in a root specific cDNA library and represent protein subunits for cytosolic glutamine synthetase isoforms. Comparative PCR amplifications of genomic DNA isolated from B. napus, B. campestris and B. oleracea followed by sequence-specific restriction analyses of the PCR products permitted the assignment of the cDNA sequences to either the A genome type (BnGSL1/BnGSR1-1/BnGSR2-1) or the C genome type (BnGSL2/BnGSR1-2/BnGSR2-2). Consequently, the ancestral GS genes of B. campestris and B. oleracea are expressed simultaneously in oilseed rape. This result was also confirmed by RFLP (restriction fragment length polymorphism) analysis of RT-PCR products. In addition, the different GS genes showed tissue specific expression patterns which are correlated with the state of development of the plant material. Especially for the GS genes encoding the cytosolic GS isoform BnGSR2, a marked increase of expression could be observed after the onset of leaf senescence.

Blotting, Northern↗

Identification and localization of a neurally expressed member of the plakoglobin/armadillo multigene family.

The plakoglobin/armadillo multigene family comprises many proteins widely differing in sizes and functions which have in common a variable number of tandemly repeated arm sequences of about 42 amino acids (aa). In a search for proteins with sequence homology to the desmosomal-plaque-associated arm-repeat-containing protein, plakophilin 1, we have identified a novel plakoglobin/armadillo protein. This new member of the multigene family is predominantly, if not exclusively, expressed in neural and neuroendocrine tissues, hence the name neural plakophilin-related arm-repeat protein (NPRAP). The murine cDNA codes for a protein of 1247 aa, with a predicted molecular weight of 135 kDa and a pI of 7.57. The orthologous human protein differs only in a few aa, indicative of the evolutionary stability of NPRAP. In human and murine cDNAs, we have found different transcripts of the NPRAP gene, suggesting that in each species the protein exists in at least two isoforms. The NPRA protein contains three different regions: a 528-aa amino-terminal "head" domain, including a potential coiled-coil-forming alpha-helix segment, a central domain with 10 imperfect arm-repeat units, and a 212-aa carboxy-terminal "tail" domain. By aa sequence, NPRAP is highly homologous to three proteins: p120cas, p0071 and ARVCP, which represent a distinct subgroup within the plakoglobin/armadillo family. By in situ hybridization and immunofluorescence microscopy using NPRAP-specific antibodies, we have demonstrated NPRAP and its mRNA in the perikarya of various kinds of CNS neurons in embryonic and adult mice, but minimal amounts have also been detected by immunoblot analysis in some other tissues containing neural or neuroendocrine elements. We have not seen significant enrichment of NPRAP at cell junctions or in nuclei. Possible NPRAP functions are discussed and the correlation of NPRAP synthesis with neuronal differentiation processes is emphasized.

Alternative Splicing↗

The olfactory multigene family.

A novel multigene family has been identified that is likely to encode odorant receptors on olfactory sensory neurons. Further studies on this gene family are likely to shed light on the molecular mechanisms underlying information coding in the mammalian olfactory system. This review is also published in Current Opinion in Neurobiology 1992, 2:282-288.

Amino Acid Sequence↗

The olfactory multigene family.

A novel multigene family has been identified that is likely to encode odorant receptors on olfactory sensory neurons. Further studies on this gene family are likely to shed light on the molecular mechanisms underlying information coding in the mammalian olfactory system. This review is also published in Current Opinion in Genetics and Development 1992, 2:467-473.

Amino Acid Sequence↗

Phylogenetic and structural relationships of the PR5 gene family reveal an ancient multigene family conserved in plants and select animal taxa.

Pathogenesis-related group 5 (PR5) plant proteins include thaumatin, osmotin, and related proteins, many of which have antimicrobial activity. The recent discovery of PR5-like (PR5-L) sequences in nematodes and insects raises questions about their evolutionary relationships. Using complete plant genome data and discovery of multiple insect PR5-L sequences, phylogenetic comparisons among plants and animals were performed. All PR5/PR5-L protein sequences were mined from genome data of a member of each of two main angiosperm groups-the eudicots (Arabidoposis thaliana) and the monocots (Oryza sativa)-and from the Caenorhabditis nematode (C. elegans and C. briggsase). Insect PR5-L sequences were mined from EST databases and GenBank submissions from four insect orders: Coleoptera (Diaprepes abbreviatus and Biphyllus lunatus), Orthoptera (Schistocerca gregaria), Hymenoptera (Lysiphlebus testaceipes), and Hemiptera (Toxoptera citricida). Parsimony and Bayesian phylogenetic analyses showed that the PR5 family is paraphyletic in plants, likely arising from 10 genes in a common ancestor to monocots and eudicots. After evolutionary divergence of monocots and eudicots, PR5 genes increased asymmetrically among the 10 clades. Insects and nematodes contain multiple sequences (seven PR5-Ls in nematodes and at least three in some insects) all related to the same plant clade, with nematode and insect sequences separating as two clades. Protein structural homology modeling showed strong similarity among animal and plant PR5/PR5-Ls, with divergence only in surface-exposed loops. Sequence and structural conservation among PR5/PR5-Ls suggests an important and conserved role throughout the evolutionary divergence of the diverse organisms from which they reside.

Amino Acid Sequence↗

A new member of the balbiani ring multigene family in the dipteran Chironomus tentans consists of a single-copy version of a unit repeated in other gene family members.

The known Balbiani ring (BR) multigene family members in the dipteran Chironomus tentans encode salivary gland secretory proteins in the size range between 38 and 1,000 kDa. The proteins interact to form protein fibers used by the aquatic larvae to spin feeding and protective larval tubes or pupation tubes. Here, we describe a new BR multigene family member, the sp17 gene, which codes for an 89-amino-acid-long protein with a relative mobility of 17k. The gene has a high content of charged amino acid residues and consists of two structurally different halves. Five regularly spaced cysteine codons are present in the 5' half while the 3' half contains five proline codons. These two different halves exhibit similarities to the C and SR regions, respectively, which form the tandemly repeated units in the about 40-kb-long BR genes and which also, in different versions, are the building blocks of all genes in the BR multigene family. In this multigene family, encoding interacting structural proteins, the long BR genes with their 125-150 tandemly arranged repeat units as well as the short sp17 gene with its single-copy version of such a repeat unit, have therefore evolved from a common ancestor.

Amino Acid Sequence↗

Linkage disequilibrium in human ribosomal genes: implications for multigene family evolution.

Members of the rDNA multigene family within a species do not evolve independently, rather, they evolve together in a concerted fashion. Between species, however, each multigene family does evolve independently indicating that mechanisms exist which will amplify and fix new mutations both within populations and within species. In order to evaluate the possible mechanisms by which mutation, amplification and fixation occur we have determined the level of linkage disequilibrium between two polymorphic sites in human ribosomal genes in five racial groups and among individuals within two of these groups. The marked linkage disequilibrium we observe within individuals suggests that sister chromatid exchanges are much more important than homologous or nonhomologous recombination events in the concerted evolution of the rDNA family and further that recent models of molecular drive may not apply to the evolution of the rDNA multigene family.

Biological Evolution↗

Structure of the multigene family of MAL loci in Saccharomyces.

Multigene families are a ubiquitous feature of eukaryotes; however, their presence in Saccharomyces is more limited. The MAL multigene family is comprised of five unlinked loci, MAL1, MAL2, MAL3, MAL4 and MAL6, any one of which is sufficient for yeast to metabolize maltose. A cloned MAL6 locus was used as a probe to facilitate the cloning of the other four functional loci as well as two partially active alleles of MAL1. Each locus could be characterized as a cluster of three genes, MALR (regulatory), MALT (maltose transport or permease) and MALS (structural or maltase), encoded by a total of about 7 kb of DNA; however, homologous sequences at each locus extend beyond the coding regions. Our results indicate that there is extensive homology among the MAL loci, especially within their maltase genes. The greatest sequence diversity occurs in their regulatory gene regions. Southern cross analyses of the cloned MAL loci indicate a single duplication of the MAL6R-homologous sequences upstream of the MAL6R gene as well as an extensive duplication of more than 10 kb at the MAL3 locus. The large repeat at the MAL3 locus results in the presence of four copies of MAL3R-homologous sequences and two copies of MAL3T-homologous sequences at that locus. Two naturally occurring inactive alleles of MAL1 show a deletion or divergence of their MALR sequences. The significance of these repeats in the evolution of the MAL multigene family is discussed.

Biological Evolution↗

Characterization of the pufferfish Takifugu rubripes apolipoprotein multigene family.

We have characterized the apolipoprotein multigene family of the pufferfish Takifugu rubripes. The pufferfish mainly contains 28-kDa, 27-kDa, and 14-kDa apolipoproteins in its plasma and was designated apo-28 kDa, apo-27 kDa, and apo-14 kDa, respectively. N-terminal amino acid sequencing revealed that pufferfish apo-28 kDa and apo-27 kDa have an identical amino acid sequence except an additional propeptide in the former; and both are homologues of apoA-I from other animals. The sequence of pufferfish apo-14 kDa is homologous to that of eel apo-14 kDa previously reported, both being apparently specific to fish. In silico screening, using the publicly available Fugu genome database confirmed the pufferfish apoA-I and apo-14 kDa genes. The database further contained the genes encoding four types of apoA-IV, one apoC-II and two types of apoE. Thus, pufferfish contains nine genes encoding apolipoprotein multigene family. Two apoA-IV and one apoE genes were tandemly arrayed and located on one scaffold. Thus two sets of these genes formed two gene clusters. The apoC-II and apo-14 kDa genes are also located on a single scaffold. apoA-I and apo-14 kDa gene transcripts were mainly expressed in liver and less abundantly in brain. The transcripts of the former gene were also observed in intestine. In contrast, the transcripts encoding four apoA-IVs, one apoC-II, and two apoEs were mainly expressed in intestine. These structural details of pufferfish apolipoproteins and tissue distribution of their gene transcripts provide a novel evidence for better understanding of evolutionary relationships of apolipoprotein multigene family.

Amino Acid Sequence↗

The evolution of the chicken sarcomeric myosin heavy chain multigene family.

This manuscript describes the chicken sarcomeric myosin heavy chain (MyHC) multigene family and how it differs from the sarcomeric MyHC multigene families of other vertebrates. Data is discussed that suggests the chicken fast MyHC multigene family has undergone recent expansion subsequent to the divergence of avians and mammals, and has been subjected to multiple gene conversion-like events. Similar to human and rodent MyHC multigene families, the chicken multigene family contains sarcomeric MyHC genes that are differentially regulated in developing embryonic, fetal, and neonatal muscles. However, unlike mammalian genes, chicken fast MyHC genes expressed in developing muscles are also expressed in mature muscle fibers as well. The potential significance of conserved and divergent sequences with the MyHC rod domain of five fast chicken isoforms that have been cloned and sequenced is also discussed.

Animals↗

Sheep CD4(+) alphabeta T cells express novel members of the T19 multigene family.

The sheep T19 multigene family contains at least 50 genes which are thought to be expressed exclusively on gammadelta T cells. The archetypal T19 molecule (represented by a full-length cattle cDNA clone termed WC1) is thought to have a relative molecular mass of about 220 000 and to contain 11 scavenger receptor cysteine rich (SRCR) repeats and a long cytoplasmic tail. In this study, purified CD4(+) and gammadeltaTCR+ sheep lymphocytes were examined by reverse transcriptase-polymerase chain reaction for the expression of T19 molecules. As expected, gammadelta T cells were found to express T19 molecules which closely resembled the archetypal form. However, CD4(+) alphabeta T cells were found to express at least two different types of T19 molecules; one resembled the previously described T19 molecules of gammadelta T cells which possessed the archetypal WC1-like structure, but a novel type of T19 variant which lacked SRCR domains 10 and 11 was also found in CD4(+) T cells but not gammadelta T cells. This novel molecule exhibited an unusual, incomplete SRCR repeat 9 joined directly to a hinge region. The transmembrane and cytoplasmic domains of this unusual T19 variant resembled the cattle T19 clone WC1, except that a complete exon within the cytoplasmic region was missing. These results, in contradistinction to existing serological data, suggest that expression of the T19 gene family is not confined to gammadelta T cells. Selected T19 genes are apparently expressed within CD4(+) T cells and possibly other lymphocytes as well.

Amino Acid Sequence↗

Interaction of selection and biased gene conversion in a multigene family.

A model of the evolutionary dynamics of a multigene family in a finite population under the joint effects of selection and (possibly biased) gene conversion is analyzed. It is assumed that the loss or fixation of a polymorphism at any particular locus in the gene family occurs on a much faster time scale than the introduction of new alleles to a monomorphic locus by gene conversion. A general formula for the fixation of a new allele throughout a multigene family for a wide class of selection functions with biased gene conversion is given for this assumption. Analysis for the case of additive selection shows that (i) unless selection is extremely weak or bias is exceptionally strong, selection usually dominates the fixation dynamics, (ii) if selection is very weak, then even a slight conversion bias can greatly alter the fixation probabilities, and (iii) if both selection and conversion bias are sufficiently small, the substitution rate of new alleles throughout a multigene family is approximately the single locus mutation rate, the same result as for neutral alleles at a single-copy gene. Finally, I analyze a fairly general class of underdominant speciation models involving multigene families, concluding for these models under weak conversion that although the probability of fixation may be relatively high, the expected time to fixation is extremely long, so that speciation by "molecular drive" is unlikely. Furthermore, speciation occurs faster by fixing underdominant alleles of the same effect at single-copy genes than by fixing the same number of loci in a single multigene family under the joint effects of selection, conversion, and drift.

Alleles↗