Showing posts with label ancient folds. Show all posts
Showing posts with label ancient folds. Show all posts

Sunday, November 30, 2008

A challenge for the biochemist- the priming problem

All cellular life forms and many DNA viruses, phages and plasmids use a primase to synthesize a short RNA primer with a free 3' OH group that is subsequently elongated by a DNA polymerase. The existence of this process is one of the most fascinating unsolved problems. Essentially, it is unclear why DNA polymerases require a primer to initiate DNA replication, while RNA polymerases do not. From an evolutionary perspective, although DNA polymerases have evolved on multiple occasions independently, and a variety of independent solutions have evolved to address the priming problem, in no case is the DNA polymerase rid of a primer. Could this reflect something more fundamental?

First, let us review the known solutions to the priming problem.
  1. In bacteria and archaeo-eukaryotes (a term clubbing archaea and eukaryotes) that possess DNA polymerases of unrelated folds, two unrelated families of primases synthesize RNA primers. Bacteria possess a DNAG-like primase of the Toprim fold, whereas in a comprehensive sequence-structure analysis, we showed that the archaeo-eukaryotic primase (AEP) belongs to the RRM fold (click here to read).
  2. In certain plasmids and phages/viruses, a primpol protein performs both the RNA polymerase (for primer synthesis) and DNA polymerase (for DNA synthesis) activities. Primpols belong to two distinct families, one that is experimentally confirmed and belongs to the archaeo-eukaryotic primase superfamily, and the second, which awaits experimental verification, belongs to the TV-Pol family.
  3. Retroelements (including retroviruses and other reverse-transcriptase based DNA mobile elements) prime DNA replication by using a tRNA that provides a free 3' OH that is used for elongation by the reverse transcriptase (RNA-dependent DNA polymerase).
  4. In adenoviruses and the φ29 family of bacteriophages,a hydroxyl group is provided by the side-chain of an amino acid of the genome attached terminal protein to which nucleotides are added by the DNA polymerase to form a new strand.
  5. Another solution to the priming problem is seen in several families of DNA viruses, such as parvoviruses, geminiviruses and circoviruses, and many phages and plasmids. All of these replicate their DNA by rolling circle replication (RCR). Here, the RCR endonuclease (RCRE) creates a nick in one of the DNA strands. The 5' end of the nicked strand is transferred to a tyrosine residue on the nuclease, and the free 3' OH group is elongated by a DNA polymerase for the new strand synthesis.
  6. The only known exception of a DNA polymerase lacking a primer, is the reverse transcriptase of the Mauriceville plasmid that uses the 3' tRNA-like structure of the parent mRNA/pRNA to de novo synthesize a daughter DNA strand.
Thus although DNA polymerases have evolved on multiple occasions independently, they don't seem to have rid themselves of the need for a primer.

Hypothesis. In our study on the evolutionary history of the archaeo-eukaryotic primases, we speculated that these observations suggests a strong constraint against ‘invention’ of de novo initiation of DNA synthesis which, probably, stems from fundamental chemical differences between ribo- and deoxyribonucleotides, rather than a frozen evolutionary accident that maintains a primer. We proposed that the inefficiency in de novo DNA synthesis by DNA polymerases may, at least in part, be due to a competing futile reaction of 3'->5' nucleotide cyclization while using deoxyribonucleotides. Given the tendency of diverse, unrelated RNA polymerases to initiate de novo strand synthesis, it seems likely that this problem does not arise with ribonucleotides. As support for this hypothesis we note that the DNA cyclases are specifically related to DNA polymerases and have evolved from the latter on multiple occasions independently.

To the best of our knowledge this hypothesis has not yet been tested. You can read about our comprehensive study on the archaeo-eukaryotic primases by clicking here.

Tuesday, November 25, 2008

One protein family, many insights: The TV-pol story

Using sensitive sequence analysis methods, we recently discovered and characterized a divergent member of the DNA polymerase I superfamily /Superfamily A DNA polymerase that we denote the Transposon-Virus polymerase (TV-Pol) family. These proteins are found in a wide range of bacteria and their prophages, phages, the chloroplast of the alga Nephroselmis, and in the Sputnik virus (virophage of Mimivirus).

Using gene neighborhood analysis we show that the TV-pol genes are components of mobile elements and might be involved in replicative transposition. As evidence, we detected a recent transposition event in Brucella melitensis 16M, of a transposon with a direct repeat that contains the TV-Pol gene, and a γδ-resolvase.More specifically, based on their frequent fusion to D5-helicases (e.g. V13 of the Sputnik virus) we speculate that TV-Pol proteins are primase-polymerases (primpols), like some members of the archaeo-eukaryotic primases (To learn more about AEP-like primpols click here).

Additional interesting insights
  • Superfamily A DNA polymerases contain a HTH domain.Upon defining the structural core of the DNA polymerase I superfamily, we noted that these proteins are distinguished by the presence of a HTH-domain within the fingers of the RRM-like palm domain. This HTH contains the highly conserved RxxxK motif characteristic of this superfamily, and potentially interacts with the elongating daughter strand.
  • The thumb, palm and fingers probably existed as independent polypeptides at an early point in the evolution of the superfamily A polymerases. Given the presence of distinct globular folds in the fingers (HTH) and palm domain (RRM) and also the displacement of the coiled coil thumb by a distinct globular domain in a TV-Pol of Gemmata obscuriglobus, an early stage in the evolution of this superfamily can be conceived where these three units were present on different polypeptides and then fused to give the Superfamily A DNA polymerases.
  • The predicted primpol activity of the TV-Pols throws light on the origins of the T7-like RNA polymerases. The T7-like DNA-dependent RNA polymerases are members of the superfamily A DNA polymerases. Their origins can now be understood in light of the discovery of TV-pols, where a primpol ancestor that had both DNA and RNA polymerase activity might have possibly contributed to the T7-like RNA polymerases. This view is also supported by experimental studies that have shown some members of the T7-like RNA polymerases to function as primases.
  • The Sputnik virophage could have evolved from a mobile element. Based on the gene contexts of the TV-Pol gene in the Sputnik virophage, we speculate that the virus may have arose from a a TV-Pol containing transposase, which subsequently acquired a DNA-packaging HerA-FtsK ATPase and virion proteins from a distinct viral source.
You can read the open access version of this study by clicking here.

Friday, October 31, 2008

What is the biochemistry of Pupylation?

Recently, a remarkable study showed that Mycobacteria have a distinct "ubiquitin-like" system in which a small protein, Pup, is transferred to the ε-amino groups of lysines in target proteins. These experiments also implicated a gene neighbor, the PafA protein, in this activity. How this was mediated was a mystery.

Using sensitive sequence and structure analysis methods, we unified the PafA proteins to the glutamine synthetase (or carboxylate-amine/ammonia ligase) superfamily. In particular the PafA proteins are closer to the γ-glutamyl-cysteine synthetases. This unification provides a simple explanation for the reaction mechanism of Pupylation by PafA (the Pup ligase).

First the Pup ligase catalyzes an ATP-dependent phosphorylation of the γ-carboxylate of glutamate followed by ligation with the ε-amino group of lysines in target proteins with the formation of an amide linkage.

In Pups with a terminal glutamine instead of a glutamate (e.g. Mycobacterial Pup), the glutamine is first deamidated and converted to glutamate. Given the similar chemistry, we propose that this reaction too might be catalyzed by the Pup ligase. Our analysis suggests that pupylation is a bacterial innovation that emerged from proteins involved in amino acid (glutamine) and cofactor (glutathione) biosynthesis. The parallels with the ubiquitination system are striking in which the ubiquitin system evolved in bacteria from a system involved in cofactor (Moco) and amino acid (cysteine) biosynthesis. Thus the similiarities in pupylation and ubiquitination represent a remarkable case of convergent evolution.

Additional points of interest
  • Pup is predicted to be a α-helical protein with an extended tail and is not related to ubiquitin.
  • The pupylation system is present in most actinobacteria, and also sporadically in verrucomicrobia, nitrospirae, deltaproteobacteria and planctomycetes. In all cases both Pup and the Pup-ligase are immediate gene neighbors.
  • Barring a few exceptions, gene neighborhoods reveal two paralogs of Pup ligases suggesting that they function as heterodimers. In species with only one copy, they would function as homodimers. Note Mycobacteria have two copies of PafA corresponding to genes Rv2097c and Rv2112c.
  • Gene neighborhoods also reveal that the actinobacterial pupylation genes are neighbors of the archaeal-type proteasomal AAA+ ATPases and proteases (NTN hydrolase superfamily) in line with prior studies that in these bacteria pupylated proteins are targeted for degradation. However, this may not be always so. The Pup ligases of deltaproteobacteria and planctomycetes are remarkable in that they have 4 transmembrane helices inserted within the core domain and are also neighbors of membrane proteins. In these bacteria, the pupylation system might target membrane proteins.
  • We also detected the prokaryotic homolog of the proteasomal chaperone PAC2 in the gene neighborhood of some actinobacterial Pupylation genes. This is the first report of a prokaryotic proteasomal chaperone and given the absence of other proteasomal chaperone subunits, it appears that PAC2 is the most ancient proteasomal chaperone (see the separate blog on PAC2).
  • Could other members of this family catalyze analogous reactions? In this quest, we detected two other previously uncharacterized families of proteins that belong to the glutamine synthetase superfamily. However, their domain contexts and gene neighborhoods suggest that they may be involved in glutathione or related peptide secondary metabolites biosynthesis.
For more details, you can read the open access version of the paper. Click here to access it. For latest updates on pupylation click here

Wednesday, October 15, 2008

Uncovering the origins and diversity of the E1-superfamily of proteins

The E1-superfamily of proteins are central to ubiquitin (Ub) conjugation, biosynthesis of cysteine, thiamine and MoCo and several secondary metabolites. Yet the diversity and evolutionary history of these proteins was poorly understood. Recently, we undertook a comprehensive study of the E1 superfamily and uncovered several interesting and surprising insights.

Watch this page for a more detailed summary. For now, please click here to access the full article.

Sunday, September 30, 2007

RAGNYA : a novel fold found in functionally diverse nucleic acid, nucleotide & peptide-binding proteins


One of our principal research objectives is to derive a natural classification of the protein universe by unifying diverse protein superfamilies. However, the detection of relationships between these superfamilies is often non-trivial due to extensive divergence or variations, like circular permutations, in their structural scaffolds. This is particularly prevalent in numerous small folds involved in binding of nucleic-acids/nucleotides. One such alpha+beta fold that we recently identified was the RAGNYA fold that includes a diverse group of proteins principally involved in nucleic acid, nucleotide or peptide interactions. Members of the fold include the Ribosomal proteins L3 and L1, the GYF domain, DNA-recombination proteins of the NinB family from caudate bacteriophages, the C-terminal DNA-interacting domain of the Y-family DNA polymerases, the uncharacterized enzyme AMMECR1, the siRNA silencing repressor of tombusviruses, tRNA Wybutosine biosynthesis enzyme Tyw3p, DNA/RNA ligases and related nucleotidyltransferases and the Enhancer of rudimentary proteins. This fold exhibits three distinct circularly permuted versions and is composed of an internal repeat of a unit with two-strands and a helix. We show that despite considerable structural diversity in the fold, its representatives show a common mode of nucleic acid or nucleotide interaction via the exposed face of the sheet.
Click here to read the paper