The CGN allows relating every amino acid residue position to equivalent positions in different Gα proteins. In short, the numbering scheme provides a relative address of every residue with respect to the domain, secondary structure element, and the specific amino acid position within the secondary structure element, similar to a postal address in the following format: (D).S.P - The top level D refers to the structural domain and is optional (G for the catalytic GTPase domain, H for the helical domain). S refers to the respective consensus secondary structure element (SSE) and P refers to the relative position within the consensus SSE in the human paralog alignment (Figure 1).
We recommend using the CGN as a superscript for individual positions of specific G protein. For instance, position Phe336 in Gαi can be represented as Phe336H5.8, which is the 8th position from the start of the consensus Helix 5 in the human reference (Domain optional: Phe336G.H5.8). The corresponding position in Gαs would be Phe376H5.8. If no particular Gα protein is referred to, PheH5.8 is sufficient to describe the position.
Insertions, such as found in Gα orthologs from different species, are identified with (D).S.P-i, where i stands for the number of inserted residues, for instance Arg334H4.27-2 which stands for the second amino acid of the insertion after helix H4 found in pufferfish (Tetraodon nigroviridis) Gαs. It should be stressed that the secondary structure elements of the CGN are those defined in a consensus secondary structure from averaging over all 80 Gα structures and not from the individual structures (see below). Although insertions occur mainly in the loop regions, the exact borders of secondary structure elements in an individual protein may be slightly different, may depend on the structure visualization software and the applied secondary structure optimizations or energy minimizations, and might vary for different individual Gα structures.
Rationale for developing the CGN Residue-based standard numbering systems to compare different members of a protein family permit the transfer of information from template to target sequences. Such common numbering systems have been described for instance for immunoglobulins1, kinases2, GPCRs3, and β-lactamases4. Many of these numbering system do not give an explicit idea of the secondary structure element in which the residue resides in. The Ballesteros-Weinstein (BW) numbering scheme permits comparing any amino acid residue in the structurally conserved transmembrane (TM) helices of GPCRs by counting from the most conserved position within each TM helix3. However, such a definition of a ‘most conserved position’ in a SSE is relative and depends on the used sequence set. Several of the numbering schemes do not include loops although residues within loops are often functionally the most relevant ones. Thus our strategy was to construct a reference system based on a large structure and sequence comparison that can instantly provide a structurally interpretable standard reference for any position within a sequence that is scalable for incorporating additional sequence in the future.
Integration of sequence data into the CGN For the integration of sequence data, we compared ~950 Gα sequences. For this reason, we first created 16 independent one-to-one ortholog alignments for each Gα type in human. These 16 independent alignments were then cross-referenced by an alignment of all ‘canonical’ 16 human Gα paralogs (Figure 2; alignments provided here). Since we find that each human Gα sequence is a good representative for each Gα type in terms of conservation of secondary structure elements and insertions/deletions, we decided to create a human-centric reference system and define the corresponding position P always relative to the human paralog alignment (Figure 1). This means if there is a deletion in one of the orthologs, position P ‘jumps’ positions in the CGN (e.g. extended h4s6 loop for Gαs sequences). Insertions in orthologs are annotated with P-i, whereas i represents the number of inserted amino acid residues after the last equivalent position in the human paralog alignment. Summarizing, the hierarchal referencing of 16 ortholog alignments to the human paralog alignment allowed us to build reliable, low-gap alignments, which permits easily relating any Gα protein position with a distantly related G protein and the method is extendable to include to any newly discovered Gα protein, even if it contains long insertions/deletions.
Integration of structure data into the CGN To include information on the secondary structure from different protein structures into the reference system, we first mapped each position in each Gα structure to the sequence alignments. The challenge is that protein structures often use author-specific individual residue numbers, as they can comprise mutations, insertions, deletions, or fusion tags required for protein purification or stability. To overcome this, we decided to use the beta version of the SIFTS consortium5 to reference the sequence of each PDB structure as accurately as possible to its respective Uniprot positions, and thus the respective sequence in each alignment. SIFTS annotates chimeras, single point mutations, and missing electron density and thus ensures higher accuracy than defining corresponding residues based on, for instance, a structural superposition of all PDBs or a Blast search6. Combined, the structure and sequence comparison allowed us to define corresponding positions in different Gα proteins. Since the secondary structure topology of Gα proteins is conserved, we defined ‘consensus secondary structure elements’ for easier interpretation of the CGN code. 36 consensus secondary-structure elements (including loops) were obtained from a secondary structure assignment and structural alignment of all 80 identified structures of G proteins (Figure 1).
1 Honegger, a. & Plückthun, a. Yet another numbering scheme for immunoglobulin variable domains: an automatic modeling and analysis tool. Journal of molecular biology 309, 657-670, doi:10.1006/jmbi.2001.4662 (2001).
2 Brooijmans, N., Chang, Y. W., Mobilio, D., Denny, R. A. & Humblet, C. An enriched structural kinase database to enable kinome-wide structure-based analyses and drug discovery. Protein science : a publication of the Protein Society 19, 763-774, doi:10.1002/pro.355 (2010).
3 Ballesteros, J. A. & Weinstein, H. Receptor Molecular Biology. Vol. 25 (Elsevier, 1995).
4 Galleni, M. et al. Standard numbering scheme for class B beta-lactamases. Antimicrobial agents and chemotherapy 45, 660-663, doi:10.1128/AAC.45.3.660-663.2001 (2001).
5 Velankar, S. et al. SIFTS: Structure Integration with Function, Taxonomy and Sequences resource. Nucleic Acids Res 41, D483-489, doi:10.1093/nar/gks1258 (2013).
6 Altschul, S. F., Gish, W., Miller, W., Myers, E. W. & Lipman, D. J. Basic local alignment search tool. J Mol Biol 215, 403-410, doi:10.1016/S0022-2836(05)80360-2 (1990).