Educational guide
Peptide Mapping - an overview
Chapters and Articles You might find these chapters and articles relevant to this topic. Peptide Mapping Peptide mapping is usually performed on an isolated protein or a protein mixture. Identifying a protein using peptide mapping requires digesting the protei
This guide cannot diagnose a condition or recommend a personal treatment plan. Discuss medical questions with a qualified professional.
Chapters and Articles
You might find these chapters and articles relevant to this topic.
Peptide Mapping
Peptide mapping is usually performed on an isolated protein or a protein mixture. Identifying a protein using peptide mapping requires digesting the protein into peptides prior to MS analysis. Although most peptide mapping experiments use trypsin to produce peptides, other enzymes (e.g., Lys-C, Glu-C, etc.) can be used, depending on the experimental requirement. Once digested, the molecular weight of the peptides are acquired.1 These experimental masses of the peptides are compared to masses generated from an in silico digest of proteins or translated nucleic acid sequences contained within a database. A protein will be identified if several of the experimental masses match those for a specific protein in the database within a certain mass tolerance (Figure 1). Because the accuracy between the experimental and in silico masses is critical to obtaining the correct protein identification, it is best to acquire the peptide map on a high mass measurement accuracy instrument, such as a time-of-flight (TOF) configuration. The greater the number of matches between the experimental and database peptide masses, the higher the confidence in the protein’s identification.
FIGURE 1. Peptide mapping for protein identification using mass spectrometry (MS). In the first step, the protein is proteolytically digested (usually with trypsin) and the experimental masses of the peptides are measured using MS. The sequences of the proteins within the selected database are digested in silico based on the specificity of the enzyme used. The masses of these peptides are calculated and theoretical mass spectra are constructed. The correct protein is identified based on the closest match between the experimental and theoretical mass spectra.
Beyond the availability of a database containing the necessary protein sequences, software is also required to turn the raw MS data into protein identifications. Fortunately, there are many freely available software programs for analyzing peptide mapping data (Table 1). At a minimum, the programs require an input that includes the experimental peptide ion list and database to compare this list to. Some programs allow additional information, such as isoelectric point (pI), molecular weight (MW), and source organism that can be used to identify the protein. Some of this additional information (e.g., pI and MW) is determined using other protein chemistry techniques primarily one- or two-dimensional gel electrophoresis. Although all of the programs listed in Table 1 are sufficient for analyzing peptide mapping data, Mascot2 stands out as the most popular, owing to its longevity, and many MS users have built a familiarity with this software. Probably the primary reasons for selecting a specific peptide mapping software program include it being part of the MS purchase, its ease of use, and its integration with other software programs used in the analysis of MS data.
TABLE 1. Software Available for Protein Identification Using Peptide Mass Fingerprinting
| Software Program | URL |
|---|---|
| MultiIdent | http://web.expasy.org |
| Mascot | http://www.matrixscience.com |
| MS-Fit | http://prospector.ucsf.edu |
| PepMAPPER | http://www.nwsr.manchester.ac.uk |
| MassWiz | http://masswiz.igib.res.in |
| Protein Lynx | http://www.matrixscience.com |
| ProFound | http://prowl.rockefeller.edu |
URL: https://www.sciencedirect.com/science/article/pii/B9780123944467000169
a Peptide Mapping/LC/MS
Peptide mapping is a powerful tool for the analysis of the primary structure of a protein.5 This method typically takes advantage of one or more specific proteases that cleave the protein into smaller peptides, which are then separated and analyzed by reversed-phase liquid chromatography (RP-HPLC), sometimes in conjunction with mass spectrometry (MS).
When peptide mapping is coupled with MS, the amino acid sequence of each peptide can be individually confirmed. In addition, covalent modifications such as glycosylation, oxidation, or deamidation can be identified in a site-specific manner (see below).
Once a peptide map is characterized using online MS, the chromatographic profile alone can serve as a routine analytical tool to monitor the protein's primary structure and covalent modifications, and is often used for batch release or stability testing of biopharmaceuticals. However, whenever in-depth characterization of a protein is needed, such as that required for comparability studies or reference material characterization, the peptide map should be coupled with MS to ensure a thorough examination of all peptides in the map.
In developing a peptide mapping procedure for characterization and/or as routine assay, many considerations need to be taken at each step. They are discussed as follows.
i Reduction and Alkylation
Cysteine residues can complicate proteolytic digestion, because of either disulfide scrambling or structural hindrance to proteolytic sites. Therefore, modification of cysteine residues by reduction and alkylation typically increases the efficiency and robustness of the proteolytic digestions. Reduction and alkylation is typically performed in a denaturing buffer (e.g., 6 M guanidine chloride) at slightly alkaline pH (7.5–8.4). Common reductants are dithiothreitol (DTT) or tris(2-carboxyethyl)phosphine (TCEP), and common alkylants include iodoacetate, iodoacetamide, or N-ethylmaleimide. The optimal conditions for reduction and alkylation (achieving complete reduction and alkylation, without overalkylating) can be evaluated by LC/MS and SDS-PAGE of the reduced/alkylated protein. After reduction and alkylation, excess alkylant may be neutralized by another addition of reducing agent to quench the alkylation reaction and avoid overalkylation.
ii Desalt and Dilution
Denaturants such as guanidine chloride used in reduction and alkylation inhibit many proteases, and must be removed or diluted out prior to the addition of the protease. Removal by buffer exchange can be achieved by dialysis, desalting filtration, or a desalting column. Dilution is simpler and less time-consuming; however, residual denaturant typically still inhibits protease activity, leading to higher levels of missed cleavages. On the other hand, the lower protease activity also may lead to less nonspecific cleavages, and a cleaner baseline in the peptide map. In our experience, the level of miscleavages tends to be reproducible, and does not affect the usefulness of the map for monitoring structural integrity and covalent modifications. Therefore, we find the dilution option to be reasonable for a routine peptide mapping method.
iii Protease Selection
Typical choices for peptide mapping are site-specific proteases, including trypsin, Lys-C, Asp-N, and Glu-C. Due to the different specificities, different proteases are expected to cleave the protein into different peptides. Very small peptides tend to be lost in the flow-through of reversed-phase HPLC (RP-HPLC), reducing the sequence coverage, so proteases that do not generate too many di- or tripeptides are preferred. Very large peptides may have either chromatographic issues (recovery, peak size), or potentially reduce the ability to detect small changes chromatographically or by MS. For example, a large, modified peptide may not fully resolve from its unmodified form. The ability to generate moderate sized proteolytic peptides (< 5000 Da) may be of particular interest in certain sensitive regions of the sequence, such as the complementarity determining region (CDR) of an antibody, where small changes may be expected to impact efficacy. On the other hand, larger peptides result in simpler maps that tend to be more robust and less time-consuming to analyze. Therefore, the choice of protease is a balancing act with the ultimate goal of good sequence coverage (> 90%), ability to detect small structural changes, and digestion robustness.
iv Digestion
Digestion is typically performed at the optimal pH of the chosen protease for 2 h to overnight. A time course study is useful to determine optimal digest time that ensures robustness of digestion, minimal artifacts (such as deamidation), and practicality. Another factor that influences completeness of digestion is enzyme:substrate ratio (E:S). Too much protease may lead to nonspecific cleavages from low levels of other protease impurities in the commercial protease preparation, while too little can lead to underdigestion. Both situations can lead to issues with the reproducibility and robustness of the peptide map. Digests are quenched by acidification and/or addition of reductant. Acidification not only stops the protease action (this is needed to ensure reproducibility of digestion time), but also reduces likelihood of deamidation and cyclization of N-terminal Gln as a result of prolonged exposure to high pH of the digestion reaction mixture.
v Chromatographic Separation
Most peptide mapping methods utilize RP-HPLC to separate the proteolytic peptides, although cation exchange chromatography (CEX) has also been used for this purpose. However, RP-HPLC is the preferred method because of its ability to directly interface with MS, making peak identification much easier. A typical peptide mapping method uses a C18 column and a gradient of water and acetonitrile with ion-pairing agents such as TFA. Initial method development involves screening various columns and gradients to identify parameters that lead to well resolved peaks with good peak area recovery. MS is often needed to aid method development, to ensure important peptides (e.g., those containing stability-indicating modifications, N-terminal and C-terminal peptides, and glycosylated peptides) are well separated from other species, and can be readily monitored.
vi Evaluation of Reproducibility and Robustness
If peptide mapping will be used for batch release or stability testing,6 the method needs to be evaluated for reproducibility and robustness (e.g., column and protease lot-to-lot consistency, sample stability over time). If any quantitative acceptance criteria are necessary, such a method also needs to undergo qualification and validation.
vii Detection and Characterization of Separated Peptides
The chromatographically separated peptides are most commonly detected by UV absorbance, especially for routine testing and method development. With a dual-wavelength detector, UV absorbance at 214 nm is universally applicable for all peptides, while absorbance at 280 nm detects only tryptophan- or tyrosine-containing peptides. The photodiode array (PDA) detector has the ability to measure the UV absorbance spectra of each peak in the chromatogram, and is sometimes useful for troubleshooting. For example, a contaminant peak representing a method-, process-, container closure-, or product-related substance may have an unusual chromophore that leads to a UV absorbance spectrum distinguishable from that of a peptide.
For identification of the peaks in a peptide map, online RP-HPLC/MS (or LC/MS) analysis is needed. Accurate mass determinations with modern high-resolution electrospray (ESI) mass spectrometers often provides the ability to distinguish and confidently assign all separated peptides from a recombinant protein in an efficient manner.7 Commercial ESI hybrid quadrupole time-of-flight (QTOF) mass spectrometers8,9 with 10,000 resolving power (i.e., mass/Δmass using bovine insulin) typically afford accurate mass measurements for peptides to within 0.003% of the respective theoretical masses in a stable, reproducible manner through the use of lock mass10 and/or flight tube temperature monitoring11 correction strategies, following external mass calibration. For mass spectral data analysis, the protein therapeutic amino acid sequence is proteolyzed in silico to generate a sorted list of respective peptides and theoretical molecular masses (monoisotopic values). For each detected peptide, the experimental mass is matched to a particular theoretical value for positive peptide identification, provided that the mass error is less than the 0.003% specified tolerance.
In some cases, where the peptide mass alone is insufficient for definitive assignment, or if site-specific information about a particular amino acid position or modification is required, the peak may be collected (i.e., the eluate containing the peptide of interest) and subjected to further analysis, such as off-line high-resolution tandem MS (MS/MS) sequencing with collision-induced dissociation (CID) or electron-transfer dissociation (ETD), Edman N-terminal sequencing, or amino acid analysis (AAA).12 The more traditional quadrupole- and ion trap-based mass spectrometers with unit mass resolution generally provide mass accuracies of 0.03% or better, which still is sufficient to distinguish and identify all separated peptides in most protein therapeutics.13 However, with this group of instruments, more confident peptide assignments are obtained by either online LC/MS in conjunction with in-source CID, provided that the peptides are well separated, or online LC/MS/MS with data-dependent scan functions for more complicated peptide separations. In this latter approach, after each recurring mass analysis, one to five abundant peptide precursor ions are selected by the acquisition software for CID in quadrupole instruments or CID and/or ETD in ion-trap instruments, based on user-defined selection criteria and the results of previous scan functions. In recent years, the use of data-independent LC/MS/MS has gained momentum14,15; with this approach, quadrupole ion selection is not used and all peptide precursor ions are subjected to CID. However, the mass spectral fragmentation data of closely spaced precursor ions in the LC separation are multiplexed and specialized, proprietary software is required to extract the respective amino acid sequences for peptide identification.
In recent years, a new generation of ultrahigh-resolution ESI-based mass spectrometers have become available that afford mass accuracies of less than 0.0005%, for even more confident peptide identification via LC/MS and LC/MS/MS.16–19 Additionally, the newer ultrahigh-resolution ESI-QTOF instruments feature greater sensitivities and faster scan rates (i.e., up to 20 spectra per second), which are well suited for fast chromatography applications with ultrahigh performance liquid chromatography (UHPLC) and capillary electrophoresis (CE).
Last, thorough identification of the peaks across the peptide map separation with MS is a very laborious, multihour, hands-on process, whether LC/MS with accurate mass measurements or LC/MS/MS with gas-phase peptide ion fragmentation is employed. Reliable peptide map informatics software, specific to protein pharmaceutical characterization, such as BiopharmaLynx and ProteinLynx Global Server (Waters Corporation), is becoming available as the biotherapeutics research and development sector matures.20 BiopharmaLynx interfaces with high-resolution Waters QTOF instruments and utilizes accurate mass measurements from LC/MS to identify the peptides represented by each peak in a peptide map—literally in minutes. Additional specificity for peptide identification is gained from amino acid sequence tags if gas-phase fragmentation is conducted in a second, alternating scan function. The BiopharmaLynx algorithm requires the target recombinant protein sequence, instrument type, protease and cleavage specificity, the number of allowed miscleavages, a specified error tolerance between 0.0005% and 0.003% (or better), and the list of possible posttranslational, storage-induced, and method-induced modifications, in order to create a comprehensive theoretical list of unmodified peptides, modified peptides, miscleaved peptides, and protease-derived peptides, including their masses, for automatically elucidating all of the peptides in an experimental peptide map separation.
URL: https://www.sciencedirect.com/science/article/pii/B9780123756800000085
Peptide-mapping is a powerful tool in unraveling the nature and identities of residues participating in covalent adduct formation. In order to determine which residue of bovine liver ECH was responsible for nucleophilic attack, [3-3H]MCPF-CoA was incubated with bovine liver ECH. The inactivated enzyme was processed with trypsin and two radiolabeled tryptic fragments were isolated by reverse-phase HPLC.14 The N-terminal sequence of both peptides were identical (Y-A-L-G-G-G-X-E-L), indicating that they were probably the result of incomplete digestion. The radioisotope was predominantly associated with residue ‘X’, whose identity could not be determined, presumably due to its covalent modification with MCPF-CoA. A sequence comparison of ECHs from difference sources indicated that the peptide sequence was part of the active site, and included Glu-115 (bovine liver ECH numbering, equivalent to Glu-144 in rat liver ECH), which was adjacent to the modified residue. Since the sequence of bovine liver ECH was not known, the identity of residue ‘X’ was tentatively assigned as Cys based on the sequences of other homologous ECHs.14 In an effort to conclusively identify the reactive nucleophile, a bovine liver cDNA library was constructed and used to isolate the coding sequence for the enzyme. Cloning and sequencing of the cDNA insert indeed confirmed that residue ‘X’ was Cys-114 (bovine liver ECH numbering; the corresponding sequence number for rat liver ECH, including the signal sequence, is Cys-143).14 Therefore, with the deduction of the mechanism of inactivation using isotopically labeled MCPF-CoA analogues and the tentative identification of the reactive nucleophile, the mechanistic picture regarding the inactivation of bovine liver ECH seemed complete.
URL: https://www.sciencedirect.com/science/article/pii/S0968089602003334
Peptide Mapping Techniques
The technique of peptide mapping has been useful in identifying subtle structural components between related proteins (31, 32). For example, peptide mapping has been elegantly used to demonstrate that mammalian β-adrenergic receptors differ in subunit structure, depending on the receptor subtype (β1 or β2). These adrenergic receptors have been shown to conserve their structures between tissues of the same animal, but they differ across species in proportion to their phylogenetic differences (33).
Peptide mapping experiments are performed following the methods of Bordier and Crettol-Jarvinen (31) with minor modifications. CRF receptors from rat anterior pituitary and frontal cortex membranes are affinity cross-linked, treated with N -glycanase, and subjected to SDS-PAGE as a first dimension as described above. The labeled bands of interest are excised with a scalpel and incubated in 125 mM Tris-HCl, 0.1% SDS, pH 6.8 at 22°C, for 45 min. One piece each of pituitary or cortical deglycosylated proteins is then apposed to the second-dimension electrophoresis gel between the glass plates of a 2 mm gel (2.5 cm, 6% stacking and 11 cm, 15% separating gel), taking care not to trap any air bubbles between the excised acrylamide pieces and the stacking gel. The pieces are then fixed onto the stacking gel by overlaying with a solution of 1% agarose in 125 mM Tris-HCl, 0.1% SDS, pH 6.8. The agarose is allowed to solidify (~20 min), at which time 1.5 ml of the specific protease solution is added in 125 mM Tris-HCl, 0.1% SDS, and 10% glycerol, pH 6.8 at 22°C. Separate gels are run on pairs of pituitary and cortical proteins using either papain or Staphylococcus aureus V8 enzymes (obtained from Worthington Biochem. Corp., Freehold, NJ; final concentrations 300 µg/ gel). Finally, the protease solutions are overlayed with running buffer as described, and the samples are electrophoresed overnight at a constant current of 2 mA. After the bromphenol blue dye reaches the separating gel (~20 hr), the fragments are separated electrophoretically at a constant current of 10 mA for 18 hr. Gels are then dried and exposed to Kodak X-AR film as described. Prestained molecular weight markers are treated in an identical manner but in the absence of any proteases. Autoradiograms are routinely scanned using the PC-based Loats system (Amersham, Arlington Heights, IL) and the optical density profiles in conjunction with the calculated migration values are used to determine differences or similarities in the peptide maps.
In performing the limited proteolysis experiments outlined above, it must be made clear that the same proteases being used to cleave the CRF receptor would also cleave between specific amino acid residues on the CRF peptide itself used to covalently label the receptor protein. For example, S. aureus V8 specifically cleaves at the carboxy end of the amino acids glutamic and aspartic acid. 125I-Labeled oCRF contains three glutamic acid residues at positions 3, 17, and 20 and three aspartic acid residues at positions 9, 25, and 39 which could all serve as substrates for this enzyme. It is expected therefore, given the relatively long incubation times required for proteolysis to take place under the conditions outlined, and the fact that the 125I-label is at the Tyr0 position on the probe, that some of the covalent label would be lost owing to the nature of the probe. Moreover, it is important to keep in mind that the actual visualization of these fragments is by autoradiography and thus only those peptide fragments still attached to the 125I label can be identified. Additionally, it must also be noted that any fragments arising solely from the digestion of the probe itself (125I-labeled oCRF) will have relatively small molecular weights (<5000) and will migrate at the bottom of the acrylamide gels. During the course of the peptide mapping studies, this was indeed what was observed. Although the 125I-labeled oCRF peptide was being proteolyzed, there was adequate incorporation of the label into peptide fragments that could be visualized, and the peptide maps of both the proteins under conditions of proteolysis by two different proteases appeared to be similar (30). Thus, in both rat anterior pituitary and cerebral cortex, the CRF ligand-binding subunit appears to have the same enzyme-cleavable sites. These data strongly suggest that CRF receptors in the two tissues share structural similarities if not identity with each other.
The data thus indicate that the ligand-binding subunits of the brain and pituitary CRF receptors reside on a polypeptide of 40,000–45,000. The heterogeneity between brain and pituitary CRF receptors appears to result from differential posttranslational glycosylation of the native protein and not from inherent differences in the protein structure. Consequently, the molecular weights of the native affinity cross-linked receptor subunits are greatly overestimated owing to this heavy glycosylation. The functional significance of the microheterogeneity observed in the carbohydrate moieties of brain and anterior pituitary CRF receptors remains to be determined. In order to determine clearly the structure of the binding site for CRF to its receptor, the CRF receptor must be purified and sequenced. In this way a detailed study of the interaction of the hormone and its binding site as well as the site of interaction between the receptor and the guanine nucleotide-binding protein can be elucidated.
URL: https://www.sciencedirect.com/science/article/pii/B9780121852597500375
Tandem Mass Spectrometry
Peptide mapping is useful only for identifying isolated proteins or maybe a simple mixture of 2 ot 3 proteins. For complex mixtures, MS2 is required.3 This method can identify isolated proteins or proteins within mixtures containing up to 100,000 different species. As with peptide mapping, bottom-up identification of a protein using MS2 requires the digestion of the proteins into peptides. In MS2, peptide are collided with an inert gas and fragmented into a series of peptide ladders via a process known as collision-induced dissociation (CID). Fortunately, CID results in the fragmentation of peptide ions primarily along its backbone. Figure 2 illustrates a MS2 spectrum for a short peptide. As indicated in this figure, the distances between various y ions is equal to the molecular mass of the specific amino acids within the peptide. That fragmentation occurs primarily across amide bonds and has allowed “rules” to be devised for creating software programs for analyzing MS2 data. Most mass spectrometers yield b and y ions, which correspond to fragmentation of the amide bonds with the charge retained at the NH2 and COOH termini, respectively. Other bonds are fragmented during CID (i.e., a, c, x, and z ions); however, these are generally less intense. As with peptide mapping, the experimental MS2 spectra are compared to a database of protein sequences using the known rules for CID fragmentation of peptides. A subtle, but important, point is that the energy put into the peptides during CID is insufficient to completely dissociate every amide bond. If every amide bond fragmented, the MS spectrum would primarily contain masses equal to those of the constituent amino acids. The resultant CID fragmentation creates “ladders” of amino acid residues originating from the peptide produced, enabling the sequence of the peptide to be read much like a DNA sequencing ladder is produced by Sanger sequencing.4 Although MS2 identification is required for identifying peptides in complex mixtures, it also provides higher confident identifications for isolated proteins than can be achieved using peptide mapping. The reason for the greater confidence afforded using MS2 is that the most distinctive characteristic that identifies a protein is its amino acid sequence, which is the type of data provided by MS2.
FIGURE 2. Tandem mass spectrometry (MS2) spectrum of peptide PNQSAFTSSGLVSK. Some of the distances between various y ions corresponding to specific amino acid residues (highlighted in bold) demonstrate the distance between these ions is equal to the molecular mass of these residues.
The raw MS2 data produced by the mass spectrometer is quite complicated. Although it is possible to determine part of the peptide’s sequence manually, this procedure is very time consuming and generally fruitful for only the highest quality MS2 spectra. Considering that many mass spectrometers can produce thousands of MS2 spectra/hour, the need for high-throughput software analysis is clearly evident. The original, and still one of the most popular, software program for analyzing MS2 data is Sequest™.5 Invented in 1994 by Jimmy Eng and John Yates, Sequest initially matches the precursor ion mass to that of peptides within a database with the same nominal mass within a specified mass accuracy. The program uses known fragmentation rules to generate theoretical MS2 spectra for the possible matches. These theoretical spectra are compared to the experimental MS2 spectrum to find the sequence that provides the best correlation. The cross correlations between the experimental and theoretical spectra are ranked and reported as a cross-correlation (Xcorr) score. In addition, the differences between the Xcorr of the first and second ranked theoretical spectra are reported and as the deltaXcorr (ΔCn).
Mascot™ is also widely used for analyzing MS2 data.2 Mascot also owes much of its popularity to the fact that it is bundled as part of the software component of many types of mass spectrometers in use today. Mascot provides a probability-based assignment by conducting a statistical evaluation of the matches between the experimental and theoretical MS2 data. Both Sequest and Mascot can be used to identify post-translational modifications (PTMs). The user must indicate the possibility of the modification on specific residue types. When searching for a phosphopeptide, for example, the potential for an additional mass of 80 Da (representing a phosphate group) is applied to each serine, threonine, and tyrosine residue. The addition is applied dynamically, allowing the software program to consider the targeted residue as either modified or unmodified. Sequest can use data only in a specific file format (.dta), and Mascot is capable of utilizing several different raw data file formats, including .dta files. Fortunately, there are scripts available for converting data from a variety of different mass spectrometer instruments into a .dta format. A list of other available software programs for analyzing MS2 data is provided in Table 2.
TABLE 2. Software Available for Protein Identification by Analysis of Tandem Mass Spectrometry Data
| Software Program | URL |
|---|---|
| Sequest | http://www.thermo.com |
| Mascot | http://www.matrixscience.com |
| MS-Tag | http://prospector.ucsf.edu |
| Pep-Frag | http://prowl.rockefeller.edu |
| OMSSA | http://pubchem.ncbi.nlm.nih.gov/omssa |
| Sonar MS/MS | http://hs2.proteome.ca/prowl/sonar/sonar_cntrl.html |
| X!Tandem | http://www.thegpm.org/tandem |
| Crux | http://noble.gs.washington.edu/proj/crux |
URL: https://www.sciencedirect.com/science/article/pii/B9780123944467000169
Peptide Mapping and Proteomic Analysis
Aspartic acid has an average occurrence of about 5% in all proteins, so Asp-N is widely used in peptide mapping and proteomic analysis [4,14,15]. Its specificity also complements those of trypsin, endoproteinase Lys-C and other proteases. For example, the average occurrence of Asp, Arg and Lys is about 5, 5, and 6% in all proteins, respectively; digestion with Asp-N, therefore, generally leads to longer and fewer peptides than tryptic cleavage [14,15]. In a comparison study of large-scale protein sequencing methods using multiple proteases, the Asp-N digestion of complex protein mixtures generated peptides of optimal length that are favorable for electron-based fragmentation detection methods, i.e. electron capture dissociation (ECD) and electron transfer dissociation (ETD) [14,15]. Asp-N can also be used for in-gel digestions alone or combined with other proteases, such as Glu-C.
URL: https://www.sciencedirect.com/science/article/pii/B978012382219200288X
Protein Identification and Structure Determination
Identifying the protein components of specific brain regions and individual nerve cells provides a basis for understanding numerous neurobiological phenomena. MS methodologies have the capability of identifying proteins by directly deducing their primary structure. The most common approach, known as peptide mass fingerprinting or a bottom-up measurement, involves digestion of a purified protein of interest or a simple protein mixture into smaller peptides using enzymes such as trypsin. The masses of the resulting peptide ions are measured, compared to predicted masses of proteolytic peptides from a database (which often are unique for a given protein), and mapped to protein sequence stretches that match experimentally observed peptide masses (Figure 2). Peptide mass fingerprinting alleviates the need for characterizing high-molecular-weight proteins and makes protein analysis amenable to several types of mass spectrometers. However, if a crude extract containing numerous proteins is digested, it places severe demands on the separation and analysis system; a sample that initially contained hundreds of protein compounds contains tens of thousands of proteolytic peptides after digestion.
Figure 2. Peptide mass fingerprint of myoglobin by matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectroscopy. Mass spectrum shows multiple peaks that represent peptides resulting from proteolytic digestion of purified horse myoglobin standard. Peptide peaks are labeled by position of the peptide amino acids in the protein sequence. Identities of measured peptides are often assigned using bioinformatics approaches such as MASCOT. This program compares the detected masses in the spectra to predicted masses of proteolytic peptides from the protein database and determines which proteins best match statistically with the experimental data. Sequence stretches underlined in blue denote the peptide identified in the spectrum, and red denotes the sites of trypsin cleavage. In this example, the resulting tryptic peptides cover 94% of the protein sequence. m/z, mass-to-charge ratio.
Existing knowledge on gene expression in the biological tissue or organism under analysis allows identification of proteins by peptide fingerprinting when combined with database searches. Numerous search algorithms are available that rank the recovered protein sequences according to their probability of producing a peptide map that matches the experimentally observed peptides. As the number of organisms with known genomes increases, peptide fingerprinting is proving to be an excellent verification tool in protein expression studies.
Instrumentation platforms that couple two or more mass analyzers in succession, known as tandem mass spectrometry (MS/MS), enable identification of a protein by fragmenting the analyte to allow the determination of its partial or complete amino acid sequence. For example, when analyzing a protein digest by MS/MS, one of the proteolytic peptides can be chosen after the first MS measurement and its molecular ions selectively fragmented by a variety of physical approaches for molecular dissociation. Ideally, fragmentation produces a ladder series of fragment ions (truncated forms of the initial peptide ions) that are analyzed by a second mass analyzer. Mass differences between fragment ions provide information on the arrangement of the amino acids within the fragmented peptide. The deduced peptide sequence, in turn, can be statistically matched to a specific part of a protein in the database. The ability to sequence peptides by fragmentation minimizes the number of proteolytic peptides needed for confident protein assignment, which extends the use of the mass fingerprinting approach to identification of proteins in more complex mixtures, as well as to proteins that produce fewer proteolytic peptides. Sequencing of peptides from unknown proteins (de novo sequencing) may lead to the characterization of new gene products and annotation of new genes.
With the advent of high-resolution mass spectrometers, a second, newer approach to protein identification has been developed, referred to as top-down MS. Unlike peptide mass fingerprinting, top-down analysis begins with direct MS detection and sequencing of an intact protein and lays a foundation for full structural protein characterization. The sequencing of intact proteins is achieved by producing various truncated forms of an entire protein ion by fragmentation inside the mass spectrometer (Figure 3). Thus far, FT-ICR MS appears well suited for top-down protein analysis because it has the required high resolving power of mass measurement and sensitivity. In addition, highly efficient fragmentation techniques that yield complementary information on whole protein fragment ions have been developed specifically for FT-ICR MS. For protein assignment, the data obtained by top-down MS/MS can be processed by search engines in a manner similar to that of the bottom-up approach already described. Although highly accurate and effective in protein identification, the top-down approach is challenging when routinely handling proteins >50 kDa, crude samples with >200 proteins, or complex posttranslational modifications such as glycosylation.
Figure 3. Schematic of the top-down approach for protein sequencing and identification. First, the mass of the whole protein ion is determined. Protein ions are then fragmented within the mass spectrometer and the accurate mass of each of the resulting fragments is measured. Patterns of mass shifts between fragment ions are indicative of the sequence in which amino acids are arranged in the original molecular ion. Analysis of fragment ion patterns allows reconstruction of the protein sequence from both or either terminus. Known variations of mass shifts in fragment ion patterns elucidate the presence of posttranslational modifications. Identification of the deduced protein sequences requires database searches. m/z, mass-to-charge ratio.
URL: https://www.sciencedirect.com/science/article/pii/B9780080450469008706
Introduction
A peptide is a molecule formed by the condensation of a small number of amino acids, of general formula
They are often obtained from the breakdown of proteins.
The analysis of peptide structure is central to the enormous advances made in the last 20 years in all aspects of biomedical sciences. The advances in peptide and protein chemistry have been accompanied by the development of a range of analytical procedures that in turn have accelerated the progress made in the understanding of, for example, hormone–receptor interactions and antigen–antibody interactions.
The vast majority of peptides currently produced are generated by solid-phase peptide synthesis (synthetic peptides) or from the enzymatic or chemical digestion of proteins. The procedures that are commonly used in the analysis of peptides can be divided into four stages: (1) purification, (2) composition and sequence analysis, (3) conformational analysis, and (4) biological analysis. For synthetic peptides, purification procedures are generally of routine nature and amino acid composition and sequence analysis are used to confirm the identity of the synthetic product. The purified peptide samples are then subjected to conformational analysis and biological evaluation to establish the structure–function relationships in the particular biological system. Peptides derived from the enzymatic or chemical digestion of protein (i.e., peptide mapping) are generally prepared in order to derive partial sequence information of a newly isolated protein or, in the case of a recombinant protein, the combined sequence analysis of all peptide fragments can be used to verify the structure of that protein. A wide variety of experimental techniques are now routinely used in the determination of peptide structure. This article deals only with the first three stages of peptide analysis, as a detailed overview of the range of biological assays commonly used to determine the activity of a peptide is outside its scope.
URL: https://www.sciencedirect.com/science/article/pii/B0123693977004416
3.9 Generating a Proteomic Map
Prior to conducting any HDX-MS experiments, a map comprising the curated list of proteolytic peptide ions must be generated from the protein sample. The peptide map is the filter through which mass spectral data are selected for HDX-MS analyses. The curated list of proteolytic peptide ions includes the verified peptide sequence, ion score, charge (z), and retention time for the UPLC apparatus. Although pepsin and the other acidic proteases described above are nonspecific in that cleavage positions cannot be predicted from a protein sequence, these cleavage trends are highly reproducible when the same conditions (pH, temperature, time, and protein concentration) are used. Under controlled conditions each unique peptide elutes from the analytical chromatography column at a consistent retention time, regardless of its deuterium content.
During generation of the peptide map, the digestion and running parameters are optimized, including the temperature of the enzymatic column and the flow rate, which in combination with the column dimensions, determines the digestion time. Proteases function most effectively at physiological temperatures; however, this efficiency must be balanced against the need to reduce back exchange of the labeled backbone amide N–D within the protein, which is decreased at reduced temperature. Additionally, the faster the flow rate, the less time the protein will be in the presence of water, minimizing back exchange; yet, decreasing the flow rate increases the time the protein is in contact with the protease, potentially increasing digestion efficiency. Interestingly, pepsin more effectively digests model proteins when the protease column is maintained at high pressures (Ahn, Jung, Wyndham, Yu, & Engen, 2012; Jones, Zhang, Vidavsky, & Gross, 2010; López-Ferrer et al., 2011). However, the construction of the column housing and stationary phase support may limit the pressure and flow rate that can be applied to the column.
Digestion efficiency is determined by observing the UPLC effluent using a tandem MS (MS/MS) instrument that employs CID to fragment the parent peptide ions. CID is an acceptable fragmentation method since these studies are performed prior to exchange, and H/D scrambling is not a factor. These experiments allow digestion conditions (temperature and flow rate), solution conditions (chaotropic and reducing agent concentrations), and the UPLC gradient to be optimized. By searching a database, such as MASCOT, the m/z values of the detected peptides and their fragments can be correlated to specific peptides within the protein sequence. Good digestions yield peptides covering 70–100% of the protein sequence and a large number of overlapping peptides. The elution time and exact mass for each peptide from the database search are then utilized by the HDX-MS software to find labeled peptides and measure the mass increase associated with deuterium uptake.
As described in the chromatography section, sample carry over complicates HDX-MS data analyses. One way to test for carry over is to repeat the tandem MS analysis immediately following peptide mapping using only water as the “sample.” Since no protein is present, peptides that are detected in a database search must have eluted from the digestion, RP-trap, or RP-analytical column. To minimize carry over, the peptide mapping can be repeated with a lower protein concentration (protein concentration affects acidic protease digestion efficiency) or additional column washes can be incorporated into the HDX-MS experiment between protein analyses (Majumdar et al., 2012).
URL: https://www.sciencedirect.com/science/article/pii/S0076687915004620
Technique 1 – Peptide Mass Fingerprinting
Protein identification by in-gel digestion and peptide mass fingerprinting is almost a routine method in MS laboratories. The protein sample is digested with a specific enzyme, e.g., trypsin, which cuts in a sequence-specific manner to produce a defined set of peptides. Trypsin is useful for mass spectrometric studies because each proteolytic fragment contains a basic arginine or lysine amino acid residue, and thus is eminently suitable for positive ionization mass spectrometric analysis. The peptides are extracted from the gel, sometimes desalted and then the digest mixture can be directly deposited onto the MALDI target plate and the rest of the sample can be stored for subsequent analysis, e.g., by ESI-MS. The peptide mixture is then analyzed by MS to produce a rather complex spectrum from which the molecular weights of all of the proteolytic fragments can be read. This spectrum, with its molecular weight information, is called a peptide map or peptide mass fingerprint (PMF). This PMF can be used to search protein databases. If the protein already exists on a database, then the peptide map is often sufficient to confirm the protein (Figure 1).
Figure 1. The mass accuracy of the peptides measured is critical in PMF; therefore, the spectra have to be internally calibrated. If trypsin is used for in-gel digestion, the enzyme will also cleave itself into fragments, which can be used for calibration. (Reprinted with permission from Nyman TA (2001) The role of mass spectrometry in proteome studies. Biomolecular Engineering 18(5): 221–227; © Elsevier.)
A search algorithm is used that carries out virtual digests of protein sequences based on the sequence-specificity of trypsin and then calculates the masses of the predicted peptides from first principles (for example, by adding up the masses of the individual atoms). The hits are generally ranked according to the number of peptides that match. Unique identification does not require the whole protein to be covered by the tryptic peptides. Usually 10–20% coverage is sufficient. The great precision of mass spectrometers nowadays allows the discrimination of even extremely similar proteins, differing in structure by only one amino acid. The mass accuracy is an important aspect of this approach and with a mass accuracy of 10–30 ppm for peptides in a MALDI-TOF system, usually four or five peptides are enough to identify the protein unambiguously. MALDI produces singly charged ions and hence the spectra are easy to interpret.
There are now available automated systems capable of cutting the spots out of the gel, digesting the samples, desalting, and then spotting onto the MALDI plates. Data processing and database searching can also be automated but in practice, when small amounts of protein digests are analyzed, the spectra have to be checked manually before database searching. Instead of excising individual spots from the gel, some researchers proteolytically digest all proteins on the gel at the same time, transfer to a membrane, and directly scan by MALDI-TOF. This molecular scanning approach takes an enormous amount of time and memory, e.g., for a 16×16 cm2 membrane, 36 days of continuous scanning and more than 40 Gb of raw data.
PMF is only successful when the digested protein exists in the protein or genomic databases, if the fingerprint is unique, and when four or more peptides are obtained by MALDI. If PMF is not successful, it will be necessary to obtain sequence information from the peptides. The most common way to do this is by ESI-MS/MS that is described below.
URL: https://www.sciencedirect.com/science/article/pii/B0123693977003654