This week I was introduced to a free MS/MS search database system operated out of the University of Ohio called Mass Matrix. This algorithm has evolved over the past several years and has been the subject of at least 7 publications. I will have to explore this further, but this algorithm is particularly interesting to me due to its extreme bias toward HCD fragmentation data that is read in high resolution in an Orbitrap FT scan mode. We discovered this bias when we took two files that were generated on an Orbitrap Elite from the same sample and conditions that were interpreted using three separate database searching algorithms. The files were generated Top25 based methods. The first method was a Top 25 CID method and the second was a Top 25 CID High-High method, where the FT MS/MS scans were accumulated at 12,000 resolution. Both methods generated around the same number of total scans ~75,000.
Both Sequest (PD 1.3) and Mascot (searched through the mysterious-to-me Elucidator platform) generated roughly the same number of peptide and protein matches for the two files (within 3%)
Mass matrix, however, demonstrated an extreme bias toward the HCD generated high resolution fragment ions, with >38% more IDs with the high-high method.
This is just another example of the importance of using multiple search algorithms to interpret results. It would be interesting to see why this bias exists in this software, because I feel that you really should be seeing more positive IDs when employing higher resolution MS/MS scans.
Friday, July 20, 2012
Sunday, July 15, 2012
SWATH: The worst idea in proteomics?
I've been reserving my opinion on the new MS/MS technique called SWATH for several weeks now. I like to think about things for a long time before I commit to them, sometimes. And I felt like I really needed to learn more about this procedure, particularly when my initial reaction was so overwhelmingly negative. After weeks of thinking about it, I'm convinced that I don't think SWATH is a good idea at all.
The SWATH technique is a data independent fragmentation method. Every ion is fragmented. The good, the bad, the intense, the weak -- every ion. In order to narrow this down, small mass ranges (or swaths) are chosen, often in 25 amu windows. Every ion coming through with a m/z of 350-375 is then fragmented and MS/MS spectra collected. The machine then goes to fragmenting all ions from 375-400. To be clear, the MS1 mass is never recorded. You simply do not know the mass of your parent ion.
In order to get over this minor obstacle of not actually knowing what you fragmented (other than its m/z, plus or minus 25 amu!), the processing program compares the MS/MS spectra generated to a pre-generated spectral library.
Okay, so maybe I'm being old fashioned, and I have been doing this a while. But not that long ago, we looked at the parent mass of every ion we fragmented and made a hypothesis of what it might be. Then we examined each ion generated in MS/MS individually and matched them to see if our hypothesis was correct. I know, mass spectrometry has been growing at an insane rate. It is almost impossible to keep up with the advances in hardware, software, and processing technologies. But at the heart of it, I personally believe that it still breaks down to: can we make a structural hypothesis based on the MS1 and support it with the MS2?
Besides these perceived shortcomings, there are some other concerns. Moving sequentially through mass ranges takes a long time. A method that requires such intense scan times can not possibly be compatible with rapid chromatography systems. The real limitation of most proteomics labs is the amount of time it takes to get data from a sample. This is most often limited by the amount of time required to get a good chromatographic separation. Ideally, these would be getting shorter all the time. I know I was thrilled when I found a superior column manufacturer that could trim 15 minutes off of my gradient time -- and here we are looking at a slower method?
The SWATH technique is a data independent fragmentation method. Every ion is fragmented. The good, the bad, the intense, the weak -- every ion. In order to narrow this down, small mass ranges (or swaths) are chosen, often in 25 amu windows. Every ion coming through with a m/z of 350-375 is then fragmented and MS/MS spectra collected. The machine then goes to fragmenting all ions from 375-400. To be clear, the MS1 mass is never recorded. You simply do not know the mass of your parent ion.
In order to get over this minor obstacle of not actually knowing what you fragmented (other than its m/z, plus or minus 25 amu!), the processing program compares the MS/MS spectra generated to a pre-generated spectral library.
Okay, so maybe I'm being old fashioned, and I have been doing this a while. But not that long ago, we looked at the parent mass of every ion we fragmented and made a hypothesis of what it might be. Then we examined each ion generated in MS/MS individually and matched them to see if our hypothesis was correct. I know, mass spectrometry has been growing at an insane rate. It is almost impossible to keep up with the advances in hardware, software, and processing technologies. But at the heart of it, I personally believe that it still breaks down to: can we make a structural hypothesis based on the MS1 and support it with the MS2?
Besides these perceived shortcomings, there are some other concerns. Moving sequentially through mass ranges takes a long time. A method that requires such intense scan times can not possibly be compatible with rapid chromatography systems. The real limitation of most proteomics labs is the amount of time it takes to get data from a sample. This is most often limited by the amount of time required to get a good chromatographic separation. Ideally, these would be getting shorter all the time. I know I was thrilled when I found a superior column manufacturer that could trim 15 minutes off of my gradient time -- and here we are looking at a slower method?
Thursday, July 5, 2012
Cool old video from the Royal Chemical Society
This is an older video produced by the Royal Chemical Society that explains mass spectrometry using a classic magnetic sector instrument. It is amazing how far we've come in the last decade or so!
3000 views in 2012
3000 views!!!
This is probably beeping my own horn, if you want, but this site has now been viewed >3,000 times since I moved it from its old domain in January of this year! Thank you so much for stopping by. Keep the questions coming, and I'll try to keep on writing!Monday, July 2, 2012
Offgel separation of depleted plasma, part 3
This is a continuation of experiments that I began to describe in April of this year
Please see Part 1 and Part 2, for more information.
Since I finished part 2 of this short experiment, 3 separate readers have written me with this exact same question: how many unique peptides were discovered in each experiment?
As (I hope) you can see in the grainy .JPEG from Excel above, the peptide OFFGEL separation provided the largest number of both unique protein groups and peptides of the 4 methods evaluated. Surprisingly, the SCX came in second, suggesting that (at least for this very simply study) methods that separated digested plasma proteins were superior to separation of plasma proteins at the un-digested level.
Disclaimer: This study was not for the purpose of discovering large numbers of proteins or peptides. This method was purely for the evaluation of 4 different separation techniques for the identification of depleted plasma proteins. All 4 methods began with the same amount of protein, either pre- or post- digestion, with the assumption that the loss during digestion whether in-gel or in-solution would be roughly equivalent. The method was a standard Top10 high/low with a dynamic exclusion at 2 events and a mass width of 0.01 Da.
As always, please email me directly if you would like further information (and sorry for the delay!). Keep those questions coming!
Please see Part 1 and Part 2, for more information.
Since I finished part 2 of this short experiment, 3 separate readers have written me with this exact same question: how many unique peptides were discovered in each experiment?
As (I hope) you can see in the grainy .JPEG from Excel above, the peptide OFFGEL separation provided the largest number of both unique protein groups and peptides of the 4 methods evaluated. Surprisingly, the SCX came in second, suggesting that (at least for this very simply study) methods that separated digested plasma proteins were superior to separation of plasma proteins at the un-digested level.
Disclaimer: This study was not for the purpose of discovering large numbers of proteins or peptides. This method was purely for the evaluation of 4 different separation techniques for the identification of depleted plasma proteins. All 4 methods began with the same amount of protein, either pre- or post- digestion, with the assumption that the loss during digestion whether in-gel or in-solution would be roughly equivalent. The method was a standard Top10 high/low with a dynamic exclusion at 2 events and a mass width of 0.01 Da.
As always, please email me directly if you would like further information (and sorry for the delay!). Keep those questions coming!
Monday, June 25, 2012
Free Graphpad calculations
This webpage came up in a conversation over the weekend. Its one of those things that has been around for so long I thought everyone used it. This site by GraphPad will do statistical calcuations (unpaired student's t test, and Welch's t test) online in about 5 seconds. Cut the columns out of an Excel spreadsheet and paste it right into one of their two columns. Forget what p value is significant or extremely significant? Who cares, GraphPad will tell you. The only drawback is that you are limited to 2 sets for comparison. You have to buy the real version for more.
Saturday, June 23, 2012
Pinpoint
Okay, so this is probably a weird entry. Weird because I don't want to lose credibility with my small audience because I'm writing about a product produced by my new employer. Regardless, I really wanted to write this entry because Pinpoint is a great piece of software. And if you're like me, you've heard of it, but don't really know what it can do.
This entry is also probably a bit premature, because it looks like I've only scratched the surface of what Pinpoint is capable of. But after several days of using it, I'm beginning to become very comfortable with the basic functions of the software.
First of all, Pinpoint is for targeted proteomic studies. Specifically, if you have a protein (or, more importantly, protienS) of interest and want to perform targeted qual/quan analysis on different samples for the specific peptides, Pinpoint is your software.
When I participated in targeted studies at the NIH, we always felt that we were limited to either peptides that we had identified in discovery runs. When we specifically went looking for a protein or peptide that had not shown up in a previous experiment we went through the following steps:
1) Looking up the protein sequence through NCBI (and trying to guess the right one, because the NCBI is almost at the point where it has too much information to sort through)
2) Taking the sequence (that is hopefully correct!) to the UCSF protein prospector and performing an in silico digest of the protein sequence. Since you can't save your settings, every modification and setting has to be re-inputted every time you reopen the webpage. (Please don't think I'm knocking the Prospector project, btw, I will never fully express my gratitude to the University of California for setting up and maintaining that site! See the lavish praise I showered on this site in my first book for more information!)
3) From the Prospector output, manually remove all singly charged, redundant, and unlikely peptides from the massive list of possible peptide masses that resulted
4) Move that list into the Always Include box within your Xcalibur method.
5) Run the sample
6) Manually extract the peaks and areas for each peptide that you found using Xcalibur and plot your standard curve using Excel
7) Make a professional looking output graph in Powerpoint. The trick is adjusting your peaks for the inevitable little shifts in retention time (or big shifts, if you are using and Eksigent...)
And this way works. I know that some of my readers are still doing it this way. After you've done it a few times, it doesn't take all day to show that your protein is upregulated after drug treatment, just most of a day.
Pinpoint is so amazing because it does all of it for you. Every bit of it. In about a minute.
This is the order of events
1) You tell Pinpoint what protein(s) you are interested in. It can be from your own FASTA file, or it will look it up for you and it inputs the sequence.
2) You tell Pinpoint what you want to digest it with. You can save these settings, so it knows you use trypsin and you expect no more than 2 missed cleavages, and you use iodoacetamide.
3) It digests your protein and gives you only the good, relevant peptides that you will be able to get information on.
4) It also predicts the RETENTION TIME of said peptides. No joke.
5) You export your file and copy it right into your Method file.
6) You make your runs
7) You drag the completed .raw files into Pinpoint
8) It plots the abundance of all of the peptides from your theoretical digest from each of the samples, lining up the peaks and making all the small adjustments in retention time. Resulting in an absolutely gorgeous output file.
Summary:
If you are doing targeted proteomics studies the way that I mentioned above, you want this software. But, seriously don't take my word for it. Go to the Thermo-BRIMS portal and download the free trial version.
This entry is also probably a bit premature, because it looks like I've only scratched the surface of what Pinpoint is capable of. But after several days of using it, I'm beginning to become very comfortable with the basic functions of the software.
First of all, Pinpoint is for targeted proteomic studies. Specifically, if you have a protein (or, more importantly, protienS) of interest and want to perform targeted qual/quan analysis on different samples for the specific peptides, Pinpoint is your software.
When I participated in targeted studies at the NIH, we always felt that we were limited to either peptides that we had identified in discovery runs. When we specifically went looking for a protein or peptide that had not shown up in a previous experiment we went through the following steps:
1) Looking up the protein sequence through NCBI (and trying to guess the right one, because the NCBI is almost at the point where it has too much information to sort through)
2) Taking the sequence (that is hopefully correct!) to the UCSF protein prospector and performing an in silico digest of the protein sequence. Since you can't save your settings, every modification and setting has to be re-inputted every time you reopen the webpage. (Please don't think I'm knocking the Prospector project, btw, I will never fully express my gratitude to the University of California for setting up and maintaining that site! See the lavish praise I showered on this site in my first book for more information!)
3) From the Prospector output, manually remove all singly charged, redundant, and unlikely peptides from the massive list of possible peptide masses that resulted
4) Move that list into the Always Include box within your Xcalibur method.
5) Run the sample
6) Manually extract the peaks and areas for each peptide that you found using Xcalibur and plot your standard curve using Excel
7) Make a professional looking output graph in Powerpoint. The trick is adjusting your peaks for the inevitable little shifts in retention time (or big shifts, if you are using and Eksigent...)
And this way works. I know that some of my readers are still doing it this way. After you've done it a few times, it doesn't take all day to show that your protein is upregulated after drug treatment, just most of a day.
Pinpoint is so amazing because it does all of it for you. Every bit of it. In about a minute.
This is the order of events
1) You tell Pinpoint what protein(s) you are interested in. It can be from your own FASTA file, or it will look it up for you and it inputs the sequence.
2) You tell Pinpoint what you want to digest it with. You can save these settings, so it knows you use trypsin and you expect no more than 2 missed cleavages, and you use iodoacetamide.
3) It digests your protein and gives you only the good, relevant peptides that you will be able to get information on.
4) It also predicts the RETENTION TIME of said peptides. No joke.
5) You export your file and copy it right into your Method file.
6) You make your runs
7) You drag the completed .raw files into Pinpoint
8) It plots the abundance of all of the peptides from your theoretical digest from each of the samples, lining up the peaks and making all the small adjustments in retention time. Resulting in an absolutely gorgeous output file.
Summary:
If you are doing targeted proteomics studies the way that I mentioned above, you want this software. But, seriously don't take my word for it. Go to the Thermo-BRIMS portal and download the free trial version.
What mass range contains the most (and best) fragment ions? Part 2
This is a continuation of an entry from a couple of weeks ago. The question is this: If I'm using an LTQ Orbitrap system, should I always be scanning in MS1 (and selecting fragment ions) with an m/z from 300 to 2000, or am I wasting valuable scan time? Should I really be scanning from 500-1,000 because there is no useful information in the low or high mass ranges? In the first analysis we looked at a file that was generated from an LTQ Orbitrap Velos using a standard Top10 CID method. What we observed in that experiment was that although some fragment ions were generated in the m/z range >1600, no peptide sequences were actually obtained from these ions.
The follow up question: Is this some artifact of the LTQ? Does the ion trap simply do a better job of fragmenting ions in the mass range I observed?
In answer to this, I obtained two files from an old colleague. In this experiment, their lab was comparing the Top 10 CID method on an Orbitrap Velos (the high/low method) to the HCD based Top 10 method (high/high). The same sample (a fraction of human serum) was ran in both methods. The same amount of sample was injected, the MS1 was set at 60,000 resolution. The HCD Top 10 ions were read in the FT at a resolution of 7,500.
The CID method identified 57 proteins from 138 high confidence peptides
The HCD method identified 44 proteins from 99 peptides.
-Since high/high method on a standard Velos is slower than the high/low method, these results are definitely in the range I would expect.
The average m/z of the identified peptides from the CID method was ~821
The average m/z of the peptides from the HCD was 769, which seems interesting until you note that the standard deviation of these two averages are ~150 amu. I was about to consider them the same, but decided to do a student's unpaired t test. The p value is 0.0162, which is technically a significant value. With error bars overlapping this much, I am willing to say that the HCD may allow the identification of peptides of lower m/z.
Here are the ID'ed peptide distributions by m/z
First of all, I want to caution that this is again, one experiment, from one run of one particular sample ran in one software (PD 1.3) versus one search engine (Sequest). But from this extremely limited dataset, it looks like using the high/high method versus the high/low method results in a somewhat different distribution of confidently identified ions.
I'm definitely not done with this concept, or even the analysis of this one pair of samples, but that is all the time I'm willing to commit to this today.
The follow up question: Is this some artifact of the LTQ? Does the ion trap simply do a better job of fragmenting ions in the mass range I observed?
In answer to this, I obtained two files from an old colleague. In this experiment, their lab was comparing the Top 10 CID method on an Orbitrap Velos (the high/low method) to the HCD based Top 10 method (high/high). The same sample (a fraction of human serum) was ran in both methods. The same amount of sample was injected, the MS1 was set at 60,000 resolution. The HCD Top 10 ions were read in the FT at a resolution of 7,500.
The CID method identified 57 proteins from 138 high confidence peptides
The HCD method identified 44 proteins from 99 peptides.
-Since high/high method on a standard Velos is slower than the high/low method, these results are definitely in the range I would expect.
The average m/z of the identified peptides from the CID method was ~821
The average m/z of the peptides from the HCD was 769, which seems interesting until you note that the standard deviation of these two averages are ~150 amu. I was about to consider them the same, but decided to do a student's unpaired t test. The p value is 0.0162, which is technically a significant value. With error bars overlapping this much, I am willing to say that the HCD may allow the identification of peptides of lower m/z.
Here are the ID'ed peptide distributions by m/z
First of all, I want to caution that this is again, one experiment, from one run of one particular sample ran in one software (PD 1.3) versus one search engine (Sequest). But from this extremely limited dataset, it looks like using the high/high method versus the high/low method results in a somewhat different distribution of confidently identified ions.
I'm definitely not done with this concept, or even the analysis of this one pair of samples, but that is all the time I'm willing to commit to this today.
Monday, June 18, 2012
Protein Identification using Top-Down
I admit it, I first downloaded this article in June's MCP because I thought that it was going to be a nice review of recent advances in top down proteomics. For anyone who hasn't done this kind of work, this is where you do not digest your proteins before performing LC-MS/MS analysis. When I was doing top down work years ago, I liked to start with 1 single protein and post-deconvolution, I could deliver results + or - 10 daltons.
Recent advances in LC separation of intact proteins as well as high resolution mass spectrometry and new fragmentation methods have allowed several recent studies to report the identification of hundreds of proteins in a single LC-MS/MS run. It isn't uncommon for labs these days to report deconvoluted mass accuracy in the parts-per-million (PPM) range.
The paper we are discussing here is not, however, a review of top-down proteomics. This paper is a description of a new piece of software for performing top-down analysis, called MS-Align+.
In order to evaluate the effectiveness of MS-Align+, the authors separately harvest proteins from yeast and a species of salmonella. The proteins are separated on long LC runs (600 minutes) on a system coupled to an LTQ Orbitrap. The MS1 spectra were obtained at 60,000 or 30,000 resolution, depending on the experiment. The intact proteins were fragmented using the HCD cell and the MS/MS spectra were also obtained in the FT cell at a resolution of 30,000.
The obtained spectra were then evaluated with MS-Align+, Mascot, OMSSA. Even though one of the primary goals of the project was to optimize MS-Align+ for simplicity and speed, the authors report that MS-Align+ compared favorably against every other algorithm they used.
Summary: Despite what the title suggests, this is not a review paper on top-down proteomics. It is a nice paper describing a new algorithm for top-down analysis. The software uses some clever mathematics to run quickly, even on outdated desktop computers. Unfortunately, due to the incredible similarity of the article's title to that of recent reviews on this subject, I fear that information on this algorithm will not disseminate as quickly as it deserves.
Wednesday, June 13, 2012
Systematic Comparison of Fractionation Methods for In-Depth Analysis of Plasma Proteomes
This paper from Cao, et al., came to my attention while digging through references for a paper I am constructing with my soon-to-be ex-lab. I've read this through 4 times, because I am absolutely amazed by how different the protocols described in this paper are from the way that I have learned to do things. These are primarily small differences, but it is the number of these alterations that astound me.
For example:
-Dimethylacrylamide (DMA) was used to alkylate proteins
-Peptides were desalted with ultra-MicroSpin columns from the Nest Group
-The Orbitrap XL performed MS1 at 60,000 resolution, and accumulated MS/MS data from the Top 6 ions
-The data was processed with Bioworks against the UniRef 100 database
-MS1 tolerance was set at 100 ppm and the MS/MS tolerance was 1 Da
-A variable modification was set for the deamidation of asparagine
-MySQL was used to compile the data
Not one of the steps listed above is how I would have done this experiment, but that obviously doesn't matter because their data is beautiful. They show thousands of unique high scoring peptides for each sample preparation method they evaluate.
That brings me back to the actual paper. The goal, as stated in the title, is to show what methods produce the best coverage of the plasma proteome. They compare 1 dimensional gels to OFFGEL fractionation to High Ph-RP-HPLC, as well as the pros and cons of collecting more fractions within each method. Overall, it is an extremely meticulous work. If you are interested in plasma proteomics you need to read this paper.
The real take-away for me, however, is how many changes can be made in MS/MS downstream processing that will still result in good peptide coverage. This also suggests how difficult it is going to be to make all of us transition over to the unified protocols we all know are necessary to really move proteomics into the promise land of cross-laboratory reproducibility.
For example:
-Dimethylacrylamide (DMA) was used to alkylate proteins
-Peptides were desalted with ultra-MicroSpin columns from the Nest Group
-The Orbitrap XL performed MS1 at 60,000 resolution, and accumulated MS/MS data from the Top 6 ions
-The data was processed with Bioworks against the UniRef 100 database
-MS1 tolerance was set at 100 ppm and the MS/MS tolerance was 1 Da
-A variable modification was set for the deamidation of asparagine
-MySQL was used to compile the data
Not one of the steps listed above is how I would have done this experiment, but that obviously doesn't matter because their data is beautiful. They show thousands of unique high scoring peptides for each sample preparation method they evaluate.
That brings me back to the actual paper. The goal, as stated in the title, is to show what methods produce the best coverage of the plasma proteome. They compare 1 dimensional gels to OFFGEL fractionation to High Ph-RP-HPLC, as well as the pros and cons of collecting more fractions within each method. Overall, it is an extremely meticulous work. If you are interested in plasma proteomics you need to read this paper.
The real take-away for me, however, is how many changes can be made in MS/MS downstream processing that will still result in good peptide coverage. This also suggests how difficult it is going to be to make all of us transition over to the unified protocols we all know are necessary to really move proteomics into the promise land of cross-laboratory reproducibility.
Subscribe to:
Posts (Atom)







