Wednesday, March 27, 2013
The ATP-binding proteome of Tuberculosis
Currently in press at MCP is a really cool project that makes amazing use of the Thermo(Pierce) kinase enrichment kits. This paper from Lisa Wolfe et al., represents the work of a group from several institutions and uses the kit to find the proteins that use ATP during the Mycobacterium tuberculosis life cycle. If you're unfamiliar with these kits, you should really check them out. What you get is this pure ATP or ADP that are tagged with biotin. You place these compounds into your experimental system and if, say, your drug treatment (as I've used it) , causes a whole lot of kinase action then the respective tag that you used will be integrated into the binding site of the protein. You then have a permanently tagged protein that you can pull down and implicate in the function you are studying. They can be a little tricky to optimize. This is a high level experiment, but to have your kinases enriched and to know that they are active can give you more information than just about any other experiment. The application of these kits is pretty much up to your imagination and your biological system.
I stole the picture above from the PierceNet website, but you can find more information here. If you are interested in tuberculosis or in a great way to apply this technology, you should definitely check out this paper!
Tuesday, March 26, 2013
ProteinLasso -- use super statistics to estimate protein interference
There is no art whatsoever on the ProteinLasso website, so I made my own icon with the help of Google Images.
Anyway, in every complex MS/MS experiment we ultimately select for fragmentation some ions that we didn't mean to. Increasing sample complexity, increasing the isolation window and decreasing the chromatographic separation all exacerbate this fact. There have been a number of different approaches to estimating or dealing with this. ProteinLasso is a new approach described in this recent paper from Ting Huang et. al., out of the Dalian University of Technology. The approach here, as far as I can tell is some high level statistics called lasso regression. My expert eye can pull a lot of very large capital letter Sigmas in both the figures and the text. The end point, however, is pretty clear -- false discovery rates calculations that appear to work with the same degree of efficiency whether the sample is simple or incredibly complex.
We'll spend some more time on FDR in the near future, but you should check out this paper. The software is also available through sourceforge if you just want to plunge right in.
Anyway, in every complex MS/MS experiment we ultimately select for fragmentation some ions that we didn't mean to. Increasing sample complexity, increasing the isolation window and decreasing the chromatographic separation all exacerbate this fact. There have been a number of different approaches to estimating or dealing with this. ProteinLasso is a new approach described in this recent paper from Ting Huang et. al., out of the Dalian University of Technology. The approach here, as far as I can tell is some high level statistics called lasso regression. My expert eye can pull a lot of very large capital letter Sigmas in both the figures and the text. The end point, however, is pretty clear -- false discovery rates calculations that appear to work with the same degree of efficiency whether the sample is simple or incredibly complex.
We'll spend some more time on FDR in the near future, but you should check out this paper. The software is also available through sourceforge if you just want to plunge right in.
Is Paris Hilton interning with Steve Gygi?
In a bit of silliness, and something that approximates news (or at least what we seem to consider it in the U.S....) I was looking at the current software offerings from Steve Gygi's lab and was surprised to recognize a face in the lab photo. Paris Hilton appears to be working in the Gygi lab, or at least stopped by for a visit. In casual conversation, you often wonder where these starlets go -- mostly to rehab, but not Paris! She's in one of the premier U.S. proteomics labs.
Monday, March 25, 2013
Excessive carry-over in your LC-MS system?
This isn't news, just another random topic for discussion (monologue). Here is the question, though: how do you know if you are having excessive carryover in your LC-MS analysis? There are lots of ways to check this, but this is what I do: I run lots of blanks and I process almost all of them.
Expanding: in between samples of importance, or in-between quantitative runs without internal controls (label-free or SRM or whatever) I inject a normal size sample load (2-10uL, depending on the LC in question). I then process that sample using an appropriate database through my normal processing scheme. Since I almost exclusively work with human samples these days, I simply use my Sapiens Uniprot FASTA (or IPI Human) with the cRAP database either tacked onto the end or searched in parallel. This gives me a pretty good metric of how clean my sample is. I don't freak out when I see some peptides. I only freak out when I see a lot of peptides. These instruments are so sensitive that they are going to find peptides floating in the air or reasonably fresh buffer. They shouldn't, however, find dozens of peptides.
If I do find enough peptides to freak out, these are my steps:
Clean the front of the mass spec. Capillary and front plate with 50% methanol should do it. If that doesn't help, it is time to approach the LC system. The order varies from person to person, but this is how I approach it:
1) Change my blank
2) Change my wash solvents
3) Change my running buffers
4) Run high solvent for a long period of time (a few hours to overnight, if I could possibly afford it. Extremely rare that I could or can)
5) If none of these help, I try injecting something harsh (a full sample loop of 20% isopropanol): Note: This may not be appropriate for all nanoLC systems. I don't think I've ever used bold print in all the years I've been writing in this silly blog. But I don't want you damaging your LC system. I don't know a lot about LCs, just enough to successfully get by and to know that I've never noticeably damaged the LC systems that I have had. That doesn't mean that you should trust me on this one. When in doubt, consult the manufacturer.
6) If this doesn't help, it is time to approach the guard column (if employed) and finally the analytical column.
End of tirade.
Saturday, March 23, 2013
Phosphorylation ends up shifting all over the place?
Ummm....
If this was coming out of other labs, maybe I would glaze right over and by this. But this is coming out of Karl Mechtler's lab. Backing up.... A paper in this month's Proteomics (Wiley) from Andreas Schmidt et al., examines arginine phosphorylation using a global approach. Arginine phosphorylation in bacteria is something that just exploded last year, with a couple really nice papers on where/how it works in Bacillus subtilis. In this paper, this group looks very closely at this and other phosphorylation sites and shows what appears to be phospho groups jumping all over the place during LC-MS/MS analysis. I would have glazed over this because I recall this being a topic of conversation about 6 years ago, phosphorylations moving around during fragmentation, but there was some solid contradictory evidence and we all moved on.
And here it is again. Just when we think we've got a system figured out. If you are doing any kind of phosphorylation analysis this is worth a read. If you are one of the many groups tracking arginine phosphorylation in bacteria (or mammals!?!?) definitely jump on this.
Beta testers needed! Proteome Discoverer tutorial vidoes
I've been kind of quiet this week. Crazy busy working on my talk for Korea HUPO and this project. A survey following this year's North American iORBI tour suggested that people would be really interested in tutorial videos on our software. This week I've been trying to generate a ton of them. Several are now finished and I'd appreciate feedback. If you are interested in being one of my test subjects on the Proteome Discoverer 1.4 tutorial videos just send me an email at: orsburn@vt.edu and I'll send you a link to the videos. Feedback (not linked to my accent or grammar!) would be greatly appreciated.
Wednesday, March 20, 2013
ABRF Results 2013
Wow, O.K., this was just brought to my attention in a comment in Sunday's post. This is a big big study of search engines and confident IDs. It is amazing how much can slip your attention in this field. So many cool things go on that it is impossible to keep of them all. This is really neat. You can find more details here. Thanks, Eric, for bringing this to my attention. I hope to speak more with you in the future.
Tuesday, March 19, 2013
New article in press -- Identification of Protein Interactions Involved in Cellular Signalling
This new article is currently in press at MCP and comes to us from Westermarck et al., as a collaboration between groups in Turku, Finland at the Institutions whose emblems are shown above. This review takes a very critical look at the technologies currently available for interrogating cellular signaling networks. It goes after the classic techniques of the geneticist, such as the yeast two hybrid assay and systems such as strep- and flag- tags. They also take a look at affinity purification coupled MS, and tandem affinity purifications. You get a nice look at the strengths and weaknesses of each assay for studying cellular networks.
Monday, March 18, 2013
Proteome Discoverer 1.4 vs MaxQuant 1.3.0.5
A couple of years ago, I wrote a short and blurb on my experience comparing MaxQuant vs Proteome Discoverer. Turns out, it may have been the most read thing I've ever written. If I'd known how many people would read it, maybe I would have done a more thorough job!
Here is my vindication, though! To celebrate last week's release of Proteome Discoverer 1.4, I took a very nice SILAC labeled data set and ran it through PD 1.4 and MaxQuant 1.3.0.5 (the newest iteration, as of this posting date).
Dataset: Human cancer cell line passaged in SILAC media with Lysine (6) and Arginine (10). The data was analyzed in one go on a Q Exactive system on a 180 minute gradient using a Top20 approach. ~600 MB file.
Software settings:
Dynamic modifications: Carbamidomethylation (C), Oxidation (M), N-acetylation, and SILAC labels
MS1 tolerance: 10 ppm (20 ppm first search for MxQ, 10 ppm for second)
MS2 tolerance: 50 ppm
FASTA: IPI Human 3.77 (originally downloaded from maxquant.org)
MaxQuant used Perseus 1.3.0.4 with an FDR of 0.01 and implemented 4 threads
PD used Sequest with the Percolator algorithm at default parameters
PC: AMD Quad Core, clocked at ~3 GHz with 8 GB of RAM
Total search time:
MaxQuant: 109 minutes
PD: 23 minutes
Results:
MaxQuant: 425 total grouped IDs, 27 of which were contaminants and 12 were reverse sequences.
386 human protein IDs
286 quantifiable
Proteome Discoverer:
465 grouped proteins
380 quantifiable.
I'll be honest. I was scared at first. MaxQuant has gone through some significant revision since I was last using it commonly. Some of the new features, such as the ability to go back and re-search spectra are crazy impressive. That team contains some of the best researchers in our field and continues to innovate how we do proteomics and process MS/MS data. However, I have met a lot of the team that writes PD and they are no slouches either.
I'm going to throw in a caveat here: I am an expert at using Proteome Discoverer. I've been using PD since version 1.0 and have been using the beta versions of PD 1.4 for about 6 months. I'm less adept with MaxQuant. For a quad core cpu, I don't know how many threads would be optimal. 4 seems the smartest, but I may have been able to optimize that number and sped it up (virtual threading, or whatever...) It may also be possible to optimize first search/second search parameters to gain more IDs. Would I have picked up almost 100 quantifiable IDs? I doubt it, but maybe the disparity wouldn't have been as large.
In time, I might do a follow-up article to this one. It would be nice to see what the overlap in ID and/or quan is like. It is a little difficult due to how differently MaxQuant and PD deal with protein grouping. My guess, however, is that the majority of IDs and quan are the same, but that Andromeda and Sequest would each add complementary data to each other.
But for now, for just pure depth of coverage and quan, Proteome Discoverer appears to be the winner, though I'd still encourage you to try running both. The worst that would happen is that you'd get more data from that MS/MS experiment.
Here is my vindication, though! To celebrate last week's release of Proteome Discoverer 1.4, I took a very nice SILAC labeled data set and ran it through PD 1.4 and MaxQuant 1.3.0.5 (the newest iteration, as of this posting date).
Dataset: Human cancer cell line passaged in SILAC media with Lysine (6) and Arginine (10). The data was analyzed in one go on a Q Exactive system on a 180 minute gradient using a Top20 approach. ~600 MB file.
Software settings:
Dynamic modifications: Carbamidomethylation (C), Oxidation (M), N-acetylation, and SILAC labels
MS1 tolerance: 10 ppm (20 ppm first search for MxQ, 10 ppm for second)
MS2 tolerance: 50 ppm
FASTA: IPI Human 3.77 (originally downloaded from maxquant.org)
MaxQuant used Perseus 1.3.0.4 with an FDR of 0.01 and implemented 4 threads
PD used Sequest with the Percolator algorithm at default parameters
PC: AMD Quad Core, clocked at ~3 GHz with 8 GB of RAM
Total search time:
MaxQuant: 109 minutes
PD: 23 minutes
Results:
MaxQuant: 425 total grouped IDs, 27 of which were contaminants and 12 were reverse sequences.
386 human protein IDs
286 quantifiable
Proteome Discoverer:
465 grouped proteins
380 quantifiable.
I'll be honest. I was scared at first. MaxQuant has gone through some significant revision since I was last using it commonly. Some of the new features, such as the ability to go back and re-search spectra are crazy impressive. That team contains some of the best researchers in our field and continues to innovate how we do proteomics and process MS/MS data. However, I have met a lot of the team that writes PD and they are no slouches either.
I'm going to throw in a caveat here: I am an expert at using Proteome Discoverer. I've been using PD since version 1.0 and have been using the beta versions of PD 1.4 for about 6 months. I'm less adept with MaxQuant. For a quad core cpu, I don't know how many threads would be optimal. 4 seems the smartest, but I may have been able to optimize that number and sped it up (virtual threading, or whatever...) It may also be possible to optimize first search/second search parameters to gain more IDs. Would I have picked up almost 100 quantifiable IDs? I doubt it, but maybe the disparity wouldn't have been as large.
In time, I might do a follow-up article to this one. It would be nice to see what the overlap in ID and/or quan is like. It is a little difficult due to how differently MaxQuant and PD deal with protein grouping. My guess, however, is that the majority of IDs and quan are the same, but that Andromeda and Sequest would each add complementary data to each other.
But for now, for just pure depth of coverage and quan, Proteome Discoverer appears to be the winner, though I'd still encourage you to try running both. The worst that would happen is that you'd get more data from that MS/MS experiment.
Saturday, March 16, 2013
How to set up a binning experiment on an Orbitrap
Recently, we've started doing quantitative experiments by looking in smaller MS1 ranges. But what if you could do discovery based proteomics in the smaller ranges? Would this help your results? Absolutely.
This idea has been touched on before by other groups, but this new paper from CE Vincent et. al., out of the University of Wisconsin really knocks it out of the park. The paper focuses on two concepts, one a little older and one brand new for doing this. They call the first one a binning experiment, and their new approach a tiling experiment. For now, I'm going to skip the tiling experiment which is the real star of this paper. This requires subtle alterations of the code controlling Xcalibur on the Orbitrap computer. The lesser note is the binning experiment, which everyone can do on every Orbitrap or LTQ system.
Above is a quick schematic I made of a binning workflow. In this you would run the same sample multiple times, but during each run you can only perform your search for your TopN to fragment from within a narrow mass range. It increases your dynamic range within that area and lets you dig a whole lot deeper. The more narrow the bin, the deeper you'll dig into your sample.
In a tiling experiment, this is taken one step further, where multiple narrow ranges are selected in one run. For example, if this is a Top20 experiment, the first ten MS/MS events can only occur from mass range 400-500 and the next ten can only occur on precursors from 500-600. Without hacking Xcalibur using the developer's kit this option simply isn't available. I would have to think though that the binning experiment would have to do a better job in obtaining sample depth than the tiling experiment, since you are concentrating more time on a more narrow range. The benefit in the tiling approach is that, if utilized well, you would get depth without re-running the sample multiple times, something most of us can't afford.
The real gold in this experiment is that this group used these approaches quantitatively. In each run, even when you were binning a different mass range, you still have the MS1 scans that you could use to for XICs for quan. So you were able to dig much deeper and still have replicates for accurate quantification.
I encourage you to take a look at this paper. And if you have questions on how to set up a binning experiment, I have a method that I successfully utilized on an Orbitrap Velos that I can make available. As always, feel free to email me at: orsburn@vt.edu
Subscribe to:
Posts (Atom)











