Sunday, October 27, 2013

Sirtuin (Sirt6/Sirt7) proteomics!


Sirtuins are genes/proteins of significant interest these days?  Why you ask?  Because they seem to be key regulators of the aging functions in Eukaryotes and a lot of us self-aware and self-centered beings out there would rather not die, nor suffer age-linked fun restrictions!
Knock out SIRTs and you totally mess up the yeast life cycle.
The same holds for mice (as shown above and stolen from the Nature paper clicking on the picture will take you to) But no one really seems to know why yet.  We have their in vitro functional activities all worked out, but there doesn't appear to be a clear link between what they're doing and the complex breakdown of senescence (programmed aging!)

Sounds like a job for proteomics!  Nuts.  I photoshopped a Q Exactive wearing a superman cape and symbol a while back for a talk I gave in Seoul, but I can't seem to find it right now.  I'll add it later!  Nevermind (I made a new one!)  I need a hobby.  This study didn't even use a QE....



Two really nice papers in press at MCP now take a swing at the functions of SIRT7 and SIRT6.
In the SIRT6 paper, "A proteomic perspective of SIRT6 phosphorylations and interactions...," by Miteva and Cristea out of Princeton, this team goes after human SIRT6 using an impressive array of molecular techniques.  The proteomics are flushed out by an Orbitrap XL and Velos.

In "SIRT7 plays a role in ribosome biogenesis and protein synthesis," Yuan-Chin Tsai, et.al., out of the same group uses a similar approach to study SIRT7 knockdowns with an Orbitrap Velos.  Again, this lab demonstrates a remarkable mastery of a wide variety of molecular skills, employing top notch microscopy, genetic techniques and proteomics to really make their case and demonstrate more about the pathways of these proteins than we knew ever knew before.

If this is your field, or you just want to see proteomics seemlessly integrated into a comprehensive pathway study, I suggest you download one or both of these nice new papers.  The links above are to the Early release abstracts and won't be there forever.





Saturday, October 26, 2013

How does Percolator affect MSAmanda search results?

This Saturday afternoon analysis is brought to you by the fact that I don't have any hobbies that you can do when it is too cold to rock climb but not cold enough to snowboard.  It is also brought about from by one of my 10 favorite questions asked by other attendees of the PD user's meeting.

Here is a paraphrase of the question:
MSAmanda gives you more peptides than Sequest using a target decoy search on high resolution MS/MS data, but how does it fare when we use Percolator?  I'm actually going to extend that question one step further, if both using MSAmanda and using Percolator give you more peptides than Sequest + Target decoy alone, are these the same peptides?  This analysis will come later.

Dataset:  A 120 minute HeLa digest run on an Orbitrap Elite using a 25 cm EasySpray column and operating in standard high-high mode employing at Top15 methodology.  So, an extremely complex high-high dataset with  nice chromatography.

Processing:
1) Sequest + target decoy
2) Sequest + percolator
3) MSAmanda + target decoy
4) MSAmanda + percolator

Conditions for processing were as similar and simple as possible, Uniprot/Swissprot database parsed on "Sapiens" alone, iodoacetamide as a static mod and M oxidation as a dynamic mod.  FDR of 0.01 as the "strict" cutoff, and those are the only peptides I looked at.  Mass tolerance of 10ppm at the MS1 level and 0.02 Da at the MS/MS

The following data is all at the Unique protein group level:
First Sequest + target decoy vs. Sequest+ percolator

As expected, Percolator ends up giving us more total unique protein groups.  No surprise there.

Question #1 then, does Percolator + MSAmanda do the same thing?


Yup!  Okay, I totally dig question #1.  And I wonder why the heck I added more work to it, because if I hadn't we're looking at an open and shut case.  MSAmanda is definitely Percolator compatible and at the end of the day, we are looking at more protein groups from MSAmanda whether we use target decoy OR percolator than we get from Sequest.

My conscience is saying that I need to check to see if these new peptides are any good (ugh...).  This is my opinion on false discovery rate calculations (feel free to look at my other discussions on this site), they're a shortcut.  Inherently, I do not trust them, and neither should you.  They are a mechanism to help you, but manual verification is ALWAYS a good idea.  Unfortunately, looking at tens of thousands of MS/MS spectra is a poor use of time.

My strategy:  manually look at a sample of the worst scoring peptides at your 1% FDR cutoff and at your 5% FDR (please excuse my shorthand, you know what I mean, or you probably wouldn't be reading this unless you were really odd.)  If you have crappy peptides at your 1% FDR, it isn't strict enough.  If you have great looking peptides all over the place at your 5% FDR, you are too stringent.  Adjust your cutoffs accordingly and re-evaluate.

This is how I do it on a Saturday afternoon on a dataset I'm not getting paid to analyze.
1) Go to the peptide tab for each analysis
2) Arrange the peptides in order of respective peptide score from worst to best
3) Double click on peptides 1,5,10,15 and 20 to reveal the XICs with the overlayed fragment matches.
4) Rapidly score them by this point system:  10 points if you would publish that peptide spectral match as it is, 5 points if you think it is okay, and -5 points if it is some junk.  Yes, I made this up, geez!  But it works.  Remind me and I'll show more evidence at some point.

Here is how they did:
Sequest + Target decoy:  45 points (one mediocre peptide match)
Sequest + Percolator:  5 points, several bad matches
MSAmanda + Target decoy:  35 points (3 mediocre ones.  not bad, but I wouldn't publish alone)
MSAmanda + Percolator:  50 points.  In this small sample set, I would trust every one of these PSMs.  Ummm...not exactly what I was expecting...but I'm cautiously excited about it!

Okay, this is getting out of hand. Now my conscience says:  Is this due to the sample size?  I looked at the next 20 spectra, spaced every 5 and I don't think so.  You'll have to take my word on it.  But at this default cutoff, there is NO doubt in my mind that the peptides scored by MSAmanda in conjunction with Percolator are significantly better than the peptides scored by Sequest + Percolator and are on par with, or are better(!?!?!), than the peptides scored by the much more conservative target decoy search.

Examples:
In this sample set, the WORST PSM scored by MSAmanda + Percolator and passing default cutoffs:


By comparison, the lowest scoring peptide from the Sequest + Percolator search that passed default FDR cutoffs.



2 y ions?  Seriously?  This is why you CAN NOT trust your default FDR cutoffs.  Take this as a shortcut.  In case you were wondering, I gave this peptide a -5!

I'm going to cut this analysis off now.  Enough data processing for a Friday evening.

Again:  Question #1, does MSAmanda work with Percolator?  My answer, based on 4 runs of 1 dataset.  Absolutely.  In fact, Percolator seems to work a whole lot better with MSAmanda than it even works for Sequest.

By the way, I'm not putting down Percolator + Sequest, I would simply tighten the FDR cutoff until I got to consistently good data.  In this example Percolator simply over-shot the mark a little and dug too hard trying to get us as many peptides as possible.  In fact, that peptide may be a good match, but it is one that I certainly would not show someone to convince them that we found their protein of interest.

Disclaimer, because I'm still a little thrown off by this:  This is one dataset.  The results are surprisingly convincing, however, and the logic is beginning to make sense to me.  Percolator is trying to dig into the data to pull out PSMs that we mistakenly threw out as false (oversimplification, but let's roll with it) and it can only do that based on the quality of the data that was originally identified.  If MsAmanda is doing a superior job of making peptide to spectral matches, Percolator has more to work with.

TL/DR:  Use MSAmanda for high resolution MS/MS spectra.  Also use Percolator, they are compatible and give you more data.  Always verify if your FDR cutoffs are giving you good data!


Friday, October 25, 2013

What happened at the PD user's meeting?


Not to rub it in if you couldn't attend, but the International Proteome Discoverer User's meeting kicked ass.  It was easily the most valuable learning experience that I had personally this year.  I'm working on hunting down the talks now and I'll provide links to them as soon as I can obtain them. I want to provide an overview of what happened.

Talk 1)  Bernard Delanghe went over basic PD functions, then plowed head on into the power of using multiple search engines, both in parallel and in series for digging into your data.  I recently touched on the extreme results the BRIMS has had when using multiple engines in parallel.  A big emphasis of Bernard's talk was the enormous value in processing speed and identification rates that occur by using multiple engines in series.  I've touched on that a little, but expect a number of experiments from me to follow.


Talk 2) Marshall Bern described Byonic, a software that I've previously beamed about, in particular how Byonic can be used effectively for glycopeptide analysis.  He also described the Byonic node that will soon be an purchasable upgrade for PD.  Expect an explosion of analyses and announcements and data from me when this launches.

Talk 3) Viktoria Dorfer gave a talk on the power of her creation, MSAmanda, in the scoring of high resolution MS/MS spectra.  An interesting note for proteomics software teams out there should be the fact that the creator of this fantastic new tool is still working on her Ph.D.  A highlight of her talk, for me was a comparison of the number of peptides that MSAmanda found for ETD spectra when compared to the other search engines.  This should be an extremely interesting observation for the Orbitrap Fusion teams out there, as the incredible speed of the Orbitrap (and the ease of obtaining high ETD signal) allows high resolution ETD spectra to be an efficient experimental design.  Expect experiments and data from me to follow.

Talk 4) Not to take anything away from the other great speakers, but this was Oliver Serang's day. In an incredibly amusing and informative talk, Oliver walked us through a new way of thinking about assigning protein identity to identified PSMs.  These new functions are to be available in PD 2.0.  I am currently awaiting library access to the papers that Oliver has written so far on this and other topics.  When these arrive, I'll spend some time trying to figure these out.  In the meantime, I highly recommend a Google Scholar search on his name.

Talk 5)  Automated spectral library generation node for Proteome Discoverer!  My good friend Maryann Vogelsang presented a node in development at BRIMS that will take your PD data and automatically generate spectral libraries, which can then be directly imported into Pinpoint!!!!!  Remember my entry that spectral libraries are about to blow up?  They are, especially with the thousands of terabytes of high resolution/high quality Orbitrap data out there that we can easily convert into high res spectral libraries?  Exciting!!!!

Talk 6) David Perlman gave a talk that I regretfully had to miss due to an important meeting.  Fortunately, I did get to meet him finally later that evening.  In case you aren't familiar, Dr. Perlman is the Director of the Proteomics and mass spectrometry core facility at Princeton University.  You can find out more about his facility here.

Talk 7) Proteome Discoverer 2.0!  Bernard showed it in action and took audience suggestions.  It is currently in alpha testing (and open on this PC I'm typing on, haha!) and it looks great.  Tons of new features are coming, but I'll go into them later.

Q/A session:  This was a great session.  I wrote down a large number of questions and many of them will be the feature of upcoming blog entries.

In sum, it was a great session.  I'm eagerly looking forward to next year's.  Expect much greater press coverage when the date is set for #4 and seriously think about showing up.  It is a valuable experience.

Thursday, October 24, 2013

Check out the new links on the side bar!


New on the side bar!  Direct links to all of the current Proteome Discoverer 1.4 videos, as well as access to the beginnings of the Orbitrap methods database.  The third one is only a placeholder, but it will my attempt to de-mystify some of the excessive terminology and get us all speaking the same language!  These new pages came about from suggestions I received during the amazing Proteome Discoverer International User's meeting.  Highlights from the meeting will follow!

Wednesday, October 23, 2013

Two new papers in press at MCP highlight the value of peptide immunoaffinity enrichment


For phosphoproteomics, this has been a by-gone concusion -- you can't get down to the majority of your PTMs without enriching at the peptide level.
Two papers in press now at MCP show the values of these approaches for other PTMS, namely ubiquitination and methylation.
 "Peptide level immunoaffinity enrichment enhances ubiquitination site identifications on individual proteins" from Anania et al., and "Immunoaffinity enrichment and mass spectrometry analysis of protein methylation" from Guo and Gu, et al., both support this fact in their PTMS of interest.

Perhaps more importantly, both papers provide a very nice and effective method for identifying high numbers of peptides with these modifications in a complex environment.  If you are interested in either of these PTMs, definitely download these papers now!

Monday, October 21, 2013

One hour yeast proteome!


Holy shit!  One hour for a proteome?  Not one hour for a protein or two and calling it "proteomics".  Full theoretical proteome coverage for a Eukaryote in one hour.
This study, in press at MCP, from Josh Coon's lab uses the Orbitrap Fusion to pull out ~4,000 yeast proteins in a 1 hour run. And reproduces it.

This is what happens when a great lab get their hands on the fastest and most sophisticated mass spec ever built.  They change our perspective of what can be done and when!  Get it here.

An extremely thorough new review of plant proteomics -- where are we now?


This review is absolutely a work of love.  Called, "A decade of plant proteomics and mass spectrometry:  Translation of technical advancements to food security and safety issues," this review is co-authored by a group representing no less than 6 different countries.  I haven't worked a lot with plants, but I've helped some people who have, so I appreciate the complexity of working with organisms with such complex and often repetitive genomes.

If this is your field, download this thorough review.  You can find it open access here.


Thank you to SpectroscopyNow for leading me to it.  You can find their overview of the review here.

Sunday, October 20, 2013

Georgia Tech Starter -- Let's crowdsource some science!


Crowdsourcing is huge right now.  We're using it to refine ideas, local speed traps on the highway (Waze!), find the best cat videos (Reddit!), and start new businesses through programs like Kickstarter.

Georgia Tech recently came up with the idea of crowdsourcing funds to get scientific research programs going.  The research that is selected for this program is extremely well filtered through a peer review process before it can be listed on the GTS page.

{Begin Ben tirade}Honestly, I consider this a little excessive.  One of the cool things about programs like Kickstarter is that we get to decide what programs we want to fund.  Having some professors pre-filter it takes a little bit of the fun out of it, in my opinion.  Their intentions are good, the peer review process is intended to filter the research to that which is most likely to succeed.  I think this intention actually allows us to forget how much we learn from scientific "failures".
{End Ben tirade}

This is a really cool program and I hope to see more of these in the future!  For more information, hit the GTS website here.

Saturday, October 19, 2013

New UCSD paper shows novel way to think about proteogenomics!



This paper, currently in press at MCP, may get my vote for bioinformatics paper of 2013.  The study comes from Natalie Castellana et al., out of UCSD.

Let me frame it, first, as I see it.  We have an organism that lacks a fully complete and annotated database.  What we do have, however, is a ton of high quality next-gen sequencing data and a few million MS/MS spectra from shotgun proteomics on this organism.  Can we possibly put the sequencing and MS/MS spectra together without having the complete sequence?

It turns out the answer is yes.  Yes we can.  In this impressive study, the team took next gen sequencing data and MS/MS spectra from corn (Zea mays) and lumped the two together, using a really impressive logical progression.  I don't want to ruin the story for you, but what if we stopped thinking that unique MS/MS spectra were the coolest part of our data?  What if we, instead, took a probability based approach and considered the repeat occurrence of spectra to be an indication of the strength that observation is true?  Obviously, we're doing some de novo type sequencing here, and considering that every peptide spectral match has a degree of uncertainty to it (and de novo even more so!) the fact that we've made that identification more than once can actually be considered a very complex functional measurement of the level of certainty of that measurement.

I'm going to stop here.  I lied.  I do want to ruin the story for you, but I am not doing the story justice.

If you are working with an organism that is not fully sequenced, or you want to but the lack of sequencing is stopping you, definitely check out this paper.  "An Automated Proteogenomic Method Utilizes Mass Spectrometry to Reveal Novel Genes in Zea mays" is available in pre-release version at MCP here.

Blood based proteomics markers demonstrated as effective diagnostics for lung cancer


A fantastic new study in this issue of Science Translational medicine demonstrates remarkable correlation between a particular combination of blood-borne peptide markers and lung cancer.  By evaluating over 300 markers in both patients with benign and cancerous lung legions, the team found 13 markers that were highly predictive of lung cancer.

I first found this article through a mention in GenomeWeb regarding plans to commercialize this and other techniques.
The original article can be found here.