Sunday, March 31, 2013

Nature special focus on Big Data. Pretty relevant to us!


What an interesting week in the big journals!  Nature is running this really great focus on Big Data and what the heck we are going to do with all of it.  I think we're all pretty familiar with these problems by now.  I am currently carrying 3 1TB hard drives with me so that I can work on my research on the road and have control data sets to show other researchers and still have space to process any data that they want.

Tranche is worthless to most people now.  I haven't been able to get even 1 single RAW file off of the site in over 6 months.  Supposedly there is data there, but I think it's probably just a room full of smoking servers.  The problem is that we are generating, on average about 1GB/hour with most of the current instrument line.  What do we do with all of this?  How do we store it safely.  Hell, how do we even process it?

(This is what our little server is running at right now....  93% load... )

We're not the only people having problems.  The Next Gen sequencers are having the exact same problem.  A researcher in Maryland that I know told me that his sequencers can generate as much data in a weekend as was produced during the entire 13 years of the Human Genome Project.  Yeah, they feel our pain.

The good news is that since we're all having problems, maybe solutions will be coming more quickly!  If you're at all concerned with where we're going, I definitely suggest that you check this out!


MaxQuant and QE phosphoproteomics data


This is going to sound more like a gripe than what I mean this to be, but has anyone out there tried to process QE AIF-NL-ddMS2 data with MaxQuant?  Holy cow.  My server has been cranking at >90% CPU and/or hard drive capacity for 14 hours now to finish 6 files that PD 1.4 with SequestHT knocked out in <4 hours.  I'll take this to the MaxQuant Google group, but since there is no place in MaxQuant where you can filter out the AIF data from the MS2 data then that is where the hangup is.  I cancelled my SIEVE and PD queues just to allow MaxQuant more power, but it doesn't seem to have helped.

Saturday, March 30, 2013

cBioPortal -- rapid access to numerous cancer genomics sets


Today's issue of Science and Science signaling have a special emphasis on Cancer Genomics.  Most of the articles, unfortunately, won't be available to actually read for a couple of days, but I did manage to find something interesting.  The cBioPortal at MSK is a database that looks through a large number of published genomics data sets and gives you information on whether your gene(s) of interest have been previously implicated in any forms of cancer.
For example, a study that  I did during one of my postdoctoral fellowship implicated a novel phosphorylation site on integrin alpha 4 (ITGA4) in the mechanism of action of a proprietary drug we were working with.  Input ITGA4 into cBioPortal and you get this nifty chart:

This shows that ITGA4 mutations, amplifications, and deletions have all been implicated in these forms of cancer.  This would have been a great slide for our monthly progress reports!  At least we have access to this great database now!  You can find more on it (and a tutorial) at the MSK website.


Friday, March 29, 2013

Culture shock: Proteomics study on CAND1 in Cell makes the news in Seoul!


There is a great proteomics paper in this month's edition of Cell.  I spent some time with one of the authors, Dr. J. Eugene Lee,  and got an amazing walkthrough on this study where they have identified a protein with a completely novel type of activity.  The paper in question, from Nathan Pierce et al., comes out of Cal Tech and describes the ability of the CAND1 protein in accelerating the dissociation of certain protein families (by over 1 million fold!).  The study is a little mind-boggling in its scope and extremely thorough.  One of the most interesting aspects of the study for me is the fact that they used a "pulse SILAC" technique.  After their protein of interest was turned on by their expression system, they changed the media for SILAC media.  This halted the production of their expression vector and left all of their protein of interest unlabeled.  Using SILAC in this manner had never occurred to me, but it makes great sense, and I can think of lots of uses for this technique.

This is a great and extremely innovative study where they demonstrate a protein function that has never before been observed.  I fully expect this system to be a staple of future biochemistry texts.

As much as I respect this study, what I almost respect more is that Dr. Lee was interviewed for television here regarding the publication of the study.  Hopefully this doesn't seem odd to some of my readers, but in the U.S. science stories almost never make the news.  In general, our media is very anti-intellectual.  When a study does make the news it is almost always about bigfoot or some conman who hunts ghosts or is searching for places in the bible that are obviously allegorical and not real places.  Large high profile NASA missions make our news but they are almost always quick blerbs, unless the mission is a failure and the media can use it to criticize science in specific and NASA in particular.  In the U.S., you don't get on TV for doing something smart, but you will be a daily feature by proudly flaunting your ignorance.  Sorry if that was a tirade.  I'm just so very impressed that there are cultures that respect scientific achievement and I'm so glad that I get to spend time in one of them.



Thursday, March 28, 2013

LuciPHOr -- phosphosite localization for the TPP


LuciPHOr is one of the gems that I ran into at KHUPO.  I spoke to one of the developers and I can't wait to try this out.  Upcoming weekend project:  same dataset LuciPHOR and the TPP vs Proteome Discoverer 1.4 with PhosphoRS 3.0.  Would LOVE to see this comparison.  Considering I'm composing a phosphoproteomics paper right now, this might be a perfect place to use it.
Anyway, LuciPHOR is a phosphosite assignment program specifically designed to work in conjunction with the transproteomic pipeline.  It is also primarily written for Linux.  This shouldn't be a problem for us Windows slaves, cause we can probably just run it through the command line once we've compiled it.  If I like it nearly as much as I expect to, I'll write up a GUI for it and post later.  Again, this might fall onto my "to do list when I finally find that hyperbolic time chamber list"  or "things for my clone to do the next time he's on a 16 hour flight and not intoxicated list."
Wow, I digress.  The really cool thing about LuciPHOr, and why I am so totally sold on it is this:  it has a FDR calculation for the localization of the phosphorylation.  They call it the "False Localization Rate".  It uses a lot of high level statistics, not my forte, but the data that I have seen looks really promising.

Korea HUPO Day 1


I just wrapped up day one at KHUPO.  What did I learn?  There is some top-notch proteomics research going on here in Korea.  Novel approaches to some old problems, ambitious tackling of some new ones, and some really really nice looking software.
I'm still a little unclear on what I am and am not allowed to write about.  The only thing that I have definite permission for is some new software, but I want to take it for a test drive first.  This is just a teaser, I guess, but it's hard to walk out of there and not write something!

Wednesday, March 27, 2013

The ATP-binding proteome of Tuberculosis


Currently in press at MCP is a really cool project that makes amazing use of the Thermo(Pierce) kinase enrichment kits.  This paper from Lisa Wolfe et al., represents the work of a group from several institutions and uses the kit to find the proteins that use ATP during the Mycobacterium tuberculosis life cycle.  If you're unfamiliar with these kits, you should really check them out.  What you get is this pure ATP or ADP that are tagged with biotin.  You place these compounds into your experimental system and if, say, your drug treatment (as I've used it) , causes a whole lot of kinase action then the respective tag that you used will be integrated into the binding site of the protein.  You then have a permanently tagged protein that you can pull down and implicate in the function you are studying.  They can be a little tricky to optimize.  This is a high level experiment, but to have your kinases enriched and to know that they are active can give you more information than just about any other experiment.  The application of these kits is pretty much up to your imagination and your biological system.

I stole the picture above from the PierceNet website, but you can find more information here.  If you are interested in tuberculosis or in a great way to apply this technology, you should definitely check out this paper!

Tuesday, March 26, 2013

ProteinLasso -- use super statistics to estimate protein interference

There is no art whatsoever on the ProteinLasso website, so I made my own icon with the help of Google Images.
Anyway, in every complex MS/MS experiment we ultimately select for fragmentation some ions that we didn't mean to.  Increasing sample complexity, increasing the isolation window and decreasing the chromatographic separation all exacerbate this fact.  There have been a number of different approaches to estimating or dealing with this.  ProteinLasso is a new approach described in this recent paper from Ting Huang et. al., out of the Dalian University of Technology.  The approach here, as far as I can tell is some high level statistics called lasso regression.  My expert eye can pull a lot of very large capital letter Sigmas in both the figures and the text.  The end point, however, is pretty clear -- false discovery rates calculations that appear to work with the same degree of efficiency whether the sample is simple or incredibly complex.
 We'll spend some more time on FDR in the near future, but you should check out this paper.  The software is also available through sourceforge if you just want to plunge right in.


Is Paris Hilton interning with Steve Gygi?


In a bit of silliness, and something that approximates news (or at least what we seem to consider it in the U.S....)   I was looking at the current software offerings from Steve Gygi's lab and was surprised to recognize a face in the lab photo.  Paris Hilton appears to be working in the Gygi lab, or at least stopped by for a visit.  In casual conversation, you often wonder where these starlets go -- mostly to rehab, but not Paris!  She's in one of the premier U.S. proteomics labs.

Monday, March 25, 2013

Excessive carry-over in your LC-MS system?


This isn't news, just another random topic for discussion (monologue).  Here is the question, though:  how do you know if you are having excessive carryover in your LC-MS analysis?  There are lots of ways to check this, but this is what I do:  I run lots of blanks and I process almost all of them.

Expanding:  in between samples of importance, or in-between quantitative runs without internal controls (label-free or SRM or whatever) I inject a normal size sample load (2-10uL, depending on the LC in question).  I then process that sample using an appropriate database through my normal processing scheme.  Since I almost exclusively work with human samples these days, I simply use my Sapiens Uniprot FASTA (or IPI Human) with the cRAP database either tacked onto the end or searched in parallel.  This gives me a pretty good metric of how clean my sample is.  I don't freak out when I see some peptides.  I only freak out when I see a lot of peptides.  These instruments are so sensitive that they are going to find peptides floating in the air or reasonably fresh buffer.  They shouldn't, however, find dozens of peptides.

If I do find enough peptides to freak out, these are my steps:
Clean the front of the mass spec.  Capillary and front plate with 50% methanol should do it.  If that doesn't help, it is time to approach the LC system.  The order varies from person to person, but this is how I approach it:

1) Change my blank
2) Change my wash solvents
3) Change my running buffers
4) Run high solvent for a long period of time (a few hours to overnight, if I could possibly afford it.  Extremely rare that I could or can)
5) If none of these help, I try injecting something harsh (a full sample loop of 20% isopropanol):  Note:  This may not be appropriate for all nanoLC systems.  I don't think I've ever used bold print in all the years I've been writing in this silly blog.  But I don't want you damaging your LC system.  I don't know a lot about LCs, just enough to successfully get by and to know that I've never noticeably damaged the LC systems that I have had.  That doesn't mean that you should trust me on this one.  When in doubt, consult the manufacturer.
6) If this doesn't help, it is time to approach the guard column (if employed) and finally the analytical column.

End of tirade.