Monday, April 30, 2018
Orphan kinase demonstrates remarkable phosphorylation control in malaria!
Ummmm...yeah...I have to come back to this one....what...?...gonna need more espresso to tackle this one....
Great TMT phosphoproteomics + knockout of a weird "orphan kinase" that has a lot of homology to all sorts of other kinases we know about = really weird downstream effects suggesting that we really have no idea what happens when protein X is phosphorylated in stage Y in P. falciparum's life process.
Intimidating because...well...we have tons of beautiful diagrams (like this) saying what and how phospho cascades work in this organism.
A relief(? no. probably not.?) because it still seems to be doing all sorts of mysterious things despite how hard really smart people are working to try to figure it out.
This is some top notch work that shows how many mysteries there still are in a disease that infected around 200 million people in 2015 -- and killed over 400,000....
Handy Venn diagram for amino acid properties
I just saw this great figure on Twitter. Shoutout to @reducentropy for posting it.
It is from this incredibly useful chapter on amino acids that you can find open access here. This is getting posted on my wall....
Sunday, April 29, 2018
Prot-SpaM -- Fast alignment of PROTEOMES for phylogeny reconstruction
It is really hard to type this because the picture is just cracking me up. Focus, Ben. Stop laughing.
Prot-SpaM!
The genomics people do things like this all the time -- take the DNA sequences and figure out how everyone is related to whom. Prot-SpaM allows us to jump in and look at the phylogeny of organisms (alignment free) from the protein sequences.
The results look seriously cool when they come out of the software but you have to zoom in like 100 times to tell that you aren't just looking at a weird smear on the page.
Those are compressed circular dendrograms!! How cool is that?
Now here is the neat thing about this -- classically this stuff is ALWAYS done with gene alignments (essentially permutations on just BLASTing the entire genomes of organisms). This requires supercomputer level resources. Prot-SpaMming doesn't. It requires far less resources AND (this makes sense to us protein people) protein sequence information appears to be a far more sensitive method for detecting relation than DNA alignment. Wins all around!
Saturday, April 28, 2018
MetaUniDec - Deconvolution software that can even handle native spectra!!!
I'm currently (well -- not right this second -- each run is only like 12 minutes -- but...) comparing a protein or 3 under both reduced (acidic conditions) and native (non acidic -- tried like
A minor problem is that there is exactly one computer with software that can handle deconvolution of native proteins in our building.
WooHoo! Thank you Deseree Reid et al., for this awesome free deconvolution software that is exactly what I need!
I was kinda bummed because the paper leads you to a download link to the software that looks like you have to use Python to use it (which isn't so much of a bummer since a guy in my lab is a Python expert) BUT then I found this link that leads you to a GUI!!
This is what it looks like in action --
See all these things it does? It has features that are not present in our commercial package -- including some things that I really wish that it did.
Now -- you do have to get the data into the correct format first -- from the manual (which is online here)
Okay -- so that's a lot of words and I don't know what most of them are. FORTUNATELY -- all this AMAZING software is looking for is a text file to deconvolute. That I can handle.
You can make Xcalibur do that for you. First, find your ugly protein peak (you don't get to see mine -- the peak is like 3 minutes wide -- it's essentially just a desalting column)
Then go to Excel and paste it!
Whoa! How cool is that? The RAW file doesn't start at exactly 2,000 (m/z)!! You'll have to believe me, but it doesn't stop at exactly 6,000 either! Quads get kind of wobbly at that kind of range, right? Makes sense to me.
Now you can make that into the text file and load it. Okay -- actually -- I'm having some trouble with the formatting on it. While my first suspicion is that this was designed for a TOF and the fact I have mass accuracy is probably the issue -- it might be the header format. It's a Saturday night and I've still got to do some sample prep so I'm going to worry about this later.
Even better? You can just directly paste the clipboard in.
It's even better than I thought!! Check this out.
Okay -- this will clearly need some optimization -- but with no use of the manual and I've got a nice peak out of this software on my first shot. For perspective, I don't know how to set the parking brake in the car I got in December. I just don't park on hills.
It was probably dumb to start with an unknown (that appears to suffer from some degradation with in-source collision energy at 90eV) so I put in one of my QC proteins. Check out this cool output deconvolution plot!
The protein should be 42,882 -- this trippy output requires some finagling (which appears to be a word?) but the software looks spot on and I haven't scratched the surface of what I can do with this awesome new tool I now have!
Friday, April 27, 2018
Mango -- Clean up chimeric spectra for crosslinking experiments!
Yes there are around 40 different software packages for crosslinking analysis. And some of them actually work!
However -- Mango (JPR paper here) looks like it meets a specific and important need -- cleaning up the chimeric (multi peptide) spectra for processing!
You know the way that ETD totally doesn't look like it worked unless you clean up the data? I think Mango is going to allow us to look at crosslinked MS/MS spectra that we couldn't extract anything from before -- and provide us with all sorts of new information.
Wednesday, April 25, 2018
MSstats -- Way more than just QC!
MSstats has shown up in my ramblings more than once here, but always in the context of MSstatsQC. I just sat through an awesome talk that demonstrated that it is capable of much more than this.
You can check out MSstats directly here.
Highlights? An R package that can take data from MaxQuant, Proteome Discoverer, Skyline, OpenMS, OpenSwath, and other stuff (its at least 7 of them) and make sense of it all with advanced statistics all over the place.
The inference algorithm appears to be called the "Accelerated time failure model". I don't know yet how it compares to the ones we more typically use (like the k nearest neighbor) but it sure sounds way cooler.
There is so much power here. Choose wisely.
Tuesday, April 24, 2018
MassIVE made some umm...massive...spectral libraries
I've got pages and pages of notes from ABRF already and as I'm sitting here trying to organize them I'll probably pull together a few blog posts out of them.
On Sunday Nuno Bandeira talked about MassIVE. Of course, I know about MassIVE. It's one place where you can deposit your RAW data so the journal editors will leave you alone about it.
However -- it's not just sitting there. Busybody bioinformaticians are combing through the data trying to find new things (succeeding) -- and they are compiling huge spectral libraries.
What do you get when you compress the most meaningful data out of 30TB of HCD fragmentation spectra? Other than a file that takes a REALLY long time to download on a hotel WiFi connection? Over 2million annotated spectra.
I may have to give up on downloading it -- or remote login to something on a much faster connection.
Now -- the question remains -- how does this help me? I was hoping that after I had it I could see if I had anything that could open it or use it as an input (MSPepSearch maybe?) ....to be continued....
Monday, April 23, 2018
MassSpectrometryMethods.org
Wow! I'm fully aware that this blog has recently been worse even than normal. Part of this is a result of this other project that I'm working on that has been revealed at ABRF today and a description of the project will show up in bioRxIV sometime this week.
If you've had the misfortune of reading this blog for a while you might remember that there was once a tab over there --> that said something like "Orbitrap Methods Database"
This was a problem for me for a lot of reasons. First of all -- I'm not going to write the best instrument method in the world. Chances are I can write you one that will work (if it's stuff I'm good at -- I'm going to write a method that will work pretty darned well) but it's crazy to think that you should be running your instruments the way I tell you to. There were other downsides and I really had to take it down -- always hoping I could bring it back somehow in a smarter way.
What about this?!?! What if I could set up a website and Google Team Drive and get a ton of instrument methods -- then -- could I convince some great mass spectrometrists to get with me a few times a year and go through the methods and pick the very best ones?!?
WAIT -- Could I also come up with a completely ridiculous way of recruiting help that requires me to put a dog in a costume and make him do a pose he really doesn't want to do...?
Was this the project I was born to do?!?
(This picture is on the ABRF poster)
Most jokes aside -- here is the idea --
If you are brand new to mass spectrometry and your instrument just got installed -- HOW DO YOU DO ANYTHING?
If you are awesome at shotgun proteomics and someone asks you to get an accurate mass of an intact antibody -- how do you get going? Do you know that in source fragmentation is a critical parameter?
Maybe you just go to Www.MassSpectrometryMethods.Org and you download the "Intact protein > 60kDa" method for your instrument. Now you've got a starting point.
Now -- back to the initial idea -- and the picture of Gusto above -- if this is just me putting methods in a Google Drive -- this idea is dumb. However -- if I can get some of the top experts in mass spectrometry to sign on and help me out -- can you imagine how cool the next release could be?
I don't know how to do lipidomics. I'm not very good at Top Down (I'd like to get better). The metabolomics methods could definitely improve. Heck -- the shotgun proteomics methods could be better.
If you'd like to help out -- check out the site and shoot me an email -- its now LCMSMethods@Gmail.com.
Friday, April 20, 2018
FRACTION OPTIMIZER -- Take the guesswork out of 2D optimization!
i had to disable my caps lock key or it would look like i was crazily shouting about how much i love this new study and software throughout this whole post. you can check out this awesome new tool here!
if you do highph offline fractionation followed by the same low ph gradient on every one of these fractions and you look at the data coming off you'll think something like "wow...i really could get more ids if i optimized each fraction separately."
but that's a lot of work (and it will impact your reproducibility if you are doing something like label free quan). plus it would be a lot of work. one run to see the relative elution times and a second with the reoptimized gradient..?...
fraction optimizer can do this for you!!!
100% recommend you check it out. you can get the software directly here.
Wednesday, April 18, 2018
PRESTO! Collect all sorts of data and let MatLab sort it out.
EDIT 4/24/18: If you don't have MatLab there is a PRESTO stand-alone! You can find it here.
I can clearly remember my excitement when I first saw 2 different "-omics" datasets integrated together (that totally worked) in a great study. It might still be in the 1,000+ entries on this scrambled blog, but who will ever know (the search bar appears to lose power as the entries get older -- which is honestly fine by me, I've said some dumb stuff over the years)
These days -- I'm still impressed -- but it's a lot more common. However, I haven't pulled one off myself yet but I'd sure love to....but where do you even start....what about here?
First impression -- wow -- this looks seriously powerful. This team pulls in a bunch of different datasets -- microarrays, proteomics, my heart jumped for just a second because I thought they were also integrating CyTOF data (I don't think they do here, they just discuss the statistics involved) and through "t-stochastic neighbor embedding" --
(PRESTO! )
-- they massively reduce a staggering amount of signals from various sets to very small and shockingly meaningful observations.
Is it a trick?
Honestly, I can't say for sure, but the results seem logical and really impressive. (Two gifs is probably too much for one post).
I start to get nervous as soon as we start talking statistics things -- but if it gets you to a small number of targets that you can validate (and they do) it's a WIN.
I made sure to mention MatLab in the title of this post. MatLab is not free software (though trials are available and many many University's have license deals for the software). There is also a home version that is roughly 1/10 the normal license price, but has some limits on it's functionality. If you have access to this software --- and have some huge and intimidating "omics" data sets you should definitely check this out!
EDIT: I accidentally reread this and I feel like I didn't emphasize this study correctly. PRESTO is a utility that these authors developed that runs in/on MATLAB. I don't mean the title of this post (which -- to some extent -- can't be changed now) to detract from the awesome amount of work PRESTO is.
Subscribe to:
Posts (Atom)


