Monday, January 16, 2023

The TransProteomic Pipeline still exists!

 


Once, long long ago there was this open-ish cloud-like data processing pipeline with a name that would, for some reason, make you think of Santa Claus and hair metal. Not the good hair metal, but the stuff that they play on cellphone commercials and in elevators. 

It was called the Trans Proteomic Pipeline and I just learned it still exists thanks to this new paper! 


There are some really helpful tables inside the paper that tell you what data it takes and what it doesn't from the "dizzying" number of mass spec data formats that we have. 

Honestly, it looks pretty comprehensive. Since the TPP is hosted on someone else's computer, you don't necessarily need a ThreadRipper on your side of the internet to do really sophisticated data processing and manipulation. 

If you are interested in checking out this pipeline and don't know where to get started, check out this page of tutorials! http://www.tppms.org/tutorials/

Wednesday, January 11, 2023

Find the super low abundance proteins with MRMHR (PRM with Zenopulsing)!

 


This isn't self-promotional because I didn't do this study. Dr. Wheeler did this study. My name is on it because I showed her how to use MaxQuant, introduced her to Skyline and...I guess funding in my name helped pay for her to wrap it up the last 6 months or so. But I guess that's how PI stuff works.


Punchline? Sometimes people who take HIV drugs have negative neurological effects like dementia. We don't know why. What Dr. Wheeler found was that if she took purified endoplasmic reticulum from different brain regions and treated them with HIV drugs, some metabolism occurred. Metabolism is supposed to happen in the liver, not the brain, right? 

While class I metabolism (P450 type oxidation) had been previously observed, class II metabolism stuff really had not been, and that was what she was seeing. And that didn't make sense. 

The QE couldn't detect enzymes capable of this process. DDA/DIA/PRM

The TIMSTOFs didn't detect enzymes capable of these processes (PASEF/diaPASEF).

Targeting on a TIMSTOF is a miserable experience. The files are huuuuuuuuuuuuge. 

Enter ZenoTOF and ZenoPulsing enabled PRM, which they call MRMHR. MRMHR produces tiny output files that are super simple to process in Skyline with ridiculously high sensitivity. 

The first author was able to demonstrate that enzymes matching the metabolites observed in the correct brain ER cell types all lined up. I can't remember if this made the paper, but mRNA could back it up as well.

I know this didn't make the paper, but she also fractionated a bunch of peptides from these brain regions and did really comprehensive proteomics using the same instrument. Using the ProteomicRuler we were able to get some ballpark estimates for how rare the enzymes she is seeing are. Current best numbers are these are at less than 300 copies per cell. Which ain't much. Files are all on MASSIVE at MSV000090576. 

Normalize and integrate your organ-on-a-chip proteomics and metabolomics!

 

Organ-on-a-chip technology can allow a better representation of in vivo conditions than 2D cell culture can. It clearly isn't as easy as 2D cell culture, but it can also be a good bit easier to scale than animal models. I didn't know it until I read this paper, but the US EPA plans to stop animal testing in their research plans in a few years and these technologies will likely play a big role.

If you are jumping into this field and want to fully characterize your drug (or in this case -- CHEMICAL WEAPON) on your simulated organs with both proteomics and metabolomics, what would you do? I'd just do what this group does!


Hey! I know some of these people...including the...senior...corresponding....author... Wait, that can't be the kid I know. Common name. 

What they do is grow human hepatocytes in scaffold bioreactors to get them to grow into 3D cell cultures. It doesn't look like any funny gelatin things are employed here (whew....all I can see when they grow cells in gels is...their gel monomers...). They take some control little organs and some other ones they treat with "VX" which, according to Google is something like a CHEMICAL WEAPON.

Since they're making tons of these little "organs" they appear to keep their digestions and extractions in plate to keep them scalable. Apparently, though, the size of the "organs" are tough to normalize and since you've got these cells all stuck on scaffolding you can't exactly weigh them. You have to basically lyse them where they are. Proteomics normalization? That's okay, we basically know how to do that (though I'd follow their directions here to make sure that buffers necessary for the "organs" don't interfere with the quan). 

Normalizing the small molecule/metabolomics input? I have no idea how I'd to that. You can normalize the TICs, etc., but that is for relatively small variability. You don't have to load too many samples for LFQ proteomics with vastly different loads on accident to discover the limitations of normalizing at the TIC level, right? What this group comes up with is a rapid method to quantify their metabolomics based on something that reminds me of the name of one of these people...


--devitolation! 

Joint pathway analysis (!!!) for the metabolomics and proteomics was done in Metaboanalyst, cause apparently it does that now

Tuesday, January 10, 2023

Baseline protein expression in 489 people!

This study might not necessarily be the most helpful to me personally, but I totally know people who will love this new resource!  


This is a clever attempt to make proteomics data more accessible to our friends in the wet labs and I wholeheartedly support that. What is the normal protein concentration distribution for a human spleen? Digging through this resource is a whole lot easier than getting a couple human spleens and digesting them and doing proteomics yourself! 

Monday, January 9, 2023

Pretend you know how to use R with MSStats Shiny!

 


Maybe my bar is unrealistically high since I happen to know a real bioinformatician or 3, but I saw a CV recently that was ROFLMAO (I genuinely don't know what that stands for) level inflated. It turns out that sometimes people using web based Shiny apps might think that since they can successfully load data into webpages that they are now experts in R based bioinformatics. Why not? There are definitely people out there successfully selling themselves as proteomics experts that have dropped off samples at cores or have loaded sample queues in instruments maintained by real experts. 

With all these great new web based tools you can keep the ruse going if you want and MSStatsShiny will certainly help! This tool has been live for a long time now (it feels like years and years) but now the paper is out. 


MSStats is something I've honestly never used except through the ShinyApp and I know that it is often considered the gold standard. 

The first thing I checked when I saw the paper was out was for DIA-NN compatibility, and it doesn't appear to be ready for that yet, but there is a ton that is IS ready for.


Including Proteome Discovererererererer! 

Look, I know that no one likes Discoverererer as much as I do, but as I'm training a bunch of new people this year, I keep showing them it because nothing gets us to the PSM level visualization as fast. 

Off topic, but if you were wondering what the upper limit is for PD, I can say for sure now that it can easily handle more than 2,000 single cell proteomes (250-ish SCoPE-MS runs not counting blanks/controls, etc.,). AND if you randomize your TMT channels and your injections for two closely related cell lines derived from the same type of cancer (different patients and immortalization methods) once you get north of 50 injections --  THEY CAN FIND EACH OTHER BY PCA ALONE. That might sound like not a big deal, but try that with 200 scSeq files from 2 different cell lines. PCA? Not a chance. You need real dimensionally reduction for that and cell cycle clustering and probably some level of stream-type trajectory analysis magic to find your cell type friends. I'm not sure PCA will even make sense of n<1000 per condition for scSeq. Proteomics data? BOOM. AND it gets better as you keep going (the divergent clusters are part of my experiment, don't stress it). The biggest question for me was whether I was going to break PD with 500 SCoPE-MS files. Half-way through it ain't even really struggling, so I like my chances. 

(If you are stupid enough to try processing 4,000 single cell proteomes in a desktop interface I strongly encourage you to turn off "Found in Samples" and "Abundances" in your PSM and Peptide Group level reports. I still strongly recommend leaving the Data Distributions post-processing node in. It is just too useful, but hide the results when looking at your full reports because it wills struggle with a matrics with 4,000 proteins X 4,000 conditions. (Imagine what that is at the PSM level. Even though it isn't maxing out my RAM something is limiting the speed at which I can sort through the data). 

Pretty sure I need to run MSStatsShiny locally for this analysis, but I might try the web interface anyway. 

Sunday, January 8, 2023

SyntheDIA -- Simulate your sample for experiment optimization!

 


Yo, you know what ain't getting cheaper? Instruments and instrument time! You'd think with all the competition out there now, we'll soon see vendors competing with pricing, and I bet we will soon. (For real, we should share pricing information, somebody put up a post on Reddit where you can put your quote information in).

One way you could save time and maybe a lot of money is not running optimization experiments for new samples (or for getting to the proteins you actually care about).

That's what SyntheDIA is! 

You can read about it here


Or you can check it out (along with some disclaimers that I didn't find immediately apparent in the paper itself) at https://synthedia.org/

If you just type it into a search engine you'll find that most of them think you are terrible at typing and you mean to swap the D to an S. 


The web interface can directly do some SyntheDIAing for you, but if you want to do more than a couple of files you can install it locally. The authors are also open to collaborating (which is a really smart thing to have listed on a cool resource). 

Thursday, January 5, 2023

Use super fancy DIA approaches for DDA with Scribe!

 


You know how we went from 
"DIA is cool, but WTF do we do with all of these crappy noisy spectra??" 
to
"We have 14 tools that do amazing neural learning machine stuff that we trust so much that we won't even look at these crappy noisy spectra anymore, they're probably okay cause this machine told us it was" 

(That was a fun ABRF core director meeting when everyone was like....ummmm....what....guess this is what we do now...???)

That stuff is probably all now very well validated, or if it isn't, it will be very hard to find anyone to blame for it now. Point of fact, though, deep learning intelligence thingies are much much better now. You can even make them write code for you

What if you applied this knowledge to D D A data? Well, that's what Scribe is! 




Wednesday, January 4, 2023

MS Imaging -- Awesome webinar Monday 1/9/2023

 


Hey you! Are you considering buying the expensive mass spec that makes the pretty pictures? Want to find out what it can actually do (beyond making pretty pictures!)

Sign up for this webinar on Monday! 

Details are here!

Tuesday, January 3, 2023

Publication ready spectra with mirror plots -- you'll want to USE this!

 


...and the winner for the hardest tool to find when using basically any internet search engine...USE

Universal

Spectrum

Explorer

I've been looking for this thing for over a year! 

You can use USE here: https://www.proteomicsdb.org/use

You can get the code to use USE locally here! https://github.com/kusterlab/universal_spectrum_explorer

You can read about it here

Think it looks like the amazing IPSA that you're already USEing? It ought to, it is an expansion on that great resource! 



Monday, January 2, 2023

Use proteomics technology around the home!

 


Happy New Year! 

I'm going to kick off 2023 by showing you my LEAST FAVORITE PHOTO OF 2022. This is a picture of an attachment on some excercise equipment that I have outside under my deck. This pretty lady and her bag of eggs were right by my face for a minute before I zoomed in on this picture and she had an unfortunate accident while I tried to carefully rehome her with an old aluminum baseball bat. 

Pug Mountain Vineyards is a spider friendly establishment but Latrodectus isn't a spider you mess around with when you have a 1 year old running around.  Fortunately, these spiders are well known to use ballooning to disseminate. They don't cluster in large populations. Any of her eggs that survived a very thorough removal process by myself and by a well compensated professional who came out twice would have only produced small spiders that would have moved on to nicer places THAN UNDER MY DECK. 

Obviously a small spider that I found on an uncharacteristically warm day right before the holidays that looked sort of like her would just be coincidental appearances. I've got bifocals and spiders are small. Just in case a tiny amount of it that remained from the unfortunate accident that befell that little organism, clearly wouldn't S-Trap out and show any Latrodec...tus....hesperus....

...so...I think I'm moving....but, hey! who said you couldn't come up with a practical application for proteomics!