Wednesday, January 29, 2020

Skyline for small molecules/metabolomics and Skyline 20.1!




Skyline has had support for small molecules and metabolites for years now-- but I still have a lot of trouble setting it up and have to bug smarter (and typically younger...) people for help a lot.  What I could use is a Step-by-Step protocol and template files I can download. 


















While I'm on the Skyline topic -- I just got this great email overnight -- Skyline 20.1 is up. 

It does require a manual download and install (which you can download here) -- but the Skyline team hasn't forgotten about proteomics.

Edit: On my laptop that has the Windows 10 disease, I did have to manually remove the last install of Skyline and reboot to install 20.1.

I've trimmed the email to remove any mention of command line Skyline and stats words I'm unfamiliar with. And highlighted my favorite parts. 

MSFragger spectral libraries!! Pull out the dark proteome and then quantify it?!?  For the biopharma groups that are finding multi attribute monitoring the most cost-effective way forward? Supported! 


Improvements since Skyline 19.1 include:
  • Prosit spectrum and iRT prediction support directly integrated into the UI
    • Building libraries for targeted peptides in a document through Peptide Settings - Library - Build button.
    • Prosit spectrum prediction viewing in the Spectrum Match plot with new right-click menus, including mirror plotting
    • Settings in Tools > Options > Prosit
  • Support for spectral library building from MS Fragger pepXML search results
  • Support for diaPASEF!
    • We have run this with 2 separate 3-organism datasets through the LFQBench statistical assessment and that works.
  • Improved ddaPASEF and initial prmPASEF support.
  • Performance gains in importing Agilent and Waters IMS data as much as 2x or more.
  • Parallel file import with proteomewide DIA in the UI or by default on the command-line has performance similar to what was previously only available from the command-line using --import-process-count. Choose "Many" on your next import or just ignore threading the next time you import from the command-line.
  • Optimized spectrum memory handling for instrument vendors with .NET data reader libraries, benefitting Agilent, SCIEX, and Thermo
  • A new "Consistency" tab in the Refine > Advanced form, supporting CV and q value cut-offs
  • New checkbox for Refine > Advanced - Results tab Max precursor peak only
  • Support for Multiple Attribute Model (MAM) grouping with Peptide.AttributeGroupID and PeptideResults.AttrributeAreaProportion
  • Added File.SampleID and .SerialNumber (of the instrument) as fields in Document Grid custom reports
  • Transitions Settings - Full-Scan - MS/MS filtering has been extended to apply to all non-MS1 spectra (e.g. MS3) as long as the MS1-level precursor matches the target precursor m/z The redundant library filtering phase of spectral library building is around 20x
  • Improved iRT calibration UI making it easy to create new sets of standards based on existing sets that can be used in spectral library building and the Import Peptide Search wizard
  • More iRT improvements including more intelligent use of 80+ CiRT peptides when CiRT is chosen during library building
  • New right-click > Quantitative menu item for changing the Quantitative property on transitions in the Targets view TIC and BPC now come from raw data files and do not need to be extracted from MS1 spectra which has performance benefits for MS1 filtering
  • New global "QC" transitions have been added such as the pressure trace
  • Calibration curve fixes to make ImCal (Isotopolog Calibration Curves) work
  • New "Calculated" annotations have been added which support storing Skyline calculated values in annotations for future use with AutoQC
  • Support KEGG IDs as molecular identifiers in small molecule targets.
  • Improved support for D used in chemical formulae in place of the Skyline default H'
  • Added support for Thermo Exploris and Eclipse instruments
  • Support for opening .skyp files downloaded directly from Panorama

Tuesday, January 28, 2020

Publicly available (unpublished?) proteomic, metabolomic and lipidomic (MERS-CoV) coronavirus data!

Wow. Do I ever love ProteomeXchange!

Skip my reading and go to MASSIVE and get proteomic data from cancer cells infected with a coronavirus and -- if you're into that sort of thing -- you can get metabolomic data here and lipidomic data here! 


The RAW files currently heating my apartment may not have been published yet, but they are publicly available and I just contacted the uploader, but I'm moving fast because this data is 1) awesome and 2) pertinent

The Wuhan Coronavirus (2019-CoV) has a very close neighbor (possibly the closest according to my rough pBLAST of the entire translated sequence, but that may just be a consequence that it emerged more recently, as sequencing technology has gotten cheaper and more common -- leading to more data) -- that is called the MERS-CoV (here is the entry from UniProt) or Middle East respiratory syndrome-related coronavirus.

The experiment is 9 files from infected Calu cells (appears to be an immortalized and/or human cancer cell line) infected with the virus and 3 files from "Mock" (presumably uninfected).

The files were acquired on an Orbitrap Velos in "high/low" mode (120k resolution MS1 and CID ion trap fragmentation). The files appear to originate from PNNL, where it is rumored they know a thing or two about running mass spectrometers.

MetaMorpheus recalibration shows the MS1 is spot on, something like -1ppm off actual when compared against human and -- get this -- I can get >70% coverage of the main capsule protein from the virus in the virus infected proteomes. This is really cool because that protein is well conserved betweent the 2 (by pBLAST score, anyway).

Update: More fast moving science!! Just because the pBLAST scores line up, it doesn't mean that the peptides do -- check this out!



Again -- big disclaimer -- this is a mass spectrometrist's blog. I know very little about viruses and is just interested in this topic!


Single cell RNASeq + Plasma Proteomics + Machine Learning!



You should check out this new preprint here! 


What a great week or 10 days for proteomics. Holy cow. January was kind of laggy and then -- BOOOOOOOOOOOM --!!

Okay -- in yet another that is going into a file called "January 2020 papers you must read!!" -- which -- is too many words for this cursed Windows 10 thing --


(BERNIE BOMB!)

Back to the paper -- if you just read the abstract, the phrasing will make you think that or friends at the Max Planck jumped on the ScoPE-MS electric Porsche into the future but you'll find inside that more standard plasma profiling (which looks a lot to me like at lot of the clinical proteomics proof - of - concept work we've seen from the Mann lab -- high fractionation, rapid HF runs [relatively affordable instrument!!] for individual patients and MBR). You can read my rambling about one of my favorite of these recent studies here.

Couple that to high throughput single cell transcriptomics and then using machine learning to link the plasma proteome features to the single cell transcripts across 31 clinically derived factors from these patients and -- it looks like the future to me, but it appears they took the Tesla.


...which...of course, that is a thing, right?

Since I'm still rambling -- this preprint was posted in medrXiV, which has some great disclaimers.




Monday, January 27, 2020

Predicting human life span with deep plasma proteomics??


...and the 2020 Grammy for most eye grabbing title goes to....

...this brand new study that is the first or second thing I get to once I'm safely behind the publisher's financial security of a university library paywall....

To be clear, I haven't read this and I'll probably doubly verify the QC/QA checks on my baloney detectors before I do. But if you think there is a force on earth that can keep me from reading this today --


...I mean...besides the paywall....$8.99....

Sunday, January 26, 2020

Wuhan Coronavirus (2019-nCoV) Complete Protein FASTA download


Edit 2/10/2020: UniProt has resource up. These are better. You can check them all out here!

I was looking for a complete protein FASTA database for the Wuhan coronavirus and came up empty.

The NCBI database was just updated yesterday (direct link here) so I pulled the newest sequences and just assembled them into a single file.

You can download the complete protein FASTA from this Google drive link here.

Hit me up if you have any issues with it.

Image above is from this preprint which was updated on 1/22 after it was ORIGINALLY POSTED ON 1/21!! This is how fast science can be, people!



Yikes -- okay, well I guess the way that blew up I wasn't the only person looking for it.

Disclaimers: I'm a loud mouthed mass spectrometrist who knows very little about viruses. I just put all the sequences NCBI translated into one file so the common proteomics software on my computers will accept it.

An Encyclopedia with Quantitative Proteomics of 375 cancer cell lines!?!?!?


Ummm.....whoa....I'm just going to leave this here. This is far too large of a resource for me to tackle on a Sunday morning.

Here is an overview and a lot of links to/around/about the study -- including an query-able -- SQL database in case you're not sure where to put 4,000 Fusion RAW files....

Correction: It's only 500 or so files. Multiplexin'

And here is a short paper about it....




Saturday, January 25, 2020

EPIFANY -- A smart and fast method for protein inference!!!



This new study in press at JPR is critically important for shotgun proteomics and smarter people than me (which means just about everyone) should really take a good look at this and 1) verify it is as good as it looks and 2) see about integrating the source code or primary logic into all sorts of other tools. (An earlier draft was also made available through biorXIV.)



Okay -- so -- shotgun proteomics is really really good at one thing -- making

Peptide
Spectral
Match(es) -- (PSM)s.

And if we're working with the best proteins in the whole entire world then each and every one of those PSMs is unique to 1 particular protein and when we identify that PSM and quantify it in a sample we have proven that particular protein is there and we can even get quantification estimates / measurements on that one protein from that one PSM. (I need a word count on sentences).

However -- from an evolutionary perspective it doesn't make a ton of sense for each protein to have developed in isolation with no relation to any other protein. So...a lot of PSMs could be derived from more than one protein. And if you only identify PSMs that could originate in more than one protein, what do you do?

You INFER the protein identity.
How do you do that?
Well -- probably by a set of mostly arbitrary rules that were chosen because....we had to do something...and it's a great idea if we keep them to ourselves...because they don't reflect well on us or our field.

The best one? When you've got equal evidence, it's probably the biggest protein in your FASTA database....(some tools use the highest percent coverage, but then you'll get all weirded out because if UniProt contains your full length variant and 4 alternative "fragment of" protein sequences you'll only ever see the fragments and then you'll be afraid your lysis method broke off all your C-termini...which...you can't rule out....see...it doesn't sound great when you say it out loud. I hate explaining it when I can tell people are paying attention. I go ahead and get the idea of a "razor" peptide out of the way next, because it's better to get two things that damage your credibility out of the way at the same time and then you can spend the rest of your talk or lecture trying to gain it back.

I'm oversimplifying a complex and varied environment of protein informatics software here. It isn't all this way. From the paper:

"Some methods tackle this problem by either ignoring shared peptides (Percolator 7,8), employing maximum parsimony principles and finding a minimal set of proteins explaining found peptides or PSMs (PIA4 ), iteratively distributing its evidence among all parents (ProteinProphet 9 ) or incorporating the evidence in a fully probabilistic manner (Fido 10, MSBayesPro11, MIPGEM12)"

The best way to do this? An exhaustive recent analysis showed on the iPRG 2016 (the big ABRF study that comes up a lot) that the full probabilistic models are the way to go. More statistics, FTW!

However -- I've only used Fido, but it required a whole lot more processing time/power than even Percolating a large dataset. And this study suggests it's not just Fido...it's a brute force approach that, in the end, may not be realistic.

EPIFANY uses some fancy statistics to achieve the same (better?) inference results, but use alternative logic (something about loopy beliefs) that massively reduce the data processing load.

Full disclaimer -- I'm still trying to figure out how to use it because it runs in KNIME and I might be too dumb for it.  I just found this cool KNIME cheatsheet thing -- with this and the full pipeline and all data available here I'm hoping to work my way through it.  [Hooooly cow. You can run it from command line....how did I miss that!?!? ]

However -- the evidence here is solid that this is a better way to infer protein identifications. The authors test it against multiple datasets including the iPRG and use all sorts of ways to infer the protein identities and EPIFANY is the best -- or close enough -- and finishes in a reasonable time.

And -- look -- even if it didn't work any better at all, wouldn't it be better for us to use the tools that at least tried to use intelligent statistics to infer our protein identities? Grant review boards are grumpy by design. We don't need to give them excuses to fund more transcriptomics.


Thursday, January 23, 2020

Determine if your methionine oxidation is from biology or an artifact!


Okay -- so, despite all appearances, methionine oxidation (Met-Ox)is actually a really important thing. Before I get distracted, you should check out this really smart way of studying whether it is a biological Met-Ox or a sample prep Met-Ox artifact here.



This is an aside, but -- holy cow -- the first 11 papers I tried to find to prove this from home were all locked behind paywalls. I had to go back to this 1997 PNAS paper for something that was open access.

Are you a US citizen and do you think that if your tax dollars funded some research then those results should have to be openly accessible to you? If so, check out this thing some guy set up....


Here is a direct access link to this petition.

With that out of the way -- back to Met-Ox. For real -- this is important. It can be used as a metric for ROS scavenging and for a long time has been thought to be impaired in a lot of diseases and may even be a generic metric of aging.  It just turns out that we don't have a great way of determining what is real Met-Ox and what is an artifact of the myriad ways our field extracts and digests proteins. And now we do! If it looks like Met-Ox might be playing a key role in your biology you can get some heavy labeled hydrogen peroxide and -- ouch -- it is surprisingly expensive, at least at the first suggestion Google had for purchasing it and find out for sure!

Wednesday, January 22, 2020

BioPlex Update Preprint -- 5,500 New Protein Interactomes -- in a new cell line!


Ummm....so on a scale of 1 to BioPlex -- how big is your big proteomics data?  Holy cow. You know, sometimes when you don't hear about these huge proteomics undertakings its easy to think "maybe they thought the first 10,000 human proteome interactomes was enough..."

NOPE. BioPlex is alive and well and providing human protein protein interaction data at a pace that doesn't quite seem possible.

Proof? Check out this new preprint!


Not familiar with BioPlex?  It is a bulldozer type approach to human protein interactions. Instead of doing something complicated and elegant -- why not just synthesize every open reading frame in humans and do an expert level immunoprecipitation -- mass spectrometry experiment on them. Yeah -- every one! BioPlex 3.0 showed about half the theoretical human proteome. For real.

It is a project so big and ambitious that is is easy to forget about. How do you take this another step forward? Well -- you throw in some different cell types. And instead of looking at a few interactomes, you look at a few THOUSAND interactomes.

What on earth do you do with all that data? Besides make the most intimidating plots of all time (which you can do online at the BioPlex Explorer, here), well -- this might be the biggest of the big data for proteomics right now. Did you need an excuse to buy that TensorFlow laptop and take that online course that keeps popping up on that sidebar you can't seem to block anymore on Reddit? To really explore this -- we're going to need those artificial learning machine things -- OR

-- the BioPlex explorer is suprisingly powerful and intuitive!

Check this out -- I've got a protein that is strongly dysregulated in a bunch of samples by both transcript and by proteome. It seems important, but it's been confusing. I'll just put that into the BioPlex explorer -- BOOM --visualizations of protein-protien interactions!


Okay -- so no surprise to me -- this thing has a done of direct interacting partners. One thing that is cool and new here is how different this family of interactors is between the BioPlex 3.0 and the new HCT interactome.

If I didn't know what this protein did BioPlex provides that information and the data is all directly exportable in several formats -- and links directly to AMIGO (which was undergoing maintenance stupid early in the morning when I was writing this)



Around these very practical resources the preprint paper makes some very impressive solutions regarding the human interactome -- and -- let's just say that the interactome doesn't shift on a small scale. The interactome appears to shift on a completely global scale. Which...has some definite ramifications, right?

How many times do you get an IP-MS (AE-MS) that is a pulldown from cell line A and cell line B? Hopefully the main characteristic of that cell line, for example, say homozygous KRAS weirdo terminus in B vs wild type in A? Hopefully that main protein is driving the change in your protein-protein interactions for your bait. But....if you've globally shifted the entire interactome? How does that change your results and confound your downstream interpretation? Way too big picture for me, but something that we need to keep in the back of our minds. Biology is complicated...

TL/DR: BioPlex is growing and is a shining example of what proteomics can be. Send this paper to every biologist you know. My guess is that it's going to be in a big journal pretty soon.