Saturday, February 1, 2020

Posting some friendly reminder from Dr. Yates.


One of the laziest posts I've ever made...but I've got a lot of stuff to do this weekend.....

ASMS ABSTRACT DEADLINE IS MONDAY -- YES. FOR THE MEETING IN JUNE!!


Friday, January 31, 2020

PCR + Mass spectrometry for Coronavirus detection!


This study is a couple of years old, but it highlights a whole clever way of detecting pathogens  -- amplifying DNA and then doing mass spectrometry of the amplicons.

Advantages:
-You can start with virtually no DNA and make a ton of it
-MALDI-TOF is fast when you've got a ton of samples to screen
-MALDI instruments, while maybe not very common in research environments, are increasingly common in clinical labs. (Big question, though, is the flexibility -- I think that a lot of these are locked down to performing one specific assay, but -- still -- those places would have staff with the technical expertise to prep the samples and run the instruments
-If you pick primers well, you'd be resistint to mutations in the viral strains

Cons:
-PCR takes time. Is it faster than it was?
-It takes some people a long time to make primers and to verify they don't cross-react. (I've heard tools have gotten better)
-MALDI is almost always connected to low power TOF devices (sensitivity, resolution and accuracy are the things that are typically the low part)

Check out this alternative technique that would alleviate some of this --



Same general idea -- 15 years ago -- this group amplified their virus DNA (SARS-coronavirus) and then did FTICR....are there a few thousand smaller, faster FTMS devices around the world right now? If you've already done PCR would you even need to couple it to HPLC? FIA-MS on an Orbitrap?

Sorry if you're tired of hearing about these viruses, but I'm motivated to read/write about it

On the topic of the 2019-nCOV (Wuhan) coronavirus -- check out this beautiful resource from NextStrain!

If you are working on this from a proteomics/mass spec/clinical detection perspective and want to talk, please reach out (orsburn@vt.edu). I'd be happy to lend a hand developing/troubleshooting assays or by (much more usefully) connecting you to people who could be of great assistance.

And -- while I'm sleepy rambling -- check this out -- ModPred (www.modpred.org) thinks palmitoylation is a dominant PTM -- if you are developing peptide specific assays -- I'd skip that entire protein terminus.



Thursday, January 30, 2020

MaxQuant.Live 1.2 is live!


I'm a little behind on this and don't have actual data to compare yet, but MaxQuant.Live 1.2 is up for download.

Important factor for many of us -- it still appears to rely on the same Foundation and Tune 2.9 (no 2020 mandatory upgrades if you were already using it in 2019).

The interface looks a little cleaner, but if  you were also hoping for some magical new data acquisition method to appear in the App store, you'd also be disappointed.

Given how active the developers have been on the awesome MaxQuant.Live Google group discussion forum, I think we're looking at a great software that still hasn't been utilized to it's potential, but has had some minor bugs ironed out.



APEX Proteomics applied to Stress Granules Provides insight into ALS progression!

A great way to see trends in proteomics is to go to ProteomeXchange and see what everyone is uploading.

The word "spatial" is completely blowing up. I think there are a couple versions of the modern "spatial" techniques. APEX, however, might be the best example. There is a great overview of this technique on the Krogan lab website that you can view here (I don't want to steal their nice image!)

Want the 6 18 second description/reminder?
1) You make a version of your protein of interest with a peroxidase on it and then you put it into your biological system and let it interact with all it's friends.
2) You put some biotin-phenols that float around in your system doing (presumably) no harm or alterations to your system
3) !!SURPRISE!! your system by adding hydrogen peroxide!
4) Nothing special happens to any other proteins (except normal H2O2 effects, I guess) but your protein of interest has that peroxidase on it --it reacts with the biotin phenols around it which causes all it's local protein friends to be labeled with biotin!
5) Pull down the biotin labeled proteins and now you know all the proteins in the general area of your protein of interest. Cool, right?

(It is a bit more complex than this. You don't want the APEX fusion being expressed all the time, for example, you can wait and activate it when your cells get ot the right stage -- you also need ot quench the reaction, but now I need to change how many seconds this took again...)

However, if you want to see it in action in a medical context -- check out this new preprint!

 
This group used multiple APEX Fusions to study stress granules (SG)s which according to the authors "form in response to a variety of cellular stresses by phase-separation of proteins associated with non-translating mRNAs" (yes, stolen from the first sentence of the abstract).

Unlike a lot of spatial labeling techniques -- this one requires nothing special from the mass spectrometer. The biotin labeled proteins are pulled down and digested. This group used a QE Plus for part of the study and it looks like they upgraded to a QE HF at some point. The LC separation was a long (3hour) gradient on 50cm columns. A relatively cycle time is maintained between the two instruments by using 17,500 resolution for MS/MS on the Plus and 30,000 resolution on the HF. The data processing here is MaxQuant and it looks very typical.

With this system set up, the experiment gets more complex, with the addition of ALS linked dipeptides that show alterations in the stress granules and the SUMoylation -- which is where the biology goes beyond me.

What I do get:
1) A fantastic application of this powerful new technique
2) A method that demonstrates that I could definitely do my part of this.
3) Important new understandings into the progression of a protein disease? 

Wednesday, January 29, 2020

Skyline for small molecules/metabolomics and Skyline 20.1!




Skyline has had support for small molecules and metabolites for years now-- but I still have a lot of trouble setting it up and have to bug smarter (and typically younger...) people for help a lot.  What I could use is a Step-by-Step protocol and template files I can download. 


















While I'm on the Skyline topic -- I just got this great email overnight -- Skyline 20.1 is up. 

It does require a manual download and install (which you can download here) -- but the Skyline team hasn't forgotten about proteomics.

Edit: On my laptop that has the Windows 10 disease, I did have to manually remove the last install of Skyline and reboot to install 20.1.

I've trimmed the email to remove any mention of command line Skyline and stats words I'm unfamiliar with. And highlighted my favorite parts. 

MSFragger spectral libraries!! Pull out the dark proteome and then quantify it?!?  For the biopharma groups that are finding multi attribute monitoring the most cost-effective way forward? Supported! 


Improvements since Skyline 19.1 include:
  • Prosit spectrum and iRT prediction support directly integrated into the UI
    • Building libraries for targeted peptides in a document through Peptide Settings - Library - Build button.
    • Prosit spectrum prediction viewing in the Spectrum Match plot with new right-click menus, including mirror plotting
    • Settings in Tools > Options > Prosit
  • Support for spectral library building from MS Fragger pepXML search results
  • Support for diaPASEF!
    • We have run this with 2 separate 3-organism datasets through the LFQBench statistical assessment and that works.
  • Improved ddaPASEF and initial prmPASEF support.
  • Performance gains in importing Agilent and Waters IMS data as much as 2x or more.
  • Parallel file import with proteomewide DIA in the UI or by default on the command-line has performance similar to what was previously only available from the command-line using --import-process-count. Choose "Many" on your next import or just ignore threading the next time you import from the command-line.
  • Optimized spectrum memory handling for instrument vendors with .NET data reader libraries, benefitting Agilent, SCIEX, and Thermo
  • A new "Consistency" tab in the Refine > Advanced form, supporting CV and q value cut-offs
  • New checkbox for Refine > Advanced - Results tab Max precursor peak only
  • Support for Multiple Attribute Model (MAM) grouping with Peptide.AttributeGroupID and PeptideResults.AttrributeAreaProportion
  • Added File.SampleID and .SerialNumber (of the instrument) as fields in Document Grid custom reports
  • Transitions Settings - Full-Scan - MS/MS filtering has been extended to apply to all non-MS1 spectra (e.g. MS3) as long as the MS1-level precursor matches the target precursor m/z The redundant library filtering phase of spectral library building is around 20x
  • Improved iRT calibration UI making it easy to create new sets of standards based on existing sets that can be used in spectral library building and the Import Peptide Search wizard
  • More iRT improvements including more intelligent use of 80+ CiRT peptides when CiRT is chosen during library building
  • New right-click > Quantitative menu item for changing the Quantitative property on transitions in the Targets view TIC and BPC now come from raw data files and do not need to be extracted from MS1 spectra which has performance benefits for MS1 filtering
  • New global "QC" transitions have been added such as the pressure trace
  • Calibration curve fixes to make ImCal (Isotopolog Calibration Curves) work
  • New "Calculated" annotations have been added which support storing Skyline calculated values in annotations for future use with AutoQC
  • Support KEGG IDs as molecular identifiers in small molecule targets.
  • Improved support for D used in chemical formulae in place of the Skyline default H'
  • Added support for Thermo Exploris and Eclipse instruments
  • Support for opening .skyp files downloaded directly from Panorama

Tuesday, January 28, 2020

Publicly available (unpublished?) proteomic, metabolomic and lipidomic (MERS-CoV) coronavirus data!

Wow. Do I ever love ProteomeXchange!

Skip my reading and go to MASSIVE and get proteomic data from cancer cells infected with a coronavirus and -- if you're into that sort of thing -- you can get metabolomic data here and lipidomic data here! 


The RAW files currently heating my apartment may not have been published yet, but they are publicly available and I just contacted the uploader, but I'm moving fast because this data is 1) awesome and 2) pertinent

The Wuhan Coronavirus (2019-CoV) has a very close neighbor (possibly the closest according to my rough pBLAST of the entire translated sequence, but that may just be a consequence that it emerged more recently, as sequencing technology has gotten cheaper and more common -- leading to more data) -- that is called the MERS-CoV (here is the entry from UniProt) or Middle East respiratory syndrome-related coronavirus.

The experiment is 9 files from infected Calu cells (appears to be an immortalized and/or human cancer cell line) infected with the virus and 3 files from "Mock" (presumably uninfected).

The files were acquired on an Orbitrap Velos in "high/low" mode (120k resolution MS1 and CID ion trap fragmentation). The files appear to originate from PNNL, where it is rumored they know a thing or two about running mass spectrometers.

MetaMorpheus recalibration shows the MS1 is spot on, something like -1ppm off actual when compared against human and -- get this -- I can get >70% coverage of the main capsule protein from the virus in the virus infected proteomes. This is really cool because that protein is well conserved betweent the 2 (by pBLAST score, anyway).

Update: More fast moving science!! Just because the pBLAST scores line up, it doesn't mean that the peptides do -- check this out!



Again -- big disclaimer -- this is a mass spectrometrist's blog. I know very little about viruses and is just interested in this topic!


Single cell RNASeq + Plasma Proteomics + Machine Learning!



You should check out this new preprint here! 


What a great week or 10 days for proteomics. Holy cow. January was kind of laggy and then -- BOOOOOOOOOOOM --!!

Okay -- in yet another that is going into a file called "January 2020 papers you must read!!" -- which -- is too many words for this cursed Windows 10 thing --


(BERNIE BOMB!)

Back to the paper -- if you just read the abstract, the phrasing will make you think that or friends at the Max Planck jumped on the ScoPE-MS electric Porsche into the future but you'll find inside that more standard plasma profiling (which looks a lot to me like at lot of the clinical proteomics proof - of - concept work we've seen from the Mann lab -- high fractionation, rapid HF runs [relatively affordable instrument!!] for individual patients and MBR). You can read my rambling about one of my favorite of these recent studies here.

Couple that to high throughput single cell transcriptomics and then using machine learning to link the plasma proteome features to the single cell transcripts across 31 clinically derived factors from these patients and -- it looks like the future to me, but it appears they took the Tesla.


...which...of course, that is a thing, right?

Since I'm still rambling -- this preprint was posted in medrXiV, which has some great disclaimers.




Monday, January 27, 2020

Predicting human life span with deep plasma proteomics??


...and the 2020 Grammy for most eye grabbing title goes to....

...this brand new study that is the first or second thing I get to once I'm safely behind the publisher's financial security of a university library paywall....

To be clear, I haven't read this and I'll probably doubly verify the QC/QA checks on my baloney detectors before I do. But if you think there is a force on earth that can keep me from reading this today --


...I mean...besides the paywall....$8.99....

Sunday, January 26, 2020

Wuhan Coronavirus (2019-nCoV) Complete Protein FASTA download


Edit 2/10/2020: UniProt has resource up. These are better. You can check them all out here!

I was looking for a complete protein FASTA database for the Wuhan coronavirus and came up empty.

The NCBI database was just updated yesterday (direct link here) so I pulled the newest sequences and just assembled them into a single file.

You can download the complete protein FASTA from this Google drive link here.

Hit me up if you have any issues with it.

Image above is from this preprint which was updated on 1/22 after it was ORIGINALLY POSTED ON 1/21!! This is how fast science can be, people!



Yikes -- okay, well I guess the way that blew up I wasn't the only person looking for it.

Disclaimers: I'm a loud mouthed mass spectrometrist who knows very little about viruses. I just put all the sequences NCBI translated into one file so the common proteomics software on my computers will accept it.

An Encyclopedia with Quantitative Proteomics of 375 cancer cell lines!?!?!?


Ummm.....whoa....I'm just going to leave this here. This is far too large of a resource for me to tackle on a Sunday morning.

Here is an overview and a lot of links to/around/about the study -- including an query-able -- SQL database in case you're not sure where to put 4,000 Fusion RAW files....

Correction: It's only 500 or so files. Multiplexin'

And here is a short paper about it....