Saturday, October 3, 2026

Identification of every protein HeLa has ever considered producing - in one single cell!


Wow. Single-cell proteomics (the hyphon appears mandatory these days) or SCP has come a looong way in a very short time. As one amazing example, in this new paper the number of proteins identified in every cell analyzed went to the absolute moon between the data presented and typing the abstract...

The paper itself is really pretty. I don't know how they generated such high resolution plots and images and convinced this specific journal to not run it through their 1988 Xerox filter. 

The results are even more impressive. In one example, this team identifies a protein that a single HeLa cell growing in a lab in central Iowa in 2008 produced in a total of 4 copies - in one of their single HeLa cells. Not only did they identify the first ever case of single-cell proteomic quantum entanglement but they also identified 18 new post-translational modifications on that protein! I know what you're thinking, this is obviously some super secret new hardware none of us will see until Houston (puke emoji, no I won't be there). You'd be wrong again! They used exactly the same instruments that lots and lots of people have (I don't know how, no one in my city can afford one. I've heard they're less expensive elsewhere? They'd have to be. We can buy 3 nice TIMSTOFs for one of these things). This lab is just way better at everything than everyone else. 

Of particular importance, the abstract clearly states 4,000 proteins in Human PBMCs! This is super important because these tiny cells are incredibly tough to work with. You know who has tried this and not gotten anywhere near this number? 

Me (no publication yet, because I suck at mass spectrometry AND I'm a slow writer) but also

PNNL/Genentech (bums) 

City of Hope (dummies with ....weird...the same hardware as this paper...) 

NorthEastern/Parallel Squared (geez...remember when they were leaders in SCP research...?) 

How did they get these amazing and undeniably field leading results? By doing weird stuff! 

Let's start here on page 10 where things seem very normal. 


On these tiny ass cells (6-8 microns most of the time) they get results that are inline with what we get here, and PNNL/Genentech, the City of Hope, and Slavov collaborative preprints demonstrate.

However, if your real goal for a proteomics study is TYPING THE BIGGEST NUMBER POSSIBLE INTO YOUR MANUSCRIPT AND GETTING AWAY WITH IT this team provides the secret trick for doing this with PBMCs. (I believe this is Figure 8b, the figure legends don't fit well on the pages in the PDF)


In these PBMC populations there are a very very small population (is that a straight line? so ...is there one cell at the very end?? is that how one of these plots work? I think it might be) of PBMCs that are bigger than a lot of the cancer cells people worth with today. 

If you don't...I dunno....sort those cells away because they're ....not...normal....PBMCs..... hell, I'll even open the question that if those are even PBMCs at all. (The term used in the paper is "minimally manipulated" or something). To be nice let's call them GODZILLA PBMCs! 

If you analyze GODZILLA PBMCs you can get a number of precursors or proteins especially if you use match between runs (don't forget to use software that doesn't do match between runs FDR.....puke emoji....angry puke emoji....) and then put this number in your abstract! 

You win! Now you get all the collaborations! 

You probably haven't noticed, but I'm annoyed about this. My biased opinion is that the people in my lab are getting very good at single-cell proteomics. It's what they do every single day. And they recently put a lot of time into cells like these and our collaborators were and are disappointed. We were getting data in line with these new preprints which made me want to give them a call and forward things over. "See? These are hard, we don't actually suck at this and you can learn things from these data." 

And I'd bet you $10 that sometime soon they're going to see this new Nature Comms paper and they're going to think that good mass spectrometrists can get 4,000 to maybe even 7,000 proteins per PBMC, since figures 1-7 are all HeLa. And I'm not the only person this is going to happen to. 

Your PBMC grant application might face similar scrutiny based on your preliminary data. Losers. These things are generally not helpful in any way to the field. They muddy the water, complicate things, set stupid and unattainable expectations, and will ultimately backfire on even the authors when they can't deliver on these numbers on the next study. 

Still, beautiful paper. Despite the Astral and Astral Zoom looking about the same in the plasma proteomics data we've posted recently, the Zoom looks better for SCP. FAIMS also appears to help in both cases. 

Whoops - Forgot one last criticism. The PRIDE repositories are weird. There are three and on the surface they look like they're just the same HeLa dilutions and single cell files repeatedly posted. I assume the PBMC data is there somewhere but they're not easy to find from the SDRFs if they are there. I'll assume the reviewers (of which, I was not one, which is probably a very good thing for this pretty paper) verified that there are actual PBMC raw files publicly available, and it's all fine. 

Friday, October 2, 2026

dNoise TIMSTOF spectra for faster processing and storage!


 Okay...so...like many people, I generally don't care that about the IMS part of the TIMSTOF thing very often. We generally use it like a Trapping TOF. And the Ion Mobility Something(?) is occasionally useful if some reviewer is like "I don't believe that peptide is there". You can go back with a kind of noisy MS1 and MS/MS spectrum or 3 but then you can also show them that some tool also predicted the ion mobilities to line up. This is rare. I generally just need a fast and sensitive instrument.

But....it is an IMS as well. And in this paper, the -now - Dr. Patrick Garrett and et al., said "wait a minute, peptides and their isotopes seem to make a very normal pattern in the ion mobility space. What if we just kept those?" 

You can find the preprint here. 


Ultimately, what you find is that, wait. Can I just copy and paste this summary image from Biorxiv? 



Heck yeah I can! Generally I have to screenshot and then import the screenshot from my desktop, which is currently just 4,000,000 screenshots. 

Do you lose some IDs? Yes, in this iteration it looks like you lose a few. But....can you now process those 500 x 4GB .d files? It sure is a lot more likely of completing before I retire! All the code is up on Github and the files come out compatible with all the downstream stuff, just smaller (of course the originals are maintained) 

Thursday, October 1, 2026

Multiplexed Deep Visual Proteomics!

 


I'm just taking screenshots of this stunning paper to make sure you read it before I get a chance to. It's not mass spec multiplexing. It's image multiplexing, but THEN deep visual proteomics! 



The return of Promis-Quan! (NCTPro)!

 


Wow. What a blast from the past. A new preprint just dropped that describes the triumphant return of what I've long thought was one of the smartest plasma proteomics methods of all time! 

Its called NCTPro, but it is very similar to the old method Promis-Quan, and it gets very similar (amazing!) protein coverage! 

This is the original Promis-Quan post (paper here) and here is the new preprint! 


Wednesday, September 30, 2026

Do you have a rat or mouse you've induced neurological symptoms in? It might be COMT!

 


Hey you! Have you drugged or poisoned or mutated a mouse or rat so that it induces Parkinson's Disease (PD) or schizophrenia (SC) symptoms in? Do you want to fix it? 

Do I ever have a veterinary drug target for you! Check out this new paper! 

DISCLAIMERS! DO NOT, DO NOT, DO NOT ATTEMPT TO EXTEND THESE OBSERVATIONS TO HUMANS. That would be silly for the following reasons: 

Rodents have SUPERCHARGED COMT proteins. Not only do they primarily produce a variant of COMT that is the very active form (there are two in humans), but even the very active form in mice has a key single amino acid variant (compared to the active human form) in it that makes it run in super overdrive at all time (when compared to humans, for sure)

As you might or might not be aware, a rodent brain and a human brain have a lot of other differences. One of the reasons SC models in mice haven't worked very well is that...well....a lot of the proteins implicated in SC don't exist in mice. Not a few. A lot. And there are entire pathways in the mouse brain that humans have evolved completely opposite mechanisms for. For real, people have put years into studying phosphorylation cascades in mice that are completely and totally irrelevant in human beings. However, if your ultimate goal is to cure a disease that occurs largely in the dorsolateral prefrontal cortex in an organism that doesn't....have one.... there is a whole lot of proteomics done here that you can totally check out! I think the proteomics is actually good. 😇😇😇 I think the attempt to extend the findings in this study to be relevant to humans is....a...stretch.... 

Monday, September 28, 2026

diaTracer vs SpectroNaut - same single cells!

 


I'll start with the dumbest analysis possible. That's probably all I have time for anyway.

I reprocessed a small population of large single human cells (TIMSTOF Ultra2 with EvoSep One 40SPD "Whisper Zoom") 41 cells that are EXTREMELY difficult to extract intact out of a cadaver.

No filters, no libraries, just the same library free approach, same FASTA database. Both came back with a little over 7,000 protein groups.

Big difference? On the same PC (an i9 desktop tower purchased last year with a ton of ram put in it before the prices went to the fucking moon) diaTracer wrapped up in about 4 hours from beginning to end. SpectroNaut was close to double that.

The reason this analysis is exciting and not depressing is how the respective softwares name protein groups and genes. SpectroNaut gives you your output as a "Protein Accession; Another protein accession" where multiple proteins could fall in that group. The way I have FragPipe 24 set up I just get a single entry.

I strongly suspect the overlap is better than this. 

Besides being fast, I went back and reprocessed both of these datasets looking for PTMs in these big single cells. And diaTracer found over 650 phosphorylation sites! SpectroNaut found ...fewer.... Are they real? I dunno! It would be super cool if they are! 

Obviously this is only cool if you're an academic who can use FragPipe, but - hot dog - I'm super excited that FragPipe can 1) process my diaPASEF single cell data and 2) it looks like it's at least very in line with software I've used to process thousands of files and have grown to trust over the years! 

Tuesday, September 22, 2026

Spatial transcriptomics is far more expensive than spatial proteomics (deep visual proteomics)!

 


Hey you! You have so many options today for doing spatial measurements. You really do. There are at least 10 different options for measuring gene products in different cells. I was at a couple of recent meetins where biologists were talking about how much these things cost, and it's impressive.

Direct quote from a core director "Each slide costs me $7,000 in reagents" for her favorite spatial transcriptomics technology. She owns all the hardware! 

I pulled some prices for internal at my University and - wow -


There are now ways to do up to 1,200 antibody probes for proteins in addition to single cell transcriptomics. This appears very new (Bruker's new acquisitions) and I've heard it is in line with the spatial transcriptomics. 

Let's break this down for deep visual proteomics (spatial proteomics by microdissection) again this is internal at my University.

For a 10x Visium 2 tissue 11 mm capture for $10,900 you get 50um pixels with 100um center-to-center (according to their website) and adding on proteins is an extra $3,900. Let's call it $15k for 2 images with 100 micron resolution.

This is whole transcript sequencing in most cases so they expect 18,000 gene products theoretically measured across the slide and it looks like 35 protein maximum right now per panel if you add that on. 

For $11,000 if you cut out 100um spots by LCMS we could (internal) do between 200 and 400 of those spots with proteomics (you'd have to find someone with an LCM to do it - I know a guy) and unless I'm doing this wrong, 220 100 micron spots (center to center) would be more than one 11 mm capture. 

A 50 micron pixel would be somewhere in the 8 cell range so we'd expect a proteomic depth of at least 7,000 proteins per "pixel" not "over the total study" or hypothetical.  That's solid depth and could easily be mined for PTMs, etc.,

And....we'd deliver a fully processed visual report in a GUI so the data could be fully mined by the end user.....

Compare that to the Xenium at 5,000 genes and $23,500! Holy shit. Xenium might be far higher resolution, I'm not sure, but if there is a takeaway here -

Spatial proteomics (or deep visual proteomics, if you prefer) is WAY WAY less expensive than spatial transcriptomics.

And....at the end of a deep visual proteomics experiment you don't have to do a western blot to see if any of those gene products actually....you know.....are real at the protein level. Which, depending in how you count the cutoffs, about half of them are actually real. 

Maybe this is a rant, but holy shit, I do love spatial transcriptomics and their pretty slides. But when I see 4 slides in a talk now I know they spent $60k - $100k on their n=1 data. And then they did IHC or western blots after. It seems like a misuse of resources. 

Monday, September 21, 2026

diaTracer - my first impressions (diaPASEF data!)


I've been generally confused about the whole "FragPipe + DIA" thing for a while. What's the DIA-NN node doing there....? Why wouldn't I just run DIA-NN...? (for examples.)

When a new way to process DIA data showed up last year, I had my hands full and just didn't get around to it. I've got SpectroNaut, BPS and automatic command line DIA-NN (academic version). Why would I need anything else? 

What if you are an academic and someone set you up a very nice interface where they'll just load you up anything that will run in Linux? And you don't want to type stuff into a command line? And you don't want to convert your Bruker files? And you don't want to pay anything for it? And you don't want to mess around with making spectral libraries yourself? 

Enter diaTracer! So I'm checking it out vs the other tools and - HOT DAMN THIS THING IS FAST! 


Maybe you can't read the box above but diaTracer took 5 minutes and then everything else took about 5 minutes. That's super fast for these files!  

Also! 
1) I didn't really have to learn anything. I already know how to use FragPipe. (I did have to download FragPipe 24 and then look very hard for the diaTracer button) 
2) I didn't even give this search a lot of resources. I pointed at 2 random single cells, gave it 18 threads (parallelization)  and I have a report in 10.2 minutes? On diaPASEF data? I can't convert these 2 files to .HTRMS to upload into SpectroNaut this fast!  
3) I'll have to compare the results (or make someone smart and meticulous do it.....) I'm just eyeballing the DIA-NN data and they seem similar. 

WHAT??? And here is one of the proteins that I'm looking for! Oh. This might be super cool. 



Sunday, September 20, 2026

May Institute Essentials! In November! Get LCMS proteomics computations statistics fast!

 


Need THE crash course in doing shotgun proteomics statistics? Enrollment is open for the upcoming May Institute (I think that is what it is called. Not the month.) 

You can find out more and register here!