Saturday, October 3, 2026

Identification of every protein HeLa has ever considered producing - in one single cell!


Wow. Single-cell proteomics (the hyphon appears mandatory these days) or SCP has come a looong way in a very short time. As one amazing example, in this new paper the number of proteins identified in every cell analyzed went to the absolute moon between the data presented and typing the abstract...

The paper itself is really pretty. I don't know how they generated such high resolution plots and images and convinced this specific journal to not run it through their 1988 Xerox filter. 

The results are even more impressive. In one example, this team identifies a protein that a single HeLa cell growing in a lab in central Iowa in 2008 produced in a total of 4 copies - in one of their single HeLa cells. Not only did they identify the first ever case of single-cell proteomic quantum entanglement but they also identified 18 new post-translational modifications on that protein! I know what you're thinking, this is obviously some super secret new hardware none of us will see until Houston (puke emoji, no I won't be there). You'd be wrong again! They used exactly the same instruments that lots and lots of people have (I don't know how, no one in my city can afford one. I've heard they're less expensive elsewhere? They'd have to be. We can buy 3 nice TIMSTOFs for one of these things). This lab is just way better at everything than everyone else. 

Of particular importance, the abstract clearly states 4,000 proteins in Human PBMCs! This is super important because these tiny cells are incredibly tough to work with. You know who has tried this and not gotten anywhere near this number? 

Me (no publication yet, because I suck at mass spectrometry AND I'm a slow writer) but also

PNNL/Genentech (bums) 

City of Hope (dummies with ....weird...the same hardware as this paper...) 

NorthEastern/Parallel Squared (geez...remember when they were leaders in SCP research...?) 

How did they get these amazing and undeniably field leading results? By doing weird stuff! 

Let's start here on page 10 where things seem very normal. 


On these tiny ass cells (6-8 microns most of the time) they get results that are inline with what we get here, and PNNL/Genentech, the City of Hope, and Slavov collaborative preprints demonstrate.

However, if your real goal for a proteomics study is TYPING THE BIGGEST NUMBER POSSIBLE INTO YOUR MANUSCRIPT AND GETTING AWAY WITH IT this team provides the secret trick for doing this with PBMCs. (I believe this is Figure 8b, the figure legends don't fit well on the pages in the PDF)


In these PBMC populations there are a very very small population (is that a straight line? so ...is there one cell at the very end?? is that how one of these plots work? I think it might be) of PBMCs that are bigger than a lot of the cancer cells people worth with today. 

If you don't...I dunno....sort those cells away because they're ....not...normal....PBMCs..... hell, I'll even open the question that if those are even PBMCs at all. (The term used in the paper is "minimally manipulated" or something). To be nice let's call them GODZILLA PBMCs! 

If you analyze GODZILLA PBMCs you can get a number of precursors or proteins especially if you use match between runs (don't forget to use software that doesn't do match between runs FDR.....puke emoji....angry puke emoji....) and then put this number in your abstract! 

You win! Now you get all the collaborations! 

You probably haven't noticed, but I'm annoyed about this. My biased opinion is that the people in my lab are getting very good at single-cell proteomics. It's what they do every single day. And they recently put a lot of time into cells like these and our collaborators were and are disappointed. We were getting data in line with these new preprints which made me want to give them a call and forward things over. "See? These are hard, we don't actually suck at this and you can learn things from these data." 

And I'd bet you $10 that sometime soon they're going to see this new Nature Comms paper and they're going to think that good mass spectrometrists can get 4,000 to maybe even 7,000 proteins per PBMC, since figures 1-7 are all HeLa. And I'm not the only person this is going to happen to. 

Your PBMC grant application might face similar scrutiny based on your preliminary data. Losers. These things are generally not helpful in any way to the field. They muddy the water, complicate things, set stupid and unattainable expectations, and will ultimately backfire on even the authors when they can't deliver on these numbers on the next study. 

Still, beautiful paper. Despite the Astral and Astral Zoom looking about the same in the plasma proteomics data we've posted recently, the Zoom looks better for SCP. FAIMS also appears to help in both cases. 

Whoops - Forgot one last criticism. The PRIDE repositories are weird. There are three and on the surface they look like they're just the same HeLa dilutions and single cell files repeatedly posted. I assume the PBMC data is there somewhere but they're not easy to find from the SDRFs if they are there. I'll assume the reviewers (of which, I was not one, which is probably a very good thing for this pretty paper) verified that there are actual PBMC raw files publicly available, and it's all fine. 

2 comments:

  1. Hi Ben,

    Thanks for featuring our work on your blog, and for the entertaining write-up! We were already hoping to see our paper make it to your blog; and had been eagerly F5’ing the website for a few days. :)

    Definitely happy that you appreciate how pretty & beautiful our paper is! We (disclaimer: not *everyone* in the JVO lab) take pride in avoiding any sort of scripting-based visualization like the plague, and rely on making graphs using boomer software like Excel, Prism, or whatever website-based GUI we can find. Then we hope it exports as PDF so we can get into Illustrator and polish everything manually. The more colorful the better. ;)

    As for your satirical approach (that some of us did not realize until ~2/3 of the way through), you make some very fair points that definitely deserve to be discussed. Perhaps even in the form of a dialogue, maybe at a conference where scientists from all over the world can give their two cents on how data should be generated, processing, visualized, and ultimately presented. In the fairest way possible. While still getting past the editors and grant reviewers.

    We would have been very happy to have your input as a reviewer, and we take such feedback seriously. An additional critical pair of eyes never hurts… well, depending on how much extra work we end up having to do.

    So, here is our unofficial (and not legally binding) rebuttal to you, Ben, reviewer #4.

    Regarding the PBMC size distribution, we went into our source data (included with the paper), and out of all the “Small” batch PBMCs, there are 17 out of 847 (~2%) that are above 15 µm in size (and they are all within 15-21 µm). And with the way Prism draws the violin plots, and our love of adding a stroke effect to every violin plot… this created a rather comical exaggeration of an extended distribution. Notably, this distribution included all of the PBMCs sorted (using our CellenOne) from that batch across the entire study, and not all of those necessarily went into the same computational analysis.

    Expanding on this, we found the largest numbers of proteins from the “Mixed” PBMC population. That is, the one where we got buffy coats from the hospital and then followed manufacturer’s instructions on the Lymphoprep™ kit. Here, 14 out of 271 (~5%) of cells were above 15 µm, with those all falling within 15-17.2 µm. We would like to point out to the reviewer that there are many studies out there which demonstrate that PBMCs can, in fact, reach up to 20 µm in size without expressing any GODZILLA genes.

    With regard to data processing, we show a median of 2702 proteins from the “Mixed” PBMCs using DIA-NN in first pass mode (Figure 8D). I.e. no computational matching or boosting whatsoever! Using MBR, this goes up to a median of 3445, whereas Spectronaut (with its great love and appreciation for FAIMS data) got a median of 4163. We debated on which number to give the spotlight, and went for something in between. 4000. And then also put “up to” in front of that number. A clever trick to push big numbers into an abstract? Rather, we believe the number is a reasonable representation of one of our key experiments. On the other hand, as the reviewer is an avid reader of proteomics papers, they may have also spotted some examples where other labs use the “up to XXXX” in relation to single data points from their study, rather than a population median… :)

    There is a character limit on this? Continuing below… :)

    ReplyDelete
  2. Continuing from above…

    Carrier-effect wise, we take pride in not needing to spike anything directly into a sample to boost numbers. Relatedly, we did run a dataset processing with 20-, 50-, and 100- PBMC carriers included (data not shown). However, we found that this actually did not increase the numbers. Perhaps because they were heterogenous carriers, and we did not pre-sort our cells to generate homogenous carriers? To avoid speculation, we left the carriers out to avoid complicating the story any further. Either way, we do not believe a rogue GODZILLA 20 µm PBMC could have inflated numbers across several analyses it was not even a computational part of.

    And now where we have to be a bit serous – the data availability point. We apologize if the data deposition structure is not clear. It is not always easy to manage submissions with 1000s of RAW files and dozens of computational searches, and we cannot readily go back and edit existing PXD IDs with updated files. We would love to have some form of FTP access where we can readily change, edit, and update everything coherently – but this is not currently available.

    We actually went through great (multi-day) efforts to manually parse every single LC/MS instrument setting on a per-RAW-file basis, and included this in “DataSet_to_rawfile_mapping_methods.txt”. Beyond that, the separate “DataSet_to_figure_mapping.txt” file specifies which RAW files correspond to which data search and experiment, and which figure panels they ultimately support. We went for this two-file approach because it provides some resilience towards having to re-order all the figures during revision. These files contain considerably more LC/MS-specific information than the automatically generated SDRF, and allow the reader to piece together which files were processed together, using which software, and how they relate to the figures.

    As for the PBMC data – it is all there on the primary PXD upload; PXD070599. Typing “PBMC” into the search box brings up all single-PBMC RAW files included into the manuscript. For sanity (and sorting) reasons, we recommend using an FTP client to access it.

    One final note on SDRF – it seems that PRIDE has started experimenting with automatically generated SDRF files, which apparently also affected our dataset. The resulting SDRF contains only a small subset of the RAW files and has substantial incorrect annotations, including incorrect cell-type assignments. We have added a note to the metadata to tell people to ignore that SDRF and point them to our own comprehensive guide to the deposited data.

    Please reach out to us if anything is unclear, we are very happy to clarify. :)

    And thanks again for giving us the spotlight!

    Ivo, Sara, and Jesper

    ReplyDelete