Target discovery

Inside-out proteins: a new source of cancer surface targets, and how AI could find more

If a protein has no signal peptide and no transmembrane domain, how does it end up on the outside of a cancer cell?

By at least two routes, both published in 2026, both independent of the canonical secretory pathway. Slezak et al. mapped roughly 140 such proteins held on the surface by heparan sulfate; Delaveris et al. showed Src kinase inverted onto it by autophagolysosomal exocytosis. The class is defined by what it lacks, which is precisely why every standard surfaceome predictor is blind to it.

By PharosBioPublished on 9 min read

Key takeaways

  • Inside-out proteins are intracellular proteins displayed on the outer surface of stressed or malignant cells.
  • Slezak et al. (2026) mapped roughly 140 high-confidence I-O proteins and validated 40 with antibodies.
  • Delaveris et al. (2026) showed Src kinase is inverted onto the surface by autophagolysosomal exocytosis.
  • The defining feature is an absence: no signal peptide, no transmembrane domain.
  • That absence is why standard surfaceome predictors cannot see this class at all.
  • Display is reversible and stress-induced, so normal cells can show it too under the right insult.
  • Binding a soluble antigen predicts surface binding poorly, which breaks the usual antibody discovery route.
  • A useful classifier has to encode mechanism, not just sequence, and be judged on what it gets wrong.

The surfaceome has been drawn from the wrong list

Almost every antibody therapeutic acts on a protein at the cell surface, and the list of candidate surface proteins has been assembled the same way for decades: look for a signal peptide, look for a transmembrane domain, and the protein that has one is on the list. It is a good rule. It is also the reason an entire category of targets has been sitting in plain sight, filed as experimental artefact.

Intracellular proteins turning up on the outside of cells has been reported for years, usually under the heading of moonlighting proteins, and usually treated with suspicion. Cell surface Hsp70, extracellular nucleolin, surface enolase: each individually documented, each easy to dismiss as contamination from lysed cells. What changed in 2026 is that two groups attacked the problem systematically, with methods designed to exclude exactly that artefact, and both found the phenomenon is real and larger than anyone was treating it as.

Definition

Inside-out (I-O) proteins are intracellular proteins displayed on the outer leaflet of the plasma membrane under cellular stress or malignancy, without the signal peptide or transmembrane domain that normally routes a protein there. They are not secreted and not inserted into the bilayer. They are held on the outside by other means, and they come back when you strip them off.

Two mechanisms, found independently

Note what holds each protein in place. On the left, nothing in the bilayer at all: the protein binds sugar chains that extend from a proteoglycan. On the right, a lipid tail in one leaflet. Neither is a transmembrane domain.

The tethered route

Slezak and colleagues used APEX2-mediated proximity biotinylation, which labels proteins within a very short radius of a membrane-impermeable probe, paired with a custom antibody generation platform. They identified roughly 140 high-confidence I-O proteins: ribosomal subunits, proteasomal components, heat shock proteins, translation machinery. Not a random slice of the proteome, but a coherent set enriched for stress-response families.

The mechanism they report is the interesting part. These proteins are not inserted into the membrane. They are tethered to it by heparan sulfate glycosaminoglycans, and heparinase II treatment eliminated virtually all of the surveyed I-O proteins from the surface. High salt did almost nothing, and removing N-linked glycans did nothing, so the interaction is specific rather than generically ionic. Then the detail that makes it a system rather than a leak: the stripped proteins returned to baseline within six hours, and only while ER-Golgi trafficking remained intact. This is a maintained equilibrium, not debris.

The inverted route

Delaveris and colleagues, working in James Wells’s lab at UCSF, came at it from a different direction with photo-proximity labeling and found something more specific and, for a drug developer, more striking. Src kinase, the proto-oncogene anchored to the inner leaflet by N-terminal myristoylation, appears on the outer surface of cancer cells, still catalytically active.

The route is autophagolysosomal exocytosis, and the paper is explicit that Src is the prototype of a family of membrane-anchored proteins moving this way rather than a one-off. It is worth being precise about the difference between the two mechanisms here, because it is easy to blur. Src is membrane-associated, but through a lipid tail buried in a single leaflet, not a domain crossing the bilayer. The proteins Slezak describes are not membrane-associated at all: they sit above the bilayer, bound to glycan chains. Two different answers to the question of what holds a protein on the outside, and neither is the one a surfaceome predictor looks for. They found extracellular Src in primary patient tumours and showed anti-Src antibody therapies killing tumour cells in culture and in mouse xenografts. Two different anchoring chemistries, two different trafficking routes, the same net result.

Why this class is invisible to the usual tools

The defining property of an I-O protein is a negative. No signal peptide. No transmembrane domain. Every computational surfaceome predictor works by detecting the positive version of exactly those features, so the entire class is not merely ranked low, it is structurally outside what the tool can express.

Transcriptomics does not rescue it either. An mRNA count tells you the protein is being made; it cannot tell you which side of the membrane it ended up on. For a target class whose whole value proposition is a change in localisation rather than a change in abundance, expression data is close to uninformative. That is a real limitation for anyone hoping to nominate these targets from public expression atlases.

There is also an uncomfortable finding for antibody discovery. Slezak and colleagues report only a weak correlation between a Fab binding the soluble antigen and the same Fab recognising it on the cell surface: for many targets, fewer than half of the Fabs that bound the purified protein showed detectable surface binding. The presentation geometry matters. A campaign run against recombinant antigen may select binders that do not work where it counts, which changes how you would design the discovery cascade for this class.

The caveat worth stating loudly

It would be easy to write these up as tumour-specific antigens and stop. The data support something more careful.

I-O proteins were undetectable on resting peripheral blood mononuclear cells and on normal tissue, which is the headline. But when Slezak and colleagues stressed normal PBMCs with GW4869, an inhibitor of neutral sphingomyelinase, the cells displayed I-O proteins measurably, and reverted to baseline within 24 hours of removing the stimulus. The authors are direct about the implication: surface display is a reversible, stress-dependent event.

What this means for a target list

Selectivity here is a stress differential, not an absence. The therapeutic window depends on tumour cells being persistently stressed while normal tissue is not, which is a claim about the patient and the indication rather than about the protein. Any programme built on this class has to characterise that window in the specific setting it intends to treat.

One member of the set has travelled furthest. Enolase-1 (ENO1) is a glycolytic enzyme with a well-documented surface pool, and antibody programmes against surface ENO1 have reached preclinical development, including work on reprogramming macrophage polarisation and inhibiting tumour growth. It is a useful proof that the category can yield a real drug programme, and a useful reminder that it has taken years per target by conventional means.

Where AI could actually contribute

The obvious move is to train a classifier that separates I-O proteins from the rest of the proteome and use it to nominate candidates for validation. The obvious move is also where most of these efforts fail, because a sequence model has no way to represent a tethering mechanism. The features have to carry the biology.

Sequence embeddings are the base layer, not the whole model. The mechanistic features are what make the predictions falsifiable for a reason.

Three of those features deserve comment. The electrostatics term exists because Slezak identified heparan sulfate as the tether, so heparin-binding character and surface charge computed from an AlphaFold model are direct proxies for the mechanism rather than generic descriptors. The lipidation term exists because Delaveris identified myristoylation as the anchor on the second route, and the paper frames Src as one of a family. And the negative feature is doing more work than it looks: encoding the absence of a signal peptide as an explicit input is what stops the model rediscovering the conventional surfaceome and calling it a result.

The harder problem is the training set. Roughly 140 validated positives is small, the negative class is the rest of the proteome and is certainly contaminated with undiscovered positives, and the label itself is assay-dependent. Those are real constraints, and a model built on them should be treated as a hypothesis generator rather than an oracle. That is why the loop matters more than the architecture: predict, rank by what is cheapest to falsify, test at the bench, and feed the graded result back, including the failures, which are the most informative data the loop produces.

This is also where the honest limits sit. A classifier can propose which proteins might reach the surface. It cannot tell you the antigen density, whether the epitope is accessible in its tethered conformation, whether the target internalises, or which modality suits it. On that last point the two papers already disagree in a useful way: for extracellular Src the reported activity favoured T-cell engager and radioligand approaches over an antibody-drug conjugate, which is exactly the kind of decision that depends on internalisation behaviour a sequence model knows nothing about.

What we would build

Our interest here is not the classifier in isolation. It is that this target class has the shape of a problem where an autonomous analysis layer earns its cost: a small validated positive set, a large candidate space, several orthogonal evidence types that have to be joined per protein, and a validation step that is slow enough to make prioritisation valuable.

Concretely, four components. Select the indications where the stress differential plausibly gives a therapeutic window, since that is the question the biology actually poses. Predict candidates with the feature stack above, stress-tested against the validated set rather than a random holdout. Assign each candidate a modality from its druggability profile, because internalisation behaviour decides between an ADC, an engager and a radioligand. Then close the loop with bench validation, so each round expands the known set instead of re-ranking the same one.

We have run versions of this shape before. The EZH2 work with the Danish Cancer Institute surfaced a driver that multiple-testing correction would have buried, using an interpretable network over four data modalities. The FRα ADC comparison is the modality-assignment problem in miniature. Neither is this problem, but both are the same discipline: make the prediction specific enough to be wrong, then go and find out.

The open questions

Being clear about what is unresolved is more useful than enthusiasm here, because this is a young field and several of these questions are load-bearing for any programme.

Open questionWhy it decides a programme
Conformation and tetheringA protein held by heparan sulfate may present a different epitope surface than the soluble form used to raise the antibody
Antigen densitySets the floor for every modality, and is the parameter most likely to differ between a cell line and a patient tumour
InternalisationDecides whether an ADC is viable at all, or whether an engager or radioligand is the right modality
Normal-tissue display under stressThe therapeutic window is a stress differential, so the relevant comparison is stressed normal tissue, not resting tissue
Which indications benefitNo characterised I-O repertoire exists for most tumour types, so indication selection is currently guesswork
Surface versus intracellular discriminationTranscriptional profiling cannot separate the two, which limits how far public expression data can take you

Glossary

TermWhat it means
Inside-out (I-O) proteinAn intracellular protein displayed on the outer leaflet of the plasma membrane without canonical targeting signals
SurfaceomeThe set of proteins present at the cell surface, and the candidate pool most antibody therapeutics are drawn from
Signal peptideThe N-terminal sequence that routes a protein into the secretory pathway; absent in this class
APEX2 proximity biotinylationAn engineered peroxidase that tags proteins within a short radius, used with a membrane-impermeable probe to label only the outside
Photo-proximity labelingLight-triggered labeling of proteins near a catalyst, giving high spatial resolution on the cell surface
Heparan sulfateA glycosaminoglycan on the cell surface; the tether holding many I-O proteins in place
N-myristoylationAttachment of a fatty acid to a protein's N-terminus, normally anchoring it to the inner leaflet
Autophagolysosomal exocytosisA secretory route in which autophagolysosome contents are released, inverting anchored proteins onto the outer surface
Stress granuleA cytoplasmic condensate formed under stress; roughly half the known I-O set overlaps its components
ESM embeddingA vector representation of a protein sequence from a protein language model, used as the base features for a classifier
Antigen densityHow many copies of a target sit on each cell, which sets the floor for what any modality can achieve
Radioligand therapyA targeting molecule carrying a radionuclide, which does not require internalisation to kill

Frequently asked questions

What are inside-out proteins?

Inside-out proteins are intracellular proteins that appear on the outer surface of the plasma membrane under stress or malignancy, without the signal peptide or transmembrane domain that normally routes a protein there. Slezak and colleagues identified roughly 140 high-confidence examples, mostly ribosomal, proteasomal, chaperone and translation factors.

How do intracellular proteins reach the cell surface without a signal peptide?

By at least two documented routes. Slezak et al. report ER-Golgi dependent trafficking with the protein tethered on the outside by heparan sulfate glycosaminoglycans rather than inserted in the bilayer. Delaveris et al. report autophagolysosomal exocytosis inverting N-myristoylated proteins such as Src onto the outer surface.

Are inside-out proteins actually tumour-specific?

Selective rather than exclusive. They were undetectable on resting peripheral blood mononuclear cells and normal tissue. But stressing normal PBMCs with GW4869 induced surface display, which reverted within 24 hours. Selectivity therefore depends on the stress differential between tumour and normal tissue, not on the protein being absent elsewhere.

Why do standard surfaceome predictors miss these proteins?

Because they screen for a positive signal, a signal peptide or a transmembrane domain, and this class is defined by lacking both. Transcriptional profiling does not help either: mRNA cannot distinguish a protein sitting on the surface from the same protein doing its normal job inside the cell.

Can machine learning predict new inside-out proteins?

Plausibly, but only if the features encode the mechanism. A sequence model alone has no way to represent heparan sulfate tethering or myristoylation anchoring. With roughly 140 validated positives the training set is also small, so the honest design is a falsification loop where bench results, including failures, feed back.

Hydra

Point it at a target class nobody has mapped

Hydra plans, runs and validates real bioinformatics analysis across roughly 100 preloaded scientific databases and 200+ codified skills. Bring a question this shaped and see what it returns.

Sources

  1. 01Slezak T, O’Leary KM, Guevara Avella T, Musial N, Li J, Andrzejczak A, Scott EF, Le DA, Kossiakoff AA (2026). Dynamic translocation of Inside-Out proteins to the cell surface underlies cellular adaptation to cancer-induced stress. PNAS 123(13): e2529493123. doi.org/10.1073/pnas.2529493123 Source for the ~140 protein set, the 40 antibody-validated targets, heparan sulfate tethering, the six-hour recovery, the PBMC and GW4869 results, and the soluble-versus-surface binding discrepancy.
  2. 02Delaveris CS, Loudermilk RP, Pandey A, et al., Leung KK, Wells JA (2026). Autophagolysosomal exocytosis inverts Src kinase onto the cell surface in cancer. Science 391(6790): eaec1778. doi.org/10.1126/science.aec1778 Source for extracellular Src, the ALE mechanism, primary tumour detection and the antibody therapy results.
  3. 03Modality preference for extracellular Src, and the observation that not every antibody raised against the soluble antigen recognises the surface form, reflect the author’s reading of the two papers together with correspondence with the authors. Treat the modality point as a working interpretation rather than a published head-to-head comparison.
  4. 04The classifier design described here is a proposal, not a reported result. No performance figures are claimed because none have been measured.

Related reading