When the cell lines are not the patients
How do you find synthetic lethal targets for patient populations that cancer cell line panels do not represent?
Cell line panels do not reproduce the genotype and disease-stage distribution of the patients they are used to model, and in myeloid malignancy the gap is severe enough to matter for target discovery. What we imagine could fill the gaps DepMap leaves: predict synthetic lethal pairs with a model whose knowledge comes from the literature, then rank those predictions against clinical sequencing and normal haematopoiesis instead of against cultured lines.
Key takeaways
- Leukaemia and lymphoma cell lines skew to late disease: 54 percent came from relapsed, refractory or terminal patients.
- DepMap holds essentially one NPM1-mutant AML line, for a genotype carried by roughly 30 percent of patients.
- Lines used as MDS models, including SKM-1 and MOLM-13, were established after transformation to acute leukaemia.
- Synthetic lethality is a claim about co-occurring events, so an unrepresentative panel cannot surface the clinically relevant pair.
- Zero-shot language models reached 0.715 to 0.768 AUROC where two published SL methods scored near chance.
- Rank candidates on patient cohorts and normal haematopoiesis rather than on which tumours grow in plastic.
- Prediction is triage, not a result: it narrows the search space and the wet lab decides what is real.
Synthetic lethality is a claim about co-occurrence, which makes the sample matter
Synthetic lethality
Two genes are synthetically lethal when losing either one alone is survivable but losing both kills the cell. In oncology that asymmetry is the therapeutic window: if the tumour has already lost the first gene, a drug against the second should kill the tumour while sparing healthy cells that still carry both.
DepMap is the most useful resource in target discovery and also a non-random sample of human cancer. Cell lines are the tumours that agreed to grow in plastic, and that is not a lottery. For many purposes the bias is tolerable. For synthetic lethality it is structural, because the whole claim rests on two events occurring together. If the co-occurrence pattern in the panel differs from the pattern in patients, the screen cannot see the interaction that matters clinically. It is not that the screen gets the answer wrong. It is that the question never appears.
The scale of the skew is documented. In a survey of 554 published leukaemia and lymphoma cell lines, Drexler and Quentmeier (2020) found that 54 percent were established from patients at relapse or at a refractory or terminal stage, against 34 percent at diagnosis or initial presentation. Of the ten bona fide Hodgkin lymphoma lines, all ten came from terminal or refractory patients and seven of ten from pleural effusions. The panel is not a cross-section of the disease. It is a cross-section of its end stage.
What is documented about those hits is uncomfortable in its own right. Ku et al. (2020) found that published SL screen hits overlap at the pathway but not the gene level, and that most SL phenotypes are strongly modulated by cellular and genetic context, offering evidence for why most reported synthetic lethals are not reproducible. Lin et al. (2019) knocked out the putative targets of ten clinical-stage cancer drugs and found efficacy unaffected, meaning the stated target was not the mechanism. That is the honest evidence base: not that cell line SL fails in patients, but that it frequently fails to reproduce at all.
Myeloid disease is the sharp case
NPM1 prevalence from Thiede et al. (2006) and TCGA-LAML (2013); DepMap line count verified against the current DepMap model list and the CCLE mutation profile, September 2026.
One cell line for the largest AML subgroup
NPM1 mutation appears in roughly 27 to 30 percent of adult AML and in 45 to 50 percent of cytogenetically normal AML. DepMap currently lists 52 AML models. Exactly one of them carries the NPM1c exon-12 frameshift: OCI-AML3, ACH-000336. The only other established NPM1c line, IMS-M2, has no DepMap or Cell Model Passport accession at all. Calling that one model is generous, because OCI-AML3 also carries DNMT3A R882C and NRAS Q61L, so any dependency found in it is confounded by two further driver lesions. Why NPM1-mutant AML resists immortalisation is not fully resolved and we will not invent a mechanism, but the consequence is not in doubt: the experimental literature for the single largest molecularly defined AML subgroup rests on two cell lines, one of which most screens never touch.
MDS is missing its own biology
SKM-1 was derived in 1989 from a 76-year-old whose MDS had already progressed to monoblastic leukaemia. MOLM-13 came from a 20-year-old with AML M5a at relapse, arising from antecedent MDS. Both are real, useful, well-characterised lines. Neither reports the ineffective haematopoiesis that defines primary MDS. A 2026 review of MDS models puts it directly: lines frequently used in MDS-oriented studies were derived after leukaemic transformation and therefore reflect clonal evolution and disease progression rather than the biology that defines the disease. MDS progenitors proliferate poorly in vitro and depend on stromal support, which is why the gap has persisted.
For a disease defined by what happens before transformation, that is the wrong tissue.
Why this is worth solving now
Myeloid care has moved from chemotherapy-only to mutation- and risk-stratified treatment, with hypomethylating agent plus targeted agent combinations at the centre: in VIALE-A, azacitidine plus venetoclax extended median overall survival to 14.7 months against 9.6. Approvals now span FLT3 (midostaurin, gilteritinib, quizartinib), IDH2 (enasidenib), IDH1 (ivosidenib, olutasidenib), BCL-2 (venetoclax) and menin (revumenib in 2024 and 2025, ziftomenib in 2025). Almost all of these are AML labels; ivosidenib is the one carrying an MDS indication. The indication is demonstrably tractable for precision therapy. What is scarce is not chemistry but targets, particularly for resistant populations whose options are exhausted.
The asymmetry worth exploiting
Many myeloid subgroups are defined by a gene that is already functionally compromised, whether mutated, epigenetically silenced, or expressed at critically low levels. Those loss-of-function events characterise specific patient populations and frequently drive treatment resistance. That gives a synthetic lethal search an unusual starting advantage: one half of the pair is already a validated biomarker with a diagnosed population attached.
The first gene is an existing, population-defining vulnerability: already characterised, already diagnosed, already used to stratify patients. The second gene, whose inhibition is selectively lethal in that background but tolerated in normal haematopoietic progenitors, is the novel target. One half of the pair is de-risked by clinical reality; the other half is where the discovery is.
A three-layer approach
Layer 1: prediction that never saw a cell line
In our preprint (Prosz et al., bioRxiv 2026; not yet peer reviewed), open-weight language models are asked to estimate the likelihood of a synthetic lethal interaction for a gene pair with no fine-tuning and no training set. The models reason from biological knowledge encoded during pretraining, and every prediction carries a written mechanistic chain a scientist can reject on its merits.
Multiple open-weight models were tested with a uniform prompt against known screens, then used to predict 398,277 clinically relevant gene pairs. From Prosz et al., bioRxiv 2026.
The benchmark is Olivieri et al. (Cell 2020), a map of the DNA damage response built from 31 CRISPR screens against 27 genotoxic agents in RPE1 cells. Two things follow from that, and both belong in the open. RPE1 is a non-cancer, near-diploid line, which makes it a genuinely independent benchmark precisely because it sits outside the cancer dependency panel. It also means this benchmark does not by itself demonstrate performance on myeloid biology.
Performance saturates around 0.73 AUC across the larger models, while both published comparators sit near chance. From Prosz et al., bioRxiv 2026.
Performance saturates around 0.73 AUC: gpt-oss-20B at 0.768, Qwen2.5-72B at 0.735, gpt-oss-120B at 0.716 and Qwen2.5-32B at 0.715, with Qwen2.5-32B chosen for the production screen on performance-to-cost grounds. The two published comparators, SLant at 0.501 and MAGICAL at 0.549, sit at or near chance on the same data. Validation against a dataset released after the training cutoff argues the result is not leakage. 398,277 pairs have been screened, with predictions public at github.com/Paureel/LLMsynthlet.
The claim here is narrow. These priors come from the literature rather than from a panel of cultured lines, so the predictions are not conditioned on which tumours grow in plastic. That is not automatically better. It is differently biased, and the biases are stated below.
Layer 2: clinical overlay, where the ranking is decided
Raw predictions are filtered and re-ranked against patient sequencing rather than cell lines: TCGA-LAML (200 de novo adult AML), Beat AML (672 specimens from 562 patients), the Papaemmanuil et al. MDS cohort (738 patients, Blood 2013) and Haferlach et al. (944 MDS patients, Leukemia 2014). Three questions decide whether a predicted pair survives.
| Question | Why it decides the ranking |
|---|---|
| Is the first gene genuinely lost, mutated or under-expressed in this population? | A vulnerability that is rare in patients cannot anchor a stratified trial, however strong the prediction |
| Is the second gene expressed and druggable in the same disease context? | A target that is not present in the tissue is not a target |
| Is it spared in normal haematopoietic progenitors? | In myeloid disease, toxicity to normal blood production is the primary safety constraint, not a later question |
Worked subgroups would be chosen with the partner. Two obvious candidates: TP53-mutant disease, at 5 to 10 percent of de novo MDS and AML, 30 to 40 percent of therapy-related cases and 70 to 80 percent of complex karyotype (where, per Bernard et al. 2020, only multi-hit status carries the adverse phenotype); and splicing-factor-mutant disease, where spliceosome mutations affect roughly 45 to 50 percent of MDS overall, with SF3B1 at 20 to 30 percent, SRSF2 at 10 to 15 percent and U2AF1 at 8 to 11 percent.
Layer 3: agentic evidence assembly, explicitly not validation
For each surviving candidate, an agent assembles a structured dossier from UniProt, STRING, Reactome, DrugBank and ClinicalTrials.gov: mechanistic rationale, existing tool compounds, active clinical programmes, known safety signals, and an audit trail. The boundary matters and we would rather overstate it: this is annotation, not evidence that a prediction is true. Its purpose is rejection speed, letting a scientist dismiss a candidate in ten minutes rather than ten days. In the preprint, this filtering scored the top 100 predicted pairs for novelty and feasibility and surfaced 11 high-novelty, high-feasibility candidates in the RPE-1 context. The failure modes such a layer must avoid are ones we have written about at length.
How you would know it works before spending wet-lab budget
The first deliverable is not a target. It is a calibration.
| Control | What it tests | Target |
|---|---|---|
| Positive: experimentally confirmed SL pairs from CRISPR screens, SynLethDB and published functional studies | Does the method recover what is already known? | recover at least 70 percent |
| Negative: known non-interactors, randomly sampled pairs, established passenger mutations | Does it reject what should be null? | no more than 30 percent false positives |
| White space: high-confidence pairs absent from existing screens | Is there anything here nobody has tested? | the actual output |
Precision, recall, F1 and AUROC get reported openly, so a partner can decide how much to trust the novel predictions from how well the method does on known biology. There is a tension in that design worth naming rather than burying: benchmarking against known SL pairs measures performance where cell line data exists, which is precisely the regime we argue is unrepresentative. It is the best available check and it is not a complete one.
Limits, stated plainly
- Literature-derived priors inherit the literature's biases: well-studied genes are over-represented, and absence of evidence reads as absence of interaction.
- Prediction is triage, never proof. Nothing here substitutes for a screen in a relevant model or a patient sample.
- The benchmark showing 0.715 to 0.768 AUROC was run in a non-cancer line against genotoxic agents, not in myeloid disease.
- Clinical cohorts carry their own ascertainment biases and mostly lack matched functional data.
- Normal-tissue tolerance is the hardest question in the whole design and the least well served by public data.
Why this is worth running
This generalises past myeloid disease to any indication where the panel does not look like the patients. We have run the analogous prioritisation problem before: our combinatorial therapy case study covers the same arithmetic at a different scale, where two genes give roughly 200 million candidate pairs and three give about 1.3 trillion, and ranking before testing is the only tractable move.
If your group works in myeloid malignancy, or anywhere the cell line panel misrepresents your patients, the interesting conversation is which subgroup to start with and which positive controls would convince you. That is a conversation we would like to have.
Which subgroup would you start with?
We are looking for partners in myeloid malignancy and in other indications where the cell line panel does not represent the patients.
Frequently asked questions
What is synthetic lethality in cancer drug discovery?
Two genes are synthetically lethal when losing either one alone is survivable but losing both kills the cell. In oncology this creates a therapeutic window: if a tumour has already lost the first gene, a drug against the second should kill the tumour while sparing healthy cells that retain both.
Why are cell line panels a problem for myeloid malignancies?
Because they under-represent the patients. Across 554 published leukaemia and lymphoma lines, 54 percent came from relapsed, refractory or terminal disease. DepMap contains essentially one NPM1-mutant AML line despite that genotype appearing in roughly 30 percent of patients, and lines used as MDS models were derived after leukaemic transformation.
Can a language model really predict synthetic lethal interactions?
To a useful degree, for triage. In our preprint, open-weight models scored 0.715 to 0.768 AUROC on an independent benchmark, while two published SL prediction methods scored 0.501 and 0.549, essentially chance. That is a ranking signal for prioritisation, not a diagnostic, and it is not a substitute for a screen.
How would you validate predictions without a cell line for the disease?
With positive controls (recovery of experimentally confirmed SL pairs), negative controls (known non-interactors and passenger mutations), and clinical overlay against patient sequencing cohorts. The honest limit: benchmarking against known pairs measures performance where cell line data exists, which is the regime we argue is unrepresentative.
What would a partner actually receive?
A narrowed list: three to five high-confidence gene pairs, each with a translational assessment, a competitive landscape, a matched patient population profile, and a readable mechanistic chain showing why the model made the call. Enough to decide where to spend experimental budget, not a claim that the biology is settled.
Sources
- 1.Prosz, A., Sztupinszki, Z., Diossy, M., Kilim, O., Zimon, B., Csabai, I., & Szallasi, Z. (2026). Zero-shot biological reasoning with open-weights large language models reproduces CRISPR screen based prediction of synthetic lethal interactions. bioRxiv preprint. Not peer reviewed.
- 2.Drexler, H. G., & Quentmeier, H. (2020). The LL-100 Cell Lines Panel: Tool for Molecular Leukemia-Lymphoma Research. Int J Mol Sci 21(16), 5800.
- 3.Ku, A. A., et al. (2020). Integration of multiple biological contexts reveals principles of synthetic lethality that affect reproducibility. Nat Commun 11, 2375.
- 4.Lin, A., Giuliano, C. J., Palladino, A., et al. (2019). Off-target toxicity is a common mechanism of action of cancer drugs undergoing clinical trials. Sci Transl Med 11(509), eaaw8412.
- 5.Olivieri, M., Cho, T., Alvarez-Quilon, A., et al. (2020). A Genetic Map of the Response to DNA Damage in Human Cells. Cell 182(2), 481-496.e21.
- 6.Benstead-Hume, G., et al. (2019). Predicting synthetic lethal interactions using conserved patterns in protein interaction networks (SLant). PLoS Comput Biol15(4), e1006888. Dey, A., Mudunuri, S., & Kiran, M. (2024). MAGICAL. PLoS Comput Biol 20(8), e1012336.
- 7.Cohorts: Ley, T. J., et al. (2013) TCGA-LAML, N Engl J Med 368, 2059-74; Tyner, J. W., et al. (2018) Beat AML, Nature 562, 526-31; Papaemmanuil, E., et al. (2013), Blood 122(22), 3616-27; Haferlach, T., et al. (2014), Leukemia 28(2), 241-7.
- 8.Myeloid context: Thiede, C., et al. (2006) Blood 107, 4011-20 (NPM1 prevalence); Daver, N. G., et al. (2022) Cancer Discov 12, 2516-29 and Bernard, E., et al. (2020) Nat Med 26, 1549-56 (TP53); DiNardo, C. D., et al. (2020) VIALE-A, N Engl J Med 383, 617-29; Maslinska-Gromadka, K., et al. (2026) Int J Mol Sci 27(2), 898 (MDS cell line models).