§ Claim under review · Research
"Researchers used an AI system to identify combinations of existing approved medicines (originally developed for conditions like high cholesterol and alcohol dependence) that show promising anticancer activity against breast cancer, with several combinations outperforming standard breast cancer drugs in lab tests (DOI: 10.1098/rsif.2024.0674)"
Verdict
Partially accurate but misleading
Confidence
HighSummary
The study is real. Researchers led from the University of Cambridge used GPT-4 to suggest pairs of cheap approved drugs, including the cholesterol drug simvastatin and the alcohol dependence drug disulfiram, and tested them on breast cancer cells in a dish. Three of the first 12 pairs scored higher than the study's approved breast cancer drug comparison pairs, and a second AI round produced three more with positive scores, which is what the paper and the university press release say. The misleading part is the word "outperforming." The score being measured only asks whether a pair works better than its own stronger ingredient, and the paper's own figures show one standard chemotherapy drug, doxorubicin, killed the cancer cells at a far lower dose than anything the AI proposed. The researchers themselves describe the results as additive rather than strongly synergistic, and the work covers one cancer cell line in culture with no animal or human testing. It is a genuine and interesting demonstration of AI generating testable ideas, but it is not evidence that these cheap drug pairs beat cancer treatments.
The readings
key figures from the evidenceHSA synergy score, simvastatin plus disulfiram, first screen
doxorubicin alone IC50 in MCF7, more potent than AI-proposed pairs
Why this verdict
Evidence
The DOI resolves to a real, open-access, peer-reviewed paper in the Journal of the Royal Society Interface, published 4 June 2025 (accepted 28 April 2025), led from the University of Cambridge with co-authors at King's College London, Goldsmiths and Arctoris Ltd.
The method: the authors prompted GPT-4 to propose pairs of FDA-approved non-cancer drugs expected to be toxic to the MCF7 breast cancer cell line but not to the non-tumorigenic breast line MCF10A, with instructions to avoid standard cancer drugs and favour cheap, approved drugs. GPT-4 also supplied two positive control pairs of approved breast cancer drugs (doxorubicin plus cyclophosphamide, and fulvestrant plus palbociclib) and two negative control pairs.
The results, in the authors' own words in the abstract: in the first round GPT-4 produced three drug combinations out of 12 tested "with synergy scores above the positive controls," and a second round seeded with those results produced three more combinations with positive synergy scores out of four tested. The three first-round winners were itraconazole plus atenolol, disulfiram plus simvastatin, and dipyridamole plus mebendazole. Simvastatin treats high cholesterol and disulfiram treats alcohol dependence, so those two drug classes named in the claim are correct.
The measured quantity is an HSA synergy score computed with SynergyFinder 3.0. Synergy here is defined in the paper as the combination doing better than its own best single agent, not better than a different drug. The whole-matrix HSA scores were small: simvastatin plus disulfiram scored 3.29 and itraconazole plus atenolol 4.83 in the first screen. The earlier preprint version states directly that no combination showed high synergy, using the conventional threshold of a score above 10. The paper describes these as additive combinations with positive synergy scores. Only in the second screen, and only within the single most synergistic 3 by 3 dose window rather than across the whole matrix, did one pair (disulfiram plus simvastatin) exceed a score of 10.
The paper's own single-drug table shows the standard cancer drugs behaving poorly in this assay: fulvestrant and cyclophosphamide both had IC50 values above the 25 µM maximum dose in both cell lines, while doxorubicin alone had an IC50 of 0.303 µM in MCF7, more potent than anything the model proposed.
The Cambridge press release compresses this to the line that three of the 12 combinations "worked better than current breast cancer drugs," which is the wording the social post follows.
Findings
✓ What's accurate 6
- The DOI is real and points to a genuine peer-reviewed paper that matches the description. Nothing here is fabricated.
- An AI system, specifically GPT-4, was used to generate the drug-pair hypotheses, and human scientists tested them in the laboratory.
- The drugs include medicines approved for other conditions, including simvastatin for high cholesterol and disulfiram for alcohol dependence, exactly as the claim says.
- The paper's abstract itself states that three of 12 first-round combinations had synergy scores above the positive controls, and that a second round produced three more with positive synergy scores out of four tested.
- The positive controls were combinations of drugs approved for breast cancer, so "standard breast cancer drugs" is a fair description of what the comparison was against.
- The iterative loop described in the post, results fed back to the model to generate new combinations, is what the paper reports.
≈ What's misleading 8
- The claim says the combinations showed "promising anticancer activity" and "outperformed" standard breast cancer drugs. The measured quantity was a synergy score, which asks only whether a pair beats its own stronger ingredient. It is not a measure of how well a treatment kills cancer cells relative to another treatment. By the paper's own single-drug data, doxorubicin alone killed MCF7 cells at an IC50 of 0.303 µM, more potent than any of the proposed pairs, so a reader who concludes these cheap pairs killed breast cancer cells better than chemotherapy drugs has been led to the wrong conclusion by a proxy metric.
- The word "outperforming" hides that the control combinations were barely active in this assay. The paper reports fulvestrant and cyclophosphamide IC50 values above the maximum 25 µM dose in both cell lines. Exceeding a near-flat comparator is a much weaker finding than the phrasing suggests.
- The scores themselves were in the range the field treats as additive rather than synergistic. The preprint version of this same work states plainly that none of the first-round combinations reached the conventional synergy threshold of above 10, and the published paper calls them additive combinations. Only one pair crossed 10, and only inside the single best 3 by 3 dose window of the second screen rather than across the full dose matrix.
- "in lab tests" does signal preclinical work, but the specific scope is narrower than most readers will assume. This was one breast cancer cell line, MCF7, compared against one normal-like breast line, in culture. There were no animal experiments and no patients, and no follow-up in vivo work was found.
- The post caption says scientists asked the AI to "examine thousands of existing medicines." The paper describes prompting GPT-4 to propose drug pairs from what it had absorbed in training. It is not a screen of thousands of catalogued medicines.
- The caption says three of the four second-round combinations "showed strong anticancer activity." The paper says those three had positive synergy scores. Positive is not strong, and the paper does not use that language.
- One of the three second-round positives was disulfiram plus fulvestrant, and fulvestrant is an approved breast cancer drug. The framing that all the successful combinations were medicines originally developed for other conditions does not hold for the second round.
- The post's framing of hidden connections humans would overlook applies to the specific pairings, which the authors report were not found in the prior breast cancer literature. The individual drugs are not obscure choices. Disulfiram in particular has been an actively studied cancer repurposing candidate for years, with its own systematic reviews and clinical trials, and the authors themselves note several of the individual drugs had already been tested against MCF7.
? What's uncertain 5
- Whether any of these effects would appear in other breast cancer cell lines, in animals, or in people. The study did not test this and no follow-up experimental work was found as of 2026-10-11.
- Whether the drug concentrations used in culture are achievable and safe in patients. The paper's assay used doses up to 25 µM, and clinical attainability is not addressed by this evidence.
- Whether an independent laboratory can reproduce the synergy scores. No independent replication was found.
- How much of the apparent success reflects the model's hypotheses versus the choice of a weak in vitro comparator, since one control component, cyclophosphamide, is a prodrug and the paper records it as inactive at the doses tested.
- Whether the authors' description of this as the first closed-loop system of its kind holds against all prior work. That is a priority claim by the authors and was not independently checked here.
Sources
8 of 8 linked to recordsAbdel-Rehim A, Zenil H, Orhobor O, Fisher M, Collins RJ, Bourne E, Fearnley GW, Tate E, Smith HX, Soldatova LN, King RD. "Scientific hypothesis generation by large language models: laboratory validation in breast cancer treatment." J. R. Soc. Interface 22(227): 20240674, published 4 June 2025, open access
Green open-access PDF of the same accepted article, Goldsmiths Research Online
arXiv preprint 2405.12258, 20 May 2024, earlier version of the same work
University of Cambridge press release, "'AI scientist' suggests combinations of widely available non-cancer drugs can kill cancer cells"
EurekAlert syndication of the Cambridge release
Society of Chemical Industry news item quoting Prof Ross King
Oxford University Research Archive record confirming journal, volume, issue, article number and dates
Background on disulfiram as a long-studied cancer repurposing candidate, systematic review