Since 2021, AlphaFold has quietly built one of biology’s most useful new public resources: a database that now includes over 17.7 million predicted protein structures spanning more than 2.4 million bacterial and archaeal genomes, freely available to any researcher. Most of the attention that database has received focuses on drug discovery — using predicted structures to design new molecules that bind bacterial proteins. A newer, less publicized use is emerging alongside it: researchers are starting to run these same structural prediction tools forward, not to design a new drug, but to forecast which specific mutations a bacterium is structurally likely to acquire next in its ongoing evolutionary arms race against antibiotics — catching resistance on the whiteboard, so to speak, before it shows up in a patient’s blood culture.
The Scientific Foundation
The foundational insight connecting protein structure to resistance prediction is that not every possible mutation to a bacterial protein is equally likely to succeed. A mutation that would let a bacterium evade an antibiotic only helps that bacterium survive if the mutated protein still folds correctly and still performs its essential biological function — meaning resistance mutations aren’t randomly scattered across a protein’s sequence, they cluster in specific structural locations where a change can alter drug binding without destroying the protein’s core function. A September 2025 preprint applied exactly this logic at scale, mapping mutations from 31,428 clinical Mycobacterium tuberculosis isolates onto AlphaFold-predicted and experimentally determined protein structures, and found clear, unsupervised clustering: the proteins with the highest concentration of mutations were, overwhelmingly, exactly the ones already flagged by the World Health Organization’s official catalogue of resistance-associated mutations, validating that structural context alone can predict which mutations matter clinically.
The genuinely forward-looking version of this idea has already moved from foundational validation into active development. A September 2025 bioRxiv preprint describes a risk-based prediction tool built on protein language models, applied specifically to WHO priority pathogens, including rifampicin-resistant M. tuberculosis and carbapenem-resistant Pseudomonas aeruginosa, that performs in silico deep mutational scanning combined with structural mapping to recover known resistance-associated regions and, notably, highlight new candidate mutations the tool flags as high-risk before they’re clinically observed. The researchers report the tool achieving competitive accuracy, F1, and Matthews correlation coefficient scores around 0.87 to 0.88 on held-out test data. This builds on earlier wet-lab deep mutational scanning work, including a 2023 Nature Communications study that systematically tested over 17,000 variants across three essential E. coli proteins, finding meaningfully different mutational tolerance between targets — a finding the researchers explicitly proposed using to steer future antibiotic development toward protein targets structurally less capable of evolving resistance in the first place.
The Cross-Domain Connection
The genuinely novel synthesis here is running a tool built for one direction of inference in reverse. AlphaFold and its successors were built and are overwhelmingly used to answer “given this protein sequence, what shape does it fold into” — a question central to drug design, since knowing a target’s shape lets chemists design molecules that bind it. The resistance-forecasting application flips that same structural machinery to ask a different, evolutionarily forward-looking question: “given this protein’s structure and its known function, which specific amino acid changes could a bacterium make that would disrupt drug binding while preserving enough of the protein’s shape to keep the bacterium alive.” That’s a subtly different computational problem, closer to a structural stability and functional-tolerance search than a pure folding prediction, and it borrows techniques, protein language models trained on evolutionary sequence data, deep mutational scanning’s systematic variant testing, structural clustering analysis, from adjacent corners of computational biology that don’t always talk to clinical microbiology directly.
What makes this cross-domain rather than merely computational is that it’s fusing three previously separate research traditions: structural biology’s AlphaFold-era prediction tools, evolutionary biology’s understanding of mutational tolerance and epistasis (how one mutation’s effect depends on others already present), and clinical microbiology’s accumulated genomic surveillance data from real patient isolates. A February 2026 review of AlphaFold’s ongoing impact notes that roughly 40 percent of new structures deposited into the Protein Data Bank between 2024 and 2025 relied on AI-driven modeling techniques descended from AlphaFold, underscoring just how much raw structural material now exists for exactly this kind of forward-looking analysis to build on, across far more pathogens than any single wet-lab deep mutational scanning effort could cover directly.
What Remains Undemonstrated
This is squarely an emerging research direction rather than a deployed clinical surveillance tool, and the September 2025 papers describing it are preprints, not yet peer-reviewed and published in final form. The risk-based AMR variant prediction tool’s stated performance metrics come from validation against already-known resistance mutations and structured test splits, not from a genuine prospective test where the tool’s predictions were made before a novel resistance mutation was later observed emerging in real clinical isolates — the single most convincing kind of evidence such a tool could eventually produce, and one that, by definition, takes years of real-world surveillance to accumulate. Deep mutational scanning itself, while systematic, remains labor- and resource-intensive to run experimentally at the scale needed to validate computational predictions across the full range of clinically important bacterial pathogens and antibiotic classes, meaning computational predictions currently outpace the experimental capacity to confirm them.
Why It Matters
Antimicrobial resistance is already a documented, accelerating public health crisis, and the traditional drug-development cycle described in the resistance-prediction literature, a compound is approved, patients are treated, resistance eventually emerges, a new drug enters development, has always run reactively, one step behind bacterial evolution. A genuinely predictive tool, even an imperfect one, that could flag which resistance mutations a given drug target is structurally vulnerable to before those mutations are observed clinically, would let drug developers deprioritize fragile targets early and let clinical surveillance systems watch specifically for predicted high-risk mutations rather than scanning blindly across an entire genome — shifting at least part of the resistance fight from purely reactive to genuinely anticipatory for the first time.
The Human Dimension
There’s a certain audacity in the idea of trying to out-think bacterial evolution using the same structural logic evolution itself operates by — treating a bacterium’s genome not as a fixed target to react to, but as a space of possible future mutations whose shapes and consequences can, in principle, be mapped in advance. It won’t stop resistance from evolving; nothing will, as Alexander Fleming warned in 1946 almost as soon as antibiotics existed at all. But for the first time, the tools exist to see a meaningful piece of that evolutionary path coming, rather than only ever documenting it after the fact.
Sources:
1. “Risk-Based Prediction of Novel AMR Variants Using Protein Language Models,” bioRxiv, September 2025 — https://www.biorxiv.org/content/10.1101/2025.09.12.672331v1.full
2. “The structural context of mutations in proteins predicts their effect on antibiotic resistance,” bioRxiv, September 2025 — https://www.biorxiv.org/content/10.1101/2025.09.23.676583.full.pdf
3. Dewachter et al., “Deep mutational scanning of essential bacterial proteins can guide antibiotic development,” Nature Communications, 2023 — https://www.nature.com/articles/s41467-023-35940-3
4. “FAQs — AlphaFold Protein Structure Database” (AllTheBacteria dataset documentation) — https://alphafold.ebi.ac.uk/faq
5. “Predicting antibiotic resistance genes and bacterial phenotypes based on protein language models,” Frontiers in Microbiology, 2025 — https://www.frontiersin.org/journals/microbiology/articles/10.3389/fmicb.2025.1628952/full
6. “The transformative impact of AI-enabled AlphaFold 3: evolution, current status, and future prospects in structural biology,” PMC, 2026 — https://pmc.ncbi.nlm.nih.gov/articles/PMC13099841/
7. “Predicting Drug Resistance Using Deep Mutational Scanning,” Molecules (MDPI) — https://www.mdpi.com/1420-3049/25/9/2265
Idea originated at artificialideas.org. Article researched and written by Claude Sonnet 4.6. Published at artificialideas.org.