Up front: I’m not a specialist. I have an undergraduate biology degree, which is nowhere near enough to form independent expert judgment here. Most of what follows comes from public papers, institutional reports, and industry writing, cross-checked against my own notes, with AI help organizing. This is a learning essay so I can understand what is actually happening in AI × biology — not a professional review. Take it with a grain of salt.
Why write it? “AI × bio” usually shows up as two incompatible storylines: “most diseases cured in a decade” versus “chatbots about to enable bioweapons.” When Google released AlphaFold, I already felt it was a turning point for biotech — and I also worried about the risks. Both sides cite real material. Readers rarely see that they share one technical stack. My read is that benefits and risks ride the same rails — structure prediction, sequence generation, lab automation, cheap synthesis — and the bottlenecks sit on different layers.
Outline:
- Definitions and jargon
- Upside: what happened, what hasn’t
- Risk: how far the evidence goes, and common overclaims
- Who does what
- Public predictions
- Takeaways and open questions
First: it is not one “biology AI”
People say “AI for biology” when they mean at least three layers:
Figure: A chat models → B biology-specific models → C synthesis and labs. Upside and risk share this line.
One line: A lowers the “do you understand” bar; B lowers the “can you design it” bar; C decides whether it becomes a physical object. Talk only about models and you overstate risk. Talk only about lab hardware and you miss how fast the digital layer moved.
A few basic terms will keep coming back:
- Wet lab: real tubes, cells, gels — as opposed to in silico simulation on a computer.
- Dual-use: the same tech can help or harm. Gene editing, virology methods, and drug design are classic cases.
- CBRN: chemical, biological, radiological, nuclear. When labs ask whether a model helps with mass-casualty weapons, bioweapon assistance usually sits in this bucket.
- Nucleic acid synthesis screening: you order custom DNA from a company; they check the order against a “sequences of concern” database before shipping. One of the most concrete physical-layer defenses.
Upside: what happened, what hasn’t
1. Structure prediction: AlphaFold as infrastructure
Proteins are amino-acid chains folded into 3D machines. Drugs often need to know the shape and which pockets bind. Experiment (X-ray, cryo-EM) is slow and expensive.
Figure: schematic, not a real AlphaFold render. The model’s job is to guess the shape on the right from the sequence on the left.
DeepMind’s AlphaFold 2 (2021) jumped structure prediction; AlphaFold 3 (2024) models interactions with DNA, RNA, and small molecules. Hassabis et al. shared the 2024 Chemistry Nobel. Usage stats move, but “default infrastructure” is fair.
Limits (jargon-heavy; dig deeper if you care):
- Static snapshot ≠ dynamics. One (or a few) guessed poses, not the real conformational ensemble; induced fit and large apo/holo shifts still fail often.
- Shape ≠ binding strength. Pose hints beat affinity ranking; AF3 alone is weak on “which ligand binds tighter” and drug-likeness.
- Disorder, antibody–antigen, some nucleic-acid / lipid complexes are uneven; disordered regions can hallucinate fake order — trust pLDDT, not blindly.
- Chemistry details break: chirality errors, atom clashes; some post-cutoff benchmarks drop hard (memorization vs generalization still debated).
- Structure ≠ a drug. Synthesis, ADMET, tox, human trials remain. Equating AlphaFold with “cancer solved” is press copy.
Isomorphic, physics refinement, wet validation — mostly patching these holes.
2. From structure to design: Isomorphic, Insilico
Isomorphic Labs (DeepMind spinout) builds a drug-design engine (IsoDDE): binding affinity, cryptic pockets — the part AlphaFold is weak at. 2026 benchmarks claim large gains over AlphaFold 3 on hard generalization tasks; partnerships with Novartis, Eli Lilly, and multi-billion funding. Hassabis has publicly hoped for AI-designed candidates entering human trials around end-2026 — a goal, not a done deal.
Closer to “already in humans”: Insilico Medicine’s rentosertib (IPF) — generative AI in target and molecule discovery; Phase IIa in Nature Medicine. Important, with boundaries:
- As of writing, no widely accepted end-to-end AI-platform drug with full FDA approval (trial success ≠ approval).
- Phase II still shows safety signals (e.g. liver-related discontinuations).
- AI mostly compresses early discovery (targets, hit compounds). It does not magically turn ~90% Phase II/III failure into ~10%.
3. Sequence models: Evo writes DNA
Not “which is more accurate than AlphaFold” — different jobs, complementary. AlphaFold takes an amino-acid sequence and guesses 3D shape (“what does this protein look like?”). Genomic language models like Evo train on DNA letters and predict/generate at the sequence layer (“what if I edit this genome / can I write DNA that looks real?”). A common pipeline: Evo proposes DNA → translate to protein → AlphaFold-check the fold → wet-lab validate. Each step can fail.
Arc Institute’s Evo / Evo2 (Evo2 in Nature, Mar 2026) trains on genomes at scale: predict mutation effects and generate new sequences. Open weights. Lab-validated CRISPR systems, toxin–antitoxin modules, etc. — selected success cases, not “whatever it emits will live.”
Limits (jargon-heavy; dig deeper if you care):
- Looks genomic ≠ works in a cell. Missing essentials, broken regulation, long-range organization unlike natural genomes — independent work shows simple classifiers can separate model-generated from natural sequence.
- Specialists often win on specialist tasks: protein DMS, distal regulatory variants — alignment+structure or chromatin models are often stabler; bigger ≠ better on every metric.
- No expression context, no phenotype. A DNA string doesn’t say which cell type, under what conditions, or whether the immune system kills it.
- Wet validation is sparse, expensive, slow. Public wins are few; failure rates and screen costs rarely make headlines. Training also excluded eukaryotic viruses and other sensitive classes — capability and safety boundaries overlap.
Virtual cells, perturbation models, generate→wet-lab→RL loops — closing “letters look right, function unknown.” Risk section comes back to Evo.
4. Virtual cells: still early
Roughly: train on single-cell data to predict “if you knock out this gene / add this drug, how does expression shift?” — closer to phenotype than raw sequence. Stanford / Genentech vision papers in Cell; Arc’s Virtual Cell Challenge; commercial claims (e.g. CellOS). A 2026 Nature Methods line of work suggests performance can plateau after modest data — not endless scaling. Far from simulating a full pathogen life cycle in silico.
5. Imaging, trial matching, navigation
Covered more in this site’s healthcare post. Short version: radiology AI and enrollment tools are further along than “AI cures aging.” Longevity escape velocity is still a concept, not a demonstrated human outcome (Kurzweil ~2029 optimism; Aubrey de Grey more often ~50% by ~2037 — personal forecasts, not trial results).
Upside table
| Fairly solid | Accelerating, not cashed | Easy to overclaim |
|---|---|---|
| Protein / complex structure prediction | AI molecules into human trials | ”Most diseases already cured” |
| Shorter early discovery (case studies) | Virtual-cell perturbation models | ”Immortality by 2030” |
| Imaging with clear ground truth | Self-driving labs into biology | AlphaFold = new drug |
Risk: how far the evidence goes
Physical layer still binds hard. “Models ace bio quizzes” ≠ “anyone can build a plague.” Tiers first:
| Tier | Meaning | ~2026 status (my synthesis) |
|---|---|---|
| 0 | Generic science tutoring | Long present |
| 1 | Post-jailbreak actionable hazardous detail for someone with some background | Basically confirmed at frontier (WMDP, red teams) |
| 2 | Speeds skilled actors through design–synthesis–experiment | Emerging (Evo red teams; RAND “assistive”) |
| 3 | Low-skill end-to-end viable | Not publicly confirmed |
“Not confirmed” ≠ “never.” It means this blog should not treat Tier 3 as a fact.
1. Chat models: from textbooks to wet-lab troubleshooting
WMDP (CAIS, 2024): multiple-choice hazardous knowledge. CAIS (Center for AI Safety) is a San Francisco nonprofit that builds dangerous-capability evals and carries scores into policy debates. WMDP is useful but conceptual — more “did you memorize this?” than “the assay just failed on the bench.”
Closer to the lab is 2025’s VCT (Virology Capabilities Test) from SecureBio + CAIS. SecureBio does biosecurity engineering (nucleic-acid screening, detection, etc.) and co-built the eval to sit closer to real wet-lab work.
What is VCT actually testing?
Questions were built with dozens of virology PhD-track scientists. 322 items, many with lab images. Not “how many RNA segments does influenza have” (Googleable), but: plaque assay looks wrong, cells misbehave, reagents are off — what most likely failed? They also filtered out items that STEM non-biologists could pass with web search.
How to read the scores
Expert virologists averaged ~22% in their own subareas (the test is hard); o3 ~44%, above ~94% of matched experts; Gemini 1.5 Pro beat expert median as early as early 2024. Both 22% and 44% mean “brutal exam,” not “omniscient model.”
Why should a non-specialist care?
Before, when a complex virology protocol stalled, you usually asked the person who trained you in that lab, or a tiny network of hands-on experts. That tacit knowledge was scarce.
VCT says: publicly available chat models can already match or beat many trained virologists at text + image troubleshooting.
In plain language: “expert-level lab assistant” moved from “know a few PhDs” toward “open a browser.”
That changes how fast knowledge spreads and how cheap trial-and-error becomes — especially for people who already have lab access, reagents, and culture conditions.
It does not mean a random person in an apartment can ChatGPT their way to a controllable global pandemic. Synthesis, labs, skill, and dissemination still sit in the product below.
What it doesn’t measure:
- Multiple-choice / troubleshooting judgment, not hands. Sterile technique, instrument feel, weird smells/colors, safety calls — not on the test. Experts ~22%, o3 ~44%: even the best public models still miss most items.
- Not an uplift study. No controlled “same people, AI vs no AI, real wet-lab success rate.” SecureBio frames these like cheap in-vitro proxies: repeatable, far from real harm.
- Most hazardous dual-use content was cut — scores understate capability on sensitive protocols and cannot be read as “already builds weapons.”
- Silent wrong answers are expensive. A confident bad tip wastes reagents/time or steers someone into a worse next step; the score doesn’t separate “mildly wrong” from “dangerously wrong.”
Real wet-lab uplift, agent+tool evals (VirBench below), physical-layer screening — that’s the gap between “aces the quiz” and “does the work.”
Public frontier models can match/exceed many experts at text+image protocol debugging. End-to-end bioweapons and low-skill end-to-end feasibility — neither is shown.
VirBench (Anthropic-adjacent): how accurately agents retrieve virus sequences from NCBI (US public gene/virus databases). Sonnet 4 ~17% bare → ~93% with deterministic gget virus. Risk is also silent failure — plausible wrong answers that poison phylogenies and epitope analysis. Tool allowlists matter, not only model IQ.
2. GeneBreaker vs Evo
GeneBreaker (2025): agents elicit pathogen-homologous outputs from Evo; ~60% ASR reported on Evo2-40B. Larger models often look worse on this dual-use metric — opposite of “scale fixes safety.”
Mechanism: propose candidates → split orders / codon tricks past naive screens → assemble in a lab. Screens catch known bad sequences; generators help find novel variants.
3. Physical layer: how hard, what can a person actually do
People hear “DNA synthesis got orders of magnitude cheaper” and jump to “building a virus got cheap.” Split two things:
Ordering a stretch of DNA ≠ making a live virus.
The default path is not “a big lab must do the whole virus for you.” It is: you (or your university/company lab) email a sequence to a commercial synthesis company (Twist, IDT, many others) — a DNA print shop — and they mail fragments back. Assembly, culture, infection work stays on your side. They sell nucleic acids, not a ready-made pathogen.
Ballpark money (DNA only, ~2025–26 list prices)
Gene fragments often start around $0.07 / base pair (Twist / IDT-class list):
| What you order | Rough cost (DNA only) |
|---|---|
| 1 kb gene fragment | ~$70 |
| 3 kb | ~$210 |
| 5 kb | ~$350 |
| Clonal gene (sequence-verified, in a plasmid) | ~$0.09/bp → ~$90 for 1 kb |
Influenza-scale genomes ~13 kb, SARS-CoV-2 ~30 kb: split into fragments, DNA itself is often hundreds to low thousands of dollars — usually not the lab budget bottleneck. Decades ago the same length was one or more orders of magnitude more expensive. That is what “costs collapsed” means.
What an individual / small lab can actually do
| Path | Reality |
|---|---|
| Commercial gene fragments | Default. Institutional accounts usually work; IGSC members screen sequences of concern. No affiliation, grey vendors, non-members are policy gaps |
| Benchtop DNA synthesizer | Exists, but most print short oligos (single-stranded snippets, often ~60–120 bp). Gene-length still needs assembly; gear+reagents often tens of thousands USD+ — usually worse than ordering |
| Laptop only, no wet lab | You can design, maybe buy some DNA; still far from a live, transmissible pathogen |
So: individuals are not “totally unable to order DNA.” What kills most attempts is everything after the package arrives.
Figure: green “order DNA” is relatively cheap and mailable; orange steps — assembly, culture, rescue, BSL — raise failure rates and access barriers.
After the DNA arrives (the hard part)
- Assembly
Vendors often ship shorter double-stranded pieces. You stitch them into the right order and backbone — commonly a plasmid (circular DNA that bacteria can copy). One wrong join, missing piece, or two-base typo and the rest is trash. AI can suggest a plan; hands still do the bench work.
Figure: plasmid schematic. Foreign inserts usually need to sit on a circle bacteria can copy. Wikimedia, CC BY-SA 2.5.
- Cells and culture
DNA does not become a virus by itself. You usually put the construct into a suitable host cell line so the cell reads it out into proteins and particles. Needs sterile technique, media, incubators, sometimes licensed cell lines. Killing a batch or contaminating a batch is normal.

Figure: cultured cells (stained). Looks calm; bad state, contamination, or wrong density stops everything downstream. Wikimedia, public domain.

Figure: flasks on a shaker — common setup when amplifying plasmids in bacteria. Wikimedia, CC BY-SA 3.0.
- Reverse genetics / rescue
For many RNA viruses, genome text is not enough — you need a protocol that “rescues” infectious particles from DNA/RNA templates in cells. Specialist craft, not free with the order.

Figure: CDC influenza research bench (atmosphere of real virology work — not a how-to). Wikimedia / CDC, public domain.
- Failure and redo
Wrong colonies, bad sequencing, low expression, unhappy cells — two weeks per cycle is not rare. That $200 of DNA can be a rounding error next to time, reagents, sequencing, and labor.

Figure: colonies on agar. Screening clones often means picking many for sequencing; wrong ones outnumber clean hits. Wikimedia, CC BY 4.0.
- Hands and tacit skill
Pipetting, sterility, spotting contamination, tuning conditions — without apprenticeship, success rates collapse.

Figure: pipetting. Looks simple; volume error, cross-contamination, and shaky technique make results look done but untrustworthy. Wikimedia, public domain.
- Compliance and access
- BSL-2/3/4: biosafety lab levels. Higher = more dangerous agents allowed, stricter doors/PPE/training. Not a garage you rent with cash.
- Select Agents (US): listed high-risk pathogens with extra possession/transport rules.
- DURC etc.: federally funded dual-use wet work can need extra review.
- Vendor nucleic-acid screening: check orders against sequences of concern — one of the best physical-layer levers, still leaky (short oligos, split orders, non-members, cross-border).

Figure: NIAID high-containment biosafety cabinet — training, badge access, and SOPs; not something you recreate from Amazon. Wikimedia / NIAID, public domain.
Screening and policy (the order-time choke)
IGSC (International Gene Synthesis Consortium) runs voluntary industry screening — incomplete coverage. US OSTP framework: federally funded buys should go through compliant providers — guidance, not a global private-order law. June 2026 screendna.org letter (Altman, Amodei, Hassabis, Suleyman + synthesis CEOs) pushes mandatory screening and logs.
What is BMIA?
BMIA = Biosecurity Modernization and Innovation Act, Senate bill S.3741, introduced Jan 2026 by Republican Tom Cotton and Democrat Amy Klobuchar. Plain aim: move “industry screens if it feels like it” toward binding federal rules.
Three chunks:
- Mandatory nucleic-acid synthesis screening — companies that make/sell synthetic DNA must check sequences (look like sequences of concern?) and customers (who is ordering?). Also anti–split-order rules — same dangerous stretch broken across vendors to dodge one screen. That is exactly the “AI designs it → order-time is still the choke” story above.
- NIST biotech governance sandbox — a testbed at NIST so screening tools and rules can iterate instead of freezing one obsolete standard.
- Clean up who owns what in the feds — force an inventory of biosecurity authorities, gaps, and an implementation plan; oversight is currently fragmented across agencies.
Status: introduced bill, not law. Introduced ≠ passed. When this post says “BMIA-class direction, no hard national mandate yet,” it means: BMIA is a main proposal in the debate, but today’s orders still live under IGSC voluntary norms + OSTP guidance for federal procurement, plus lab/state rules. Gaps remain: non-members, cross-border arbitrage, short oligos, benchtops.
NIH, BSL, and Select Agents govern physical access and funded wet work — not sequences drawn on a laptop.
Product, not a single switch
Risk is closer to a product than “stronger model → automatic catastrophe”:
VCT moves knowledge/troubleshooting; Evo moves design; synthesis screening moves “can you buy a bad fragment?”; BSL and cell craft move “does it become live and spread?” Chat users scaling from thousands to hundreds of millions grows advice access — not pandemic odds by the same factor — unless synthesis and labs fail too.
The intuitive takeaway:
- No-lab amateurs: cheaper DNA ≠ plague DIY; assembly–cells–skill–compliance still bind hard.
- Skilled actors with lab access: models mainly cut design/troubleshooting friction; physical layer still there, lower friction.
- Best cheap policy lever: still order-time screening + KYC — the narrow door from design to matter.
Summary
- Skilled actors with lab access: risk up a lot; models cut trial-and-error.
- No-lab amateurs: still hard; don’t buy “orders of magnitude” headlines wholesale — that drop is DNA unit price, not end-to-end weapon difficulty.
- Future experiment-running agents: a tier above today’s chatbots; early.
- Accidents: more people doing dangerous work can raise accident surface, not only malice.
Who does what
Research / evals
| Who | One-line who | Role here |
|---|---|---|
| CAIS | AI safety eval / advocacy nonprofit | WMDP, VCT |
| SecureBio | Biosecurity engineering org | VCT; nucleic acids / detection |
| RAND | Think tank, expert Delphi | ”Still mostly assistive through ~2027” |
| Arc Institute | Biology research institute | Evo, virtual cells, open models |
| Anthropic / OpenAI / DeepMind | Frontier labs | Safety frameworks + CBRN red team / refusal |
Industry
Isomorphic, Insilico, Recursion, Xaira, XtalPi, etc.; China industrialization signals (check dates/headlines). Incubators like Halcyon treat biosecurity as its own direction (e.g. Red Queen Bio: AI + wet lab for defense).
Governance (mostly US)
| Layer | Targets | Covers pure in-silico design? |
|---|---|---|
| Lab RSP / Preparedness | Model outputs, deploy gates | Partly |
| State laws (RAISE, SB 53) | Catastrophic-risk process | Indirectly |
| OSTP / IGSC screening | Order-time | No |
| NIH DURC / BSL / Select Agents | Funded wet work + pathogens | Mostly no |
Gap in one sentence: design is digital; interception is often at order-time and the lab door. Open DNA models (Evo) vs gated weights (AlphaFold 3) is the open-science vs misuse tradeoff.
Public timelines
Dario Amodei (Anthropic)
In Machines of Loving Grace, biology is a prime upside of powerful AI (compressing scientific decades). In The Adolescence of Technology (2026), he says internal mid-2025 measures suggest models may already raise success odds on some bioweapon-relevant tasks by roughly 2–3×, and that models may be nearing help for STEM-but-not-biology people through more of the pipeline. Anthropic released some Claude models under higher safety levels and runs bio-weapon classifiers at non-trivial inference cost. That is one lab’s internal red-team narrative — hard to independently reproduce — but directionally aligned with public VCT/WMDP.
RAND Delphi (~2025–2027)
Biology and AI experts asked where hard constraints sit (see RBA4087-1). Near-term read: AI mainly assists experts, cannot yet independently design novel pathogens or make pure novices succeed end-to-end; longer horizons are separate. Policy push: coordinated gene-synthesis screening, cloud-lab identity checks, pandemic preparedness (diagnostics, vaccine platforms) as backstop.
Demis Hassabis / Isomorphic
Public timelines have slipped; recent hope is AI-designed candidates entering clinic around end-2026. Industry-optimistic target.
Longevity camp
Kurzweil’s longevity escape velocity, de Grey’s SENS/LEV — biological frameworks exist; hard human endpoints do not match the slogans. Not the same claim as “AI shortens early drug discovery.”
My own forecast notes (subjective): skilled-actor + AI assist window ~2026–27; low-skill end-to-end still unconfirmed; extinction-scale bio as a fat tail, not my modal path. Details in the site p(doom) evidence chain (bio especially the CBRN node). Those numbers are elicitations, not measurements.
Takeaways
Biology research is clearly dual-use.
- Upside is front-loaded; the back end (humans, regulators, failure rates) is still hard.
- Near-term cheap governance levers are probably synthesis screening + KYC + evals.
Worth watching next:
- SecureBio-style real wet-lab uplift studies: does AI assist actually raise novice/expert success rates?
- Can agents stably chain design → order → lab gear?
- When does function-based synthesis screening scale past sequence homology?
Sources
- AlphaFold 3: https://www.nature.com/articles/s41586-024-07487-w
- Isomorphic Drug Design Engine: https://www.isomorphiclabs.com/articles/the-isomorphic-labs-drug-design-engine-unlocks-a-new-frontier
- Insilico rentosertib Phase IIa: https://www.nature.com/articles/s41591-025-03743-2
- Evo 2 (Nature 2026): https://www.nature.com/articles/s41586-026-10176-5
- VCT: https://www.virologytest.ai/ · arXiv:2504.16137
- GeneBreaker: https://arxiv.org/abs/2505.23839
- WMDP: https://www.wmdp.ai/
- VirBench: https://arxiv.org/html/2606.06749 · https://www.anthropic.com/research/agents-in-biology
- RAND RBA4087-1: https://www.rand.org/pubs/research_briefs/RBA4087-1.html
- screendna letter: https://screendna.org/
- Amodei, Adolescence of Technology: https://darioamodei.com/essay/the-adolescence-of-technology
- Amodei, Machines of Loving Grace: https://www.darioamodei.com/essay/machines-of-loving-grace
- Twist gene synthesis pricing (list ~7¢/bp fragments): https://www.twistbioscience.com/products/genes/gene-synthesis
- IDT gene fragments (~7¢/bp, 2025): https://www.businesswire.com/news/home/20250710425195/en/IDT-Launches-All-Inclusive-Gene-Fragments-From-7-Cents-Per-Base-With-25-Lower-Turnaround-Times-and-Next-Day-Shipping-Powered-by-Its-Cutting-Edge-Synthetic-Biology-Facility