← Back to all writing

AI × biology: a long beginner map

July 16, 2026

Up front: I’m not a specialist. I have an undergraduate biology degree, which is nowhere near enough to form independent expert judgment here. Most of what follows comes from public papers, institutional reports, and industry writing, cross-checked against my own notes, with AI help organizing. This is a learning essay so I can understand what is actually happening in AI × biology — not a professional review. Take it with a grain of salt.


Why write it? “AI × bio” usually shows up as two incompatible storylines: “most diseases cured in a decade” versus “chatbots about to enable bioweapons.” When Google released AlphaFold, I already felt it was a turning point for biotech — and I also worried about the risks. Both sides cite real material. Readers rarely see that they share one technical stack. My read is that benefits and risks ride the same rails — structure prediction, sequence generation, lab automation, cheap synthesis — and the bottlenecks sit on different layers.

Outline:

  1. Definitions and jargon
  2. Upside: what happened, what hasn’t
  3. Risk: how far the evidence goes, and common overclaims
  4. Who does what
  5. Public predictions
  6. Takeaways and open questions

First: it is not one “biology AI”

People say “AI for biology” when they mean at least three layers:

AI × biology three-layer stack

Figure: A chat models → B biology-specific models → C synthesis and labs. Upside and risk share this line.

One line: A lowers the “do you understand” bar; B lowers the “can you design it” bar; C decides whether it becomes a physical object. Talk only about models and you overstate risk. Talk only about lab hardware and you miss how fast the digital layer moved.

A few basic terms will keep coming back:

  • Wet lab: real tubes, cells, gels — as opposed to in silico simulation on a computer.
  • Dual-use: the same tech can help or harm. Gene editing, virology methods, and drug design are classic cases.
  • CBRN: chemical, biological, radiological, nuclear. When labs ask whether a model helps with mass-casualty weapons, bioweapon assistance usually sits in this bucket.
  • Nucleic acid synthesis screening: you order custom DNA from a company; they check the order against a “sequences of concern” database before shipping. One of the most concrete physical-layer defenses.

Upside: what happened, what hasn’t

1. Structure prediction: AlphaFold as infrastructure

Proteins are amino-acid chains folded into 3D machines. Drugs often need to know the shape and which pockets bind. Experiment (X-ray, cryo-EM) is slow and expensive.

Schematic: 1D sequence → 3D fold

Figure: schematic, not a real AlphaFold render. The model’s job is to guess the shape on the right from the sequence on the left.

DeepMind’s AlphaFold 2 (2021) jumped structure prediction; AlphaFold 3 (2024) models interactions with DNA, RNA, and small molecules. Hassabis et al. shared the 2024 Chemistry Nobel. Usage stats move, but “default infrastructure” is fair.

Limits (jargon-heavy; dig deeper if you care):

  • Static snapshot ≠ dynamics. One (or a few) guessed poses, not the real conformational ensemble; induced fit and large apo/holo shifts still fail often.
  • Shape ≠ binding strength. Pose hints beat affinity ranking; AF3 alone is weak on “which ligand binds tighter” and drug-likeness.
  • Disorder, antibody–antigen, some nucleic-acid / lipid complexes are uneven; disordered regions can hallucinate fake order — trust pLDDT, not blindly.
  • Chemistry details break: chirality errors, atom clashes; some post-cutoff benchmarks drop hard (memorization vs generalization still debated).
  • Structure ≠ a drug. Synthesis, ADMET, tox, human trials remain. Equating AlphaFold with “cancer solved” is press copy.

Isomorphic, physics refinement, wet validation — mostly patching these holes.

2. From structure to design: Isomorphic, Insilico

Isomorphic Labs (DeepMind spinout) builds a drug-design engine (IsoDDE): binding affinity, cryptic pockets — the part AlphaFold is weak at. 2026 benchmarks claim large gains over AlphaFold 3 on hard generalization tasks; partnerships with Novartis, Eli Lilly, and multi-billion funding. Hassabis has publicly hoped for AI-designed candidates entering human trials around end-2026 — a goal, not a done deal.

Closer to “already in humans”: Insilico Medicine’s rentosertib (IPF) — generative AI in target and molecule discovery; Phase IIa in Nature Medicine. Important, with boundaries:

  • As of writing, no widely accepted end-to-end AI-platform drug with full FDA approval (trial success ≠ approval).
  • Phase II still shows safety signals (e.g. liver-related discontinuations).
  • AI mostly compresses early discovery (targets, hit compounds). It does not magically turn ~90% Phase II/III failure into ~10%.

3. Sequence models: Evo writes DNA

Not “which is more accurate than AlphaFold” — different jobs, complementary. AlphaFold takes an amino-acid sequence and guesses 3D shape (“what does this protein look like?”). Genomic language models like Evo train on DNA letters and predict/generate at the sequence layer (“what if I edit this genome / can I write DNA that looks real?”). A common pipeline: Evo proposes DNA → translate to protein → AlphaFold-check the fold → wet-lab validate. Each step can fail.

Arc Institute’s Evo / Evo2 (Evo2 in Nature, Mar 2026) trains on genomes at scale: predict mutation effects and generate new sequences. Open weights. Lab-validated CRISPR systems, toxin–antitoxin modules, etc. — selected success cases, not “whatever it emits will live.”

Limits (jargon-heavy; dig deeper if you care):

  • Looks genomic ≠ works in a cell. Missing essentials, broken regulation, long-range organization unlike natural genomes — independent work shows simple classifiers can separate model-generated from natural sequence.
  • Specialists often win on specialist tasks: protein DMS, distal regulatory variants — alignment+structure or chromatin models are often stabler; bigger ≠ better on every metric.
  • No expression context, no phenotype. A DNA string doesn’t say which cell type, under what conditions, or whether the immune system kills it.
  • Wet validation is sparse, expensive, slow. Public wins are few; failure rates and screen costs rarely make headlines. Training also excluded eukaryotic viruses and other sensitive classes — capability and safety boundaries overlap.

Virtual cells, perturbation models, generate→wet-lab→RL loops — closing “letters look right, function unknown.” Risk section comes back to Evo.

4. Virtual cells: still early

Roughly: train on single-cell data to predict “if you knock out this gene / add this drug, how does expression shift?” — closer to phenotype than raw sequence. Stanford / Genentech vision papers in Cell; Arc’s Virtual Cell Challenge; commercial claims (e.g. CellOS). A 2026 Nature Methods line of work suggests performance can plateau after modest data — not endless scaling. Far from simulating a full pathogen life cycle in silico.

5. Imaging, trial matching, navigation

Covered more in this site’s healthcare post. Short version: radiology AI and enrollment tools are further along than “AI cures aging.” Longevity escape velocity is still a concept, not a demonstrated human outcome (Kurzweil ~2029 optimism; Aubrey de Grey more often ~50% by ~2037 — personal forecasts, not trial results).

Upside table

Fairly solidAccelerating, not cashedEasy to overclaim
Protein / complex structure predictionAI molecules into human trials”Most diseases already cured”
Shorter early discovery (case studies)Virtual-cell perturbation models”Immortality by 2030”
Imaging with clear ground truthSelf-driving labs into biologyAlphaFold = new drug

Risk: how far the evidence goes

Physical layer still binds hard. “Models ace bio quizzes” ≠ “anyone can build a plague.” Tiers first:

TierMeaning~2026 status (my synthesis)
0Generic science tutoringLong present
1Post-jailbreak actionable hazardous detail for someone with some backgroundBasically confirmed at frontier (WMDP, red teams)
2Speeds skilled actors through design–synthesis–experimentEmerging (Evo red teams; RAND “assistive”)
3Low-skill end-to-end viableNot publicly confirmed

“Not confirmed” ≠ “never.” It means this blog should not treat Tier 3 as a fact.

1. Chat models: from textbooks to wet-lab troubleshooting

WMDP (CAIS, 2024): multiple-choice hazardous knowledge. CAIS (Center for AI Safety) is a San Francisco nonprofit that builds dangerous-capability evals and carries scores into policy debates. WMDP is useful but conceptual — more “did you memorize this?” than “the assay just failed on the bench.”

Closer to the lab is 2025’s VCT (Virology Capabilities Test) from SecureBio + CAIS. SecureBio does biosecurity engineering (nucleic-acid screening, detection, etc.) and co-built the eval to sit closer to real wet-lab work.

What VCT measures vs does not

What is VCT actually testing?
Questions were built with dozens of virology PhD-track scientists. 322 items, many with lab images. Not “how many RNA segments does influenza have” (Googleable), but: plaque assay looks wrong, cells misbehave, reagents are off — what most likely failed? They also filtered out items that STEM non-biologists could pass with web search.

How to read the scores
Expert virologists averaged ~22% in their own subareas (the test is hard); o3 ~44%, above ~94% of matched experts; Gemini 1.5 Pro beat expert median as early as early 2024. Both 22% and 44% mean “brutal exam,” not “omniscient model.”

Why should a non-specialist care?
Before, when a complex virology protocol stalled, you usually asked the person who trained you in that lab, or a tiny network of hands-on experts. That tacit knowledge was scarce.
VCT says: publicly available chat models can already match or beat many trained virologists at text + image troubleshooting.
In plain language: “expert-level lab assistant” moved from “know a few PhDs” toward “open a browser.”
That changes how fast knowledge spreads and how cheap trial-and-error becomes — especially for people who already have lab access, reagents, and culture conditions.
It does not mean a random person in an apartment can ChatGPT their way to a controllable global pandemic. Synthesis, labs, skill, and dissemination still sit in the product below.

What it doesn’t measure:

  • Multiple-choice / troubleshooting judgment, not hands. Sterile technique, instrument feel, weird smells/colors, safety calls — not on the test. Experts ~22%, o3 ~44%: even the best public models still miss most items.
  • Not an uplift study. No controlled “same people, AI vs no AI, real wet-lab success rate.” SecureBio frames these like cheap in-vitro proxies: repeatable, far from real harm.
  • Most hazardous dual-use content was cut — scores understate capability on sensitive protocols and cannot be read as “already builds weapons.”
  • Silent wrong answers are expensive. A confident bad tip wastes reagents/time or steers someone into a worse next step; the score doesn’t separate “mildly wrong” from “dangerously wrong.”

Real wet-lab uplift, agent+tool evals (VirBench below), physical-layer screening — that’s the gap between “aces the quiz” and “does the work.”

Public frontier models can match/exceed many experts at text+image protocol debugging. End-to-end bioweapons and low-skill end-to-end feasibility — neither is shown.

VirBench (Anthropic-adjacent): how accurately agents retrieve virus sequences from NCBI (US public gene/virus databases). Sonnet 4 ~17% bare → ~93% with deterministic gget virus. Risk is also silent failure — plausible wrong answers that poison phylogenies and epitope analysis. Tool allowlists matter, not only model IQ.

2. GeneBreaker vs Evo

GeneBreaker (2025): agents elicit pathogen-homologous outputs from Evo; ~60% ASR reported on Evo2-40B. Larger models often look worse on this dual-use metric — opposite of “scale fixes safety.”

Mechanism: propose candidates → split orders / codon tricks past naive screens → assemble in a lab. Screens catch known bad sequences; generators help find novel variants.

3. Physical layer: how hard, what can a person actually do

People hear “DNA synthesis got orders of magnitude cheaper” and jump to “building a virus got cheap.” Split two things:

Ordering a stretch of DNAmaking a live virus.

The default path is not “a big lab must do the whole virus for you.” It is: you (or your university/company lab) email a sequence to a commercial synthesis company (Twist, IDT, many others) — a DNA print shop — and they mail fragments back. Assembly, culture, infection work stays on your side. They sell nucleic acids, not a ready-made pathogen.

Ballpark money (DNA only, ~2025–26 list prices)

Gene fragments often start around $0.07 / base pair (Twist / IDT-class list):

What you orderRough cost (DNA only)
1 kb gene fragment~$70
3 kb~$210
5 kb~$350
Clonal gene (sequence-verified, in a plasmid)~$0.09/bp → ~$90 for 1 kb

Influenza-scale genomes ~13 kb, SARS-CoV-2 ~30 kb: split into fragments, DNA itself is often hundreds to low thousands of dollars — usually not the lab budget bottleneck. Decades ago the same length was one or more orders of magnitude more expensive. That is what “costs collapsed” means.

What an individual / small lab can actually do

PathReality
Commercial gene fragmentsDefault. Institutional accounts usually work; IGSC members screen sequences of concern. No affiliation, grey vendors, non-members are policy gaps
Benchtop DNA synthesizerExists, but most print short oligos (single-stranded snippets, often ~60–120 bp). Gene-length still needs assembly; gear+reagents often tens of thousands USD+ — usually worse than ordering
Laptop only, no wet labYou can design, maybe buy some DNA; still far from a live, transmissible pathogen

So: individuals are not “totally unable to order DNA.” What kills most attempts is everything after the package arrives.

After DNA arrives: harder at each step

Figure: green “order DNA” is relatively cheap and mailable; orange steps — assembly, culture, rescue, BSL — raise failure rates and access barriers.

After the DNA arrives (the hard part)

  1. Assembly
    Vendors often ship shorter double-stranded pieces. You stitch them into the right order and backbone — commonly a plasmid (circular DNA that bacteria can copy). One wrong join, missing piece, or two-base typo and the rest is trash. AI can suggest a plan; hands still do the bench work.

Plasmid schematic: circular DNA vector

Figure: plasmid schematic. Foreign inserts usually need to sit on a circle bacteria can copy. Wikimedia, CC BY-SA 2.5.

  1. Cells and culture
    DNA does not become a virus by itself. You usually put the construct into a suitable host cell line so the cell reads it out into proteins and particles. Needs sterile technique, media, incubators, sometimes licensed cell lines. Killing a batch or contaminating a batch is normal.

Cultured cells under a microscope

Figure: cultured cells (stained). Looks calm; bad state, contamination, or wrong density stops everything downstream. Wikimedia, public domain.

Culture flasks on a lab shaker

Figure: flasks on a shaker — common setup when amplifying plasmids in bacteria. Wikimedia, CC BY-SA 3.0.

  1. Reverse genetics / rescue
    For many RNA viruses, genome text is not enough — you need a protocol that “rescues” infectious particles from DNA/RNA templates in cells. Specialist craft, not free with the order.

Virus research bench

Figure: CDC influenza research bench (atmosphere of real virology work — not a how-to). Wikimedia / CDC, public domain.

  1. Failure and redo
    Wrong colonies, bad sequencing, low expression, unhappy cells — two weeks per cycle is not rare. That $200 of DNA can be a rounding error next to time, reagents, sequencing, and labor.

Bacterial colonies on an agar plate

Figure: colonies on agar. Screening clones often means picking many for sequencing; wrong ones outnumber clean hits. Wikimedia, CC BY 4.0.

  1. Hands and tacit skill
    Pipetting, sterility, spotting contamination, tuning conditions — without apprenticeship, success rates collapse.

Pipetting in a lab

Figure: pipetting. Looks simple; volume error, cross-contamination, and shaky technique make results look done but untrustworthy. Wikimedia, public domain.

  1. Compliance and access
    • BSL-2/3/4: biosafety lab levels. Higher = more dangerous agents allowed, stricter doors/PPE/training. Not a garage you rent with cash.
    • Select Agents (US): listed high-risk pathogens with extra possession/transport rules.
    • DURC etc.: federally funded dual-use wet work can need extra review.
    • Vendor nucleic-acid screening: check orders against sequences of concern — one of the best physical-layer levers, still leaky (short oligos, split orders, non-members, cross-border).

High-containment biosafety cabinet

Figure: NIAID high-containment biosafety cabinet — training, badge access, and SOPs; not something you recreate from Amazon. Wikimedia / NIAID, public domain.

Screening and policy (the order-time choke)

IGSC (International Gene Synthesis Consortium) runs voluntary industry screening — incomplete coverage. US OSTP framework: federally funded buys should go through compliant providers — guidance, not a global private-order law. June 2026 screendna.org letter (Altman, Amodei, Hassabis, Suleyman + synthesis CEOs) pushes mandatory screening and logs.

What is BMIA?
BMIA = Biosecurity Modernization and Innovation Act, Senate bill S.3741, introduced Jan 2026 by Republican Tom Cotton and Democrat Amy Klobuchar. Plain aim: move “industry screens if it feels like it” toward binding federal rules.

Three chunks:

  1. Mandatory nucleic-acid synthesis screening — companies that make/sell synthetic DNA must check sequences (look like sequences of concern?) and customers (who is ordering?). Also anti–split-order rules — same dangerous stretch broken across vendors to dodge one screen. That is exactly the “AI designs it → order-time is still the choke” story above.
  2. NIST biotech governance sandbox — a testbed at NIST so screening tools and rules can iterate instead of freezing one obsolete standard.
  3. Clean up who owns what in the feds — force an inventory of biosecurity authorities, gaps, and an implementation plan; oversight is currently fragmented across agencies.

Status: introduced bill, not law. Introduced ≠ passed. When this post says “BMIA-class direction, no hard national mandate yet,” it means: BMIA is a main proposal in the debate, but today’s orders still live under IGSC voluntary norms + OSTP guidance for federal procurement, plus lab/state rules. Gaps remain: non-members, cross-border arbitrage, short oligos, benchtops.

NIH, BSL, and Select Agents govern physical access and funded wet work — not sequences drawn on a laptop.

Product, not a single switch

Risk is closer to a product than “stronger model → automatic catastrophe”:

Misuse risk is a product

VCT moves knowledge/troubleshooting; Evo moves design; synthesis screening moves “can you buy a bad fragment?”; BSL and cell craft move “does it become live and spread?” Chat users scaling from thousands to hundreds of millions grows advice access — not pandemic odds by the same factor — unless synthesis and labs fail too.

The intuitive takeaway:

  • No-lab amateurs: cheaper DNA ≠ plague DIY; assembly–cells–skill–compliance still bind hard.
  • Skilled actors with lab access: models mainly cut design/troubleshooting friction; physical layer still there, lower friction.
  • Best cheap policy lever: still order-time screening + KYC — the narrow door from design to matter.

Summary

  • Skilled actors with lab access: risk up a lot; models cut trial-and-error.
  • No-lab amateurs: still hard; don’t buy “orders of magnitude” headlines wholesale — that drop is DNA unit price, not end-to-end weapon difficulty.
  • Future experiment-running agents: a tier above today’s chatbots; early.
  • Accidents: more people doing dangerous work can raise accident surface, not only malice.

Who does what

Research / evals

WhoOne-line whoRole here
CAISAI safety eval / advocacy nonprofitWMDP, VCT
SecureBioBiosecurity engineering orgVCT; nucleic acids / detection
RANDThink tank, expert Delphi”Still mostly assistive through ~2027”
Arc InstituteBiology research instituteEvo, virtual cells, open models
Anthropic / OpenAI / DeepMindFrontier labsSafety frameworks + CBRN red team / refusal

Industry

Isomorphic, Insilico, Recursion, Xaira, XtalPi, etc.; China industrialization signals (check dates/headlines). Incubators like Halcyon treat biosecurity as its own direction (e.g. Red Queen Bio: AI + wet lab for defense).

Governance (mostly US)

LayerTargetsCovers pure in-silico design?
Lab RSP / PreparednessModel outputs, deploy gatesPartly
State laws (RAISE, SB 53)Catastrophic-risk processIndirectly
OSTP / IGSC screeningOrder-timeNo
NIH DURC / BSL / Select AgentsFunded wet work + pathogensMostly no

Gap in one sentence: design is digital; interception is often at order-time and the lab door. Open DNA models (Evo) vs gated weights (AlphaFold 3) is the open-science vs misuse tradeoff.


Public timelines

Dario Amodei (Anthropic)
In Machines of Loving Grace, biology is a prime upside of powerful AI (compressing scientific decades). In The Adolescence of Technology (2026), he says internal mid-2025 measures suggest models may already raise success odds on some bioweapon-relevant tasks by roughly 2–3×, and that models may be nearing help for STEM-but-not-biology people through more of the pipeline. Anthropic released some Claude models under higher safety levels and runs bio-weapon classifiers at non-trivial inference cost. That is one lab’s internal red-team narrative — hard to independently reproduce — but directionally aligned with public VCT/WMDP.

RAND Delphi (~2025–2027)
Biology and AI experts asked where hard constraints sit (see RBA4087-1). Near-term read: AI mainly assists experts, cannot yet independently design novel pathogens or make pure novices succeed end-to-end; longer horizons are separate. Policy push: coordinated gene-synthesis screening, cloud-lab identity checks, pandemic preparedness (diagnostics, vaccine platforms) as backstop.

Demis Hassabis / Isomorphic
Public timelines have slipped; recent hope is AI-designed candidates entering clinic around end-2026. Industry-optimistic target.

Longevity camp
Kurzweil’s longevity escape velocity, de Grey’s SENS/LEV — biological frameworks exist; hard human endpoints do not match the slogans. Not the same claim as “AI shortens early drug discovery.”

My own forecast notes (subjective): skilled-actor + AI assist window ~2026–27; low-skill end-to-end still unconfirmed; extinction-scale bio as a fat tail, not my modal path. Details in the site p(doom) evidence chain (bio especially the CBRN node). Those numbers are elicitations, not measurements.


Takeaways

Biology research is clearly dual-use.

  1. Upside is front-loaded; the back end (humans, regulators, failure rates) is still hard.
  2. Near-term cheap governance levers are probably synthesis screening + KYC + evals.

Worth watching next:

  • SecureBio-style real wet-lab uplift studies: does AI assist actually raise novice/expert success rates?
  • Can agents stably chain design → order → lab gear?
  • When does function-based synthesis screening scale past sequence homology?

Sources