Researchers from Stanford University and the Arc Institute have achieved a historic milestone by using a generative AI model to design complete viral genomes from scratch. This breakthrough marks a pivot from traditional biological discovery to an era of computational design, where AI dictates the architecture of life.
Evo: The Large Language Model for Genomics
At the heart of this research is Evo, a specialized generative model that functions similarly to a Large Language Model (LLM) but operates on the "language" of DNA. Unlike text-based models, Evo was trained on a massive dataset comprising roughly nine trillion nucleotides sourced from millions of organisms, including animals, plants, microbes, and viruses.
To refine its precision, the researchers employed a two-stage training process. After learning the foundational patterns of the tree of life, Evo underwent a second, specialized round of training using the 11 genes of the phiX174 bacteriophage and approximately 15,000 of its closest relatives. This allowed the model to master the specific nuances required to design functional, replicating biological entities.
From 700,000 Proposals to Functional Viruses
The scale of Evo's generative capabilities is immense. The model proposed 700,000 potential viral genomes. The research team then narrowed these down through rigorous selection, chemically synthesizing 285 sequences as DNA strands and inserting them into E. coli bacteria.
The results were highly successful: 16 of the synthesized sequences produced viruses capable of replication. Notably, these AI-generated viruses were not merely "sickly" or weakened versions of existing species; they were as robust as natural viruses, with some replicating even faster than the original phiX174. Experts noted that the AI even introduced unexpected changes to gene orders and arrangements that human scientists had not previously considered.
The Dual-Edge Sword: Medical Breakthroughs vs. Biosafety Gaps
The implications for biotechnology are profound. AI-designed viruses could revolutionize phage therapy—using viruses to kill multidrug-resistant bacteria—and enhance gene therapy by creating more efficient delivery vehicles for genetic material.
However, the development has highlighted a critical "disconnect" in global biosafety regulations. Current U.S. National Institutes of Health (NIH) policies focus on prohibiting experiments that make known pathogens more dangerous, but they do not explicitly cover purely computational work—the design of viral DNA on a computer. This creates a regulatory vacuum for "unknown" risks: if an AI designs a novel, highly transmissible pathogen that has never existed in nature, existing guardrails may fail to identify the threat.
To mitigate this, the Stanford team implemented proactive safety measures by ensuring Evo was never trained on human pathogens or related animal/plant viruses, effectively preventing the model from being able to generate human-infecting genomes.
Why the result matters for medicine
Phage therapy, which employs viruses to target antibiotic-resistant bacteria, has long been hampered by the limited toolbox of natural phages. An AI that can conjure new, efficient viral capsids could expand that toolbox dramatically. Similarly, gene-therapy vectors—often engineered viruses that ferry therapeutic DNA into cells—might become more precise and less immunogenic if designed from scratch rather than modified from existing strains.
The regulatory blind spot exposed
U.S. funding agencies currently restrict experiments that make known pathogens more dangerous, but they do not explicitly address the creation of novel pathogens through purely computational design. Evo’s work sidestepped that clause because the model never saw human-infecting viruses during training; the researchers deliberately excluded any sequences that could give rise to human pathogens. Nonetheless, the possibility remains that an AI could, without such safeguards, devise a virus with no natural precedent—a “synthetic unknown” that existing review boards might not flag.
Этот разрыв не носит чисто теоретический характер. Если модели позволят обучаться на более широком наборе вирусных данных, она может усвоить базовые принципы высокой трансмиссивности или уклонения от иммунного ответа, а затем выдать геном, который после синтеза может представлять угрозу пандемического масштаба. Текущие механизмы надзора сосредоточены на манипуляциях в «мокрой» лаборатории; им не хватает четких полномочий для оценки рисков проектирования in silico, которое может осуществляться удаленно и в больших масштабах.
Созданные вирусы были бактериофагами — вирусами, которые поражают бактерии, а не людей. Их полезность в основном ограничена контролем бактерий и лабораторными исследованиями, что снижает непосредственные опасения для общественного здравоохранения.
Итог
Evo демонстрирует, что ИИ может перейти от предсказания структур белков к проектированию целых функциональных вирусных геномов. Эта технология обещает создание более быстрых и универсальных инструментов для борьбы с лекарственно-устойчивыми бактериями и доставки генной терапии, но она также заставляет регуляторов столкнуться с новым классом рисков: цифровыми чертежами жизни, которых никогда не существовало в природе.
