Researchers from Stanford University and the Arc Institute have achieved a historic milestone by using a generative AI model to design complete viral genomes from scratch. This breakthrough marks a pivot from traditional biological discovery to an era of computational design, where AI dictates the architecture of life.

Evo: The Large Language Model for Genomics

At the heart of this research is Evo, a specialized generative model that functions similarly to a Large Language Model (LLM) but operates on the "language" of DNA. Unlike text-based models, Evo was trained on a massive dataset comprising roughly nine trillion nucleotides sourced from millions of organisms, including animals, plants, microbes, and viruses.

To refine its precision, the researchers employed a two-stage training process. After learning the foundational patterns of the tree of life, Evo underwent a second, specialized round of training using the 11 genes of the phiX174 bacteriophage and approximately 15,000 of its closest relatives. This allowed the model to master the specific nuances required to design functional, replicating biological entities.

From 700,000 Proposals to Functional Viruses

The scale of Evo's generative capabilities is immense. The model proposed 700,000 potential viral genomes. The research team then narrowed these down through rigorous selection, chemically synthesizing 285 sequences as DNA strands and inserting them into E. coli bacteria.

The results were highly successful: 16 of the synthesized sequences produced viruses capable of replication. Notably, these AI-generated viruses were not merely "sickly" or weakened versions of existing species; they were as robust as natural viruses, with some replicating even faster than the original phiX174. Experts noted that the AI even introduced unexpected changes to gene orders and arrangements that human scientists had not previously considered.

The Dual-Edge Sword: Medical Breakthroughs vs. Biosafety Gaps

The implications for biotechnology are profound. AI-designed viruses could revolutionize phage therapy—using viruses to kill multidrug-resistant bacteria—and enhance gene therapy by creating more efficient delivery vehicles for genetic material.

However, the development has highlighted a critical "disconnect" in global biosafety regulations. Current U.S. National Institutes of Health (NIH) policies focus on prohibiting experiments that make known pathogens more dangerous, but they do not explicitly cover purely computational work—the design of viral DNA on a computer. This creates a regulatory vacuum for "unknown" risks: if an AI designs a novel, highly transmissible pathogen that has never existed in nature, existing guardrails may fail to identify the threat.

To mitigate this, the Stanford team implemented proactive safety measures by ensuring Evo was never trained on human pathogens or related animal/plant viruses, effectively preventing the model from being able to generate human-infecting genomes.

Why the result matters for medicine

Phage therapy, which employs viruses to target antibiotic-resistant bacteria, has long been hampered by the limited toolbox of natural phages. An AI that can conjure new, efficient viral capsids could expand that toolbox dramatically. Similarly, gene-therapy vectors—often engineered viruses that ferry therapeutic DNA into cells—might become more precise and less immunogenic if designed from scratch rather than modified from existing strains.

The regulatory blind spot exposed

U.S. funding agencies currently restrict experiments that make known pathogens more dangerous, but they do not explicitly address the creation of novel pathogens through purely computational design. Evo’s work sidestepped that clause because the model never saw human-infecting viruses during training; the researchers deliberately excluded any sequences that could give rise to human pathogens. Nonetheless, the possibility remains that an AI could, without such safeguards, devise a virus with no natural precedent—a “synthetic unknown” that existing review boards might not flag.

Luka ta nie ma charakteru wyłącznie teoretycznego. Gdyby modelowi pozwolono trenować na szerszym zbiorze danych wirusowych, mógłby on nauczyć się podstawowych elementów wysokiej zakaźności lub unikania odpowiedzi immunologicznej, a następnie wygenerować genom, który po zsyntetyzowaniu mógłby stanowić zagrożenie na skalę pandemii. Obecne mechanizmy nadzoru koncentrują się na manipulacjach w laboratoriach typu wet-lab; brakuje im jasnego mandatu do oceny ryzyka związanego z projektowaniem in silico, które może być przeprowadzane zdalnie i na dużą skalę.

Wyprodukowane wirusy to bakteriofagi — wirusy atakujące bakterie, a nie ludzi. Ich zastosowanie ogranicza się głównie do kontroli bakterii i badań laboratoryjnych, co minimalizuje bezpośrednie obawy dla zdrowia publicznego.

Podsumowanie

Evo pokazuje, że sztuczna inteligencja może przejść od przewidywania struktur białkowych do projektowania całych, funkcjonalnych genomów wirusowych. Technologia ta obiecuje szybsze i bardziej wszechstronne narzędzia do walki z bakteriami lekoopornymi oraz do dostarczania terapii genowych, ale zmusza również regulatorów do zmierzenia się z nową klasą ryzyka: cyfrowymi schematami życia, które nigdy nie istniały w naturze.