The federal government rarely rushes into untested technology. Public health departments, in particular, can spend years validating a tool before it ever touches patient records or outbreak data. So when U.S. health agencies announced they will run live trials with generative AI systems built by OpenAI and Anthropic, the signal was clear: large language models have moved from consumer experiment to serious institutional candidate.
What PULSE Actually Does
The new effort is called the Public Health Use Case and Learning Scaling Engine—PULSE for short. It is a testing framework, not a product rollout. Rather than dropping chatbots into health department email systems and hoping for the best, PULSE treats generative AI the same way public health treats a new vaccine. It creates controlled conditions where state, local, tribal, and territorial agencies can trial these tools while rigorously watching for side effects.
The program brings together four distinct partners. The Coalition for Health AI (CHAI) supplies healthcare-specific governance and safety standards. OpenAI and Anthropic contribute the frontier language models. Accenture provides the integration expertise needed to connect these systems to existing government infrastructure, which often runs on decades-old databases and strict procurement rules.
PULSE is not asking whether AI can write emails or summarize a meeting transcript. It is asking whether these models can reliably assist with three core tasks: epidemiological surveillance, resource allocation, and public health communication. Each of these sounds abstract, but together they describe the daily burden of running a health department. Surveillance means sifting through lab reports, hospital admissions, and school absentee data to spot an outbreak before it balloons. Resource allocation means deciding, in real time, where to send limited vaccine doses or antiviral treatments during a severe flu season. Public health communication means translating complex federal guidance into plain language for dozens of different communities, often while the clock is ticking on a foodborne illness investigation. If AI can help with any of these without creating extra work or introducing errors, the efficiency gains could be substantial.
Why Ten Jurisdictions Changes Everything
The program will run trials across ten jurisdictions drawn from states, localities, tribal nations, and U.S. territories. That deliberate spread is the entire point. A state health department in a densely populated region, staffed with hundreds of epidemiologists and running fiber-optic networks, faces entirely different constraints than a territorial agency tracking dengue with a skeleton crew and intermittent connectivity.
Tribal inclusion carries particular weight. Tribal health agencies have historically faced chronic underfunding and acute data sovereignty concerns. By including tribal jurisdictions, PULSE acknowledges that any AI tool claiming national utility must handle specialized populations, respect distinct governance structures, and function in settings with unique environmental and social health burdens. If a model cannot perform under these conditions, it does not deserve a federal endorsement.
Territorial agencies add another necessary stress test. Public health officials in U.S. territories manage disease patterns and supply chains that differ sharply from the mainland. An AI assistant that assumes perfect internet connectivity, or that relies on ICD-10 coding habits common in large urban hospitals, will simply fail when a hurricane knocks out infrastructure. The ten-jurisdiction design forces the program to collect hard evidence on whether these models adapt to imperfect environments or fall apart.
From Chatbots to Mission-Critical Infrastructure
For the past two years, the public conversation around large language models has centered on coding shortcuts, marketing copy, and image generation. PULSE deliberately shifts the focus to mission-critical infrastructure. When a health department uses an AI model to parse infectious disease reports, a mistake is not a quirky hallucination. It is a potential public safety failure.
That reality makes this pilot a precedent-setter for three persistent technical problems: accuracy, hallucination mitigation, and data privacy. In this context, hallucination might mean a model fabricating a drug interaction, inventing a nonexistent outbreak cluster, or misstating a reporting regulation. PULSE is structured to measure how often these errors occur and whether technical guardrails can catch them before they reach a decision-maker.
Privacy carries equal weight. Public health agencies handle protected health information governed by HIPAA and additional state-level statutes. Any AI system touching this data must demonstrate exactly what it stores, where training data flows, and who can later access the outputs. The involvement of CHAI signals that these trials will test not just the raw intelligence of OpenAI and Anthropic models, but their ability to operate within strict institutional governance frameworks and leave clean audit trails.
A Blueprint for Developers and Regulators
For developers and startup founders building healthcare AI, PULSE offers something the market urgently needs: a visible, government-backed standard. Young companies often struggle to sell into public health systems because procurement officers have no common checklist for evaluating AI safety. If PULSE produces transparent criteria for measuring bias, accuracy, and privacy in administrative health settings, those criteria could quickly become the default benchmark for vendors nationwide.
The program also hints at how large language models may need to evolve to serve government. Consumer chatbots are optimized for fluent, helpful responses. Government health work requires citations, calibrated uncertainty, and institutional memory. A model advising a nurse epidemiologist must be willing to say "the evidence is unclear" rather than fabricating a confident answer. Watching how OpenAI and Anthropic adjust their prompting strategies and retrieval systems for this environment will reveal whether the underlying technology can actually meet bureaucratic expectations.
If the pilots succeed, the technical and governance insights could travel well beyond U.S. borders. Public health departments worldwide face parallel data burdens and staffing shortages. A validated framework for deploying LLMs in epidemiology and health communication could become an international reference point, much like FHIR standards did for electronic health records.
The Real Takeaway
PULSE is not a promise that artificial intelligence will fix public health. It is an admission that the old ways of processing health information are too slow and too labor-intensive, and that any replacement must be proven under realistic, varied conditions before it scales nationally. By combining CHAI's safety frameworks with frontier models from OpenAI and Anthropic, and by forcing those tools to prove themselves in ten genuinely different jurisdictions, the program treats generative AI with the skepticism it deserves while giving it the testing ground it needs. For health officials, developers, and the communities they serve, that balance is exactly what responsible adoption looks like.
