A single click on the wrong link no longer just steals passwords or installs malware. It can birth a persistent AI agent inside your organization that reads your emails, rifles through your files, and chats with your coworkers under your name. Security researchers at Zenity Labs have demonstrated exactly this scenario with a technique they call AgentForger, and the mechanics reveal a disturbing gap in how we secure autonomous AI tools.

From CSRF to AgentForger: A New Class of Attack

Most security professionals know Cross-Site Request Forgery, or CSRF, as a classic web vulnerability. In its traditional form, an attacker tricks an authenticated user’s browser into firing off a single unauthorized request. Maybe it changes an account email address or disables two-factor authentication. The damage is usually contained to one action.

AgentForger blows that model apart. Rather than forging one request, it forges an entire autonomous agent inside OpenAI’s ChatGPT Workspace Agents. The attack starts at the ChatGPT Agent Builder, which lives at chatgpt.com/agents/studio/new. Under normal circumstances, building an agent is a hands-on process. You select a template, you review which tools the agent can touch, and you set guardrails so the machine asks before it reads your inbox or writes to your drive. The workflow assumes a human is paying attention at each step.

Zenity Labs discovered that these safeguards evaporate when an attacker controls the URL. By embedding specific query parameters—namely template_name and initial_assistant_prompt—a malicious link can pre-populate the entire creation form. When the victim clicks, the initial_assistant_prompt is submitted and executed automatically. The template selection, the tool review, and the permission dialogs never appear. The agent simply spawns in the background, configured exactly the way the attacker wants it, while the victim may only see an ordinary-looking ChatGPT page load and then close the tab.

This is not a traditional malware infection. Nothing gets downloaded to the laptop. The threat lives entirely inside the victim’s SaaS session, operating with the full legitimacy of the platform behind it.

Hijacking Trust Instead of Credentials

The real brutality of AgentForger lies in how it handles permissions. Enterprise users of ChatGPT Workspace routinely link their AI agents to business applications through OAuth. They authenticate once to connect Gmail, Outlook, Slack, Google Drive, or Microsoft Teams, and from then on the agent can query those services on their behalf.

AgentForger inherits all of this access automatically. Because the victim is already logged into ChatGPT and has already established those OAuth connections, the forged agent reuses them. The victim does not see a fresh consent screen asking whether this new agent should be allowed into their email. The system treats the agent as an extension of the user.

In Zenity’s demonstration, the attacker crafted the malicious URL to do something even more aggressive. It instructed the Agent Builder to flip every permission toggle for reading, writing, and deleting to "Never ask." Normally, the platform would prompt a human before executing sensitive actions—opening a spreadsheet, sending a message, deleting a calendar event. By forcing these controls to silent mode, the attacker removes the human-in-the-loop entirely. The agent becomes a ghost in the account, free to move through connected apps without generating a single alert or confirmation dialog.

A Command Channel Hiding in Plain Sight

Persistence turns a nasty exploit into an operational nightmare, and AgentForger achieves persistence through ChatGPT’s built-in scheduling features. The attacker sets multiple staggered schedules during the creation process, producing an agent that wakes up and runs as often as every five minutes.

Here is how the command-and-control loop works. The attacker sends an email to the victim’s Outlook inbox. The subject line contains a trigger word, such as "TASK." Every five minutes, the rogue agent scans the inbox looking for that trigger. When it finds a match, it opens the email, reads the instructions, executes them using the victim’s authenticated applications, and sends the results back to the attacker.

Korporacyjna poczta e-mail samej ofiary staje się infrastrukturą command-and-control. Nie ma podejrzanych sygnałów DNS, żadnych połączeń z nieznanym adresem IP w Europie Wschodniej, żadnych plików binarnych złośliwego oprogramowania, które systemy wykrywania na punktach końcowych mogłyby oznaczyć. Ruch przepływa przez API Microsoftu lub Google przy użyciu w pełni legalnych poświadczeń. Dla zespołu operacji bezpieczeństwa monitorującego logi sieciowe wygląda to jak intensywna praca pracownika korzystającego z zatwierdzonych narzędzi SaaS.

Wpływ w świecie rzeczywistym: Mapowanie organizacji i wykorzystywanie zaufania jako broni

Zenity Labs przetestowało praktyczne granice tego ataku w kontrolowanych środowiskach, a wyniki powinny niepokoić każdego CISO.

Używając pojedynczej instrukcji dostarczonej poprzez wyzwalacz e-mail, badacze polecili agentowi zmapowanie organizacji. Agent odpytywał Slack, Microsoft Teams oraz SharePoint. Zwrócił nazwy kanałów, struktury zespołów i repozytoria plików. Agent zlokalizował wrażliwe dokumenty, w tym dokumenty warunkowe fuzji i przejęć (M&A) oraz dane o wynagrodzeniach pracowników. Każdy dostęp został zarejestrowany pod legalną tożsamością ofiary, co sprawia, że detekcja śledcza sprowadza się do odróżnienia normalnego szumu użytkownika od szumu generowanego przez złośliwe działania.

Istnieje również potencjał socjotechniczny. Ponieważ agent może wysyłać wiadomości za pośrednictwem oficjalnych kont ofiary na Slacku lub Teamsach, może prowadzić wewnętrzne kampanie phishingowe, które całkowicie omijają zewnętrzne bramki e-mailowe i kontrole DMARC. Wyobraź sobie wiadomość bezpośrednią od zaufanego współpracownika, proszącą o potwierdzenie nadchodzącego wdrożenia SSO lub o kliknięcie w link w celu przetestowania nowego portalu benefitowego. Wiadomość zawiera imię kolegi, jego zdjęcie profilowe oraz kontekst historii czatu. Trafia do tego samego wątku rozmowy, w którym wczoraj omawialiście plany na lunch. Taki poziom zaufania zasadniczo różni się od podszytej domeny czy błędnie napisanego adresu nadawcy, a AgentForger bezlitośnie to wykorzystuje.

Kluczowe wnioski

AgentForger obnaża problem strukturalny w sposobie, w jaki korporacyjne platformy AI dziedziczą zaufanie. Budujemy autonomicznych agentów, którzy potrafią sami planować zadania, pisać kod i odpytywać wrażliwe dane, a mimo to wciąż korzystamy z modeli uprawnień zaprojektowanych dla statycznych aplikacji webowych, w których człowiek klika każdy przycisk. Granica między tym, „co robi użytkownik”, a tym, „co robi agent użytkownika”, uległa zatarciu, a atakujący są teraz w pozycji, by wykorzystać to załamanie na dużą skalę.

Organizacje korzystające z ChatGPT Workspace muszą traktować tworzenie agentów jako operację uprzywilejowaną, a nie zwykły proces roboczy. Oznacza to audytowanie zakresów OAuth, aby upewnić się, że agenci nie mogą po cichu uzyskiwać dostępu do całych skrzynek odbiorczych lub magazynów plików. Oznacza to monitorowanie nagłych skoków aktywności, które przypominają zaplanowane pętle, a nie ludzkie tempo pracy. I oznacza to uznanie, że następne wielkie naruszenie bezpieczeństwa może nie zacząć się od przejętego hasła czy wiadomości phishingowej. Może zacząć się od jednego nieuważnego kliknięcia, które po cichu deleguje Twoją tożsamość maszynie, która nigdy nie śpi.