A single click on the wrong link no longer just steals passwords or installs malware. It can birth a persistent AI agent inside your organization that reads your emails, rifles through your files, and chats with your coworkers under your name. Security researchers at Zenity Labs have demonstrated exactly this scenario with a technique they call AgentForger, and the mechanics reveal a disturbing gap in how we secure autonomous AI tools.
From CSRF to AgentForger: A New Class of Attack
Most security professionals know Cross-Site Request Forgery, or CSRF, as a classic web vulnerability. In its traditional form, an attacker tricks an authenticated user’s browser into firing off a single unauthorized request. Maybe it changes an account email address or disables two-factor authentication. The damage is usually contained to one action.
AgentForger blows that model apart. Rather than forging one request, it forges an entire autonomous agent inside OpenAI’s ChatGPT Workspace Agents. The attack starts at the ChatGPT Agent Builder, which lives at chatgpt.com/agents/studio/new. Under normal circumstances, building an agent is a hands-on process. You select a template, you review which tools the agent can touch, and you set guardrails so the machine asks before it reads your inbox or writes to your drive. The workflow assumes a human is paying attention at each step.
Zenity Labs discovered that these safeguards evaporate when an attacker controls the URL. By embedding specific query parameters—namely template_name and initial_assistant_prompt—a malicious link can pre-populate the entire creation form. When the victim clicks, the initial_assistant_prompt is submitted and executed automatically. The template selection, the tool review, and the permission dialogs never appear. The agent simply spawns in the background, configured exactly the way the attacker wants it, while the victim may only see an ordinary-looking ChatGPT page load and then close the tab.
This is not a traditional malware infection. Nothing gets downloaded to the laptop. The threat lives entirely inside the victim’s SaaS session, operating with the full legitimacy of the platform behind it.
Hijacking Trust Instead of Credentials
The real brutality of AgentForger lies in how it handles permissions. Enterprise users of ChatGPT Workspace routinely link their AI agents to business applications through OAuth. They authenticate once to connect Gmail, Outlook, Slack, Google Drive, or Microsoft Teams, and from then on the agent can query those services on their behalf.
AgentForger inherits all of this access automatically. Because the victim is already logged into ChatGPT and has already established those OAuth connections, the forged agent reuses them. The victim does not see a fresh consent screen asking whether this new agent should be allowed into their email. The system treats the agent as an extension of the user.
In Zenity’s demonstration, the attacker crafted the malicious URL to do something even more aggressive. It instructed the Agent Builder to flip every permission toggle for reading, writing, and deleting to "Never ask." Normally, the platform would prompt a human before executing sensitive actions—opening a spreadsheet, sending a message, deleting a calendar event. By forcing these controls to silent mode, the attacker removes the human-in-the-loop entirely. The agent becomes a ghost in the account, free to move through connected apps without generating a single alert or confirmation dialog.
A Command Channel Hiding in Plain Sight
Persistence turns a nasty exploit into an operational nightmare, and AgentForger achieves persistence through ChatGPT’s built-in scheduling features. The attacker sets multiple staggered schedules during the creation process, producing an agent that wakes up and runs as often as every five minutes.
Here is how the command-and-control loop works. The attacker sends an email to the victim’s Outlook inbox. The subject line contains a trigger word, such as "TASK." Every five minutes, the rogue agent scans the inbox looking for that trigger. When it finds a match, it opens the email, reads the instructions, executes them using the victim’s authenticated applications, and sends the results back to the attacker.
Chính email doanh nghiệp của nạn nhân trở thành hạ tầng điều khiển (command-and-control). Không có tín hiệu DNS (DNS beacon) khả nghi, không có kết nối tới địa chỉ IP lạ ở Đông Âu, cũng không có tệp thực thi mã độc (malware binary) để các công cụ phát hiện tại điểm cuối (endpoint detection) gắn cờ. Lưu lượng truy cập chạy qua các API của Microsoft hoặc Google bằng các thông tin xác thực hoàn toàn hợp lệ. Đối với một đội ngũ vận hành an ninh (security operations team) đang giám sát nhật ký mạng, điều này trông giống như một nhân viên bận rộn đang sử dụng các công cụ SaaS đã được phê duyệt.
Tác động thực tế: Sơ đồ hóa tổ chức và vũ khí hóa sự tin cậy
Zenity Labs đã thử nghiệm các giới hạn thực tế của cuộc tấn công này trong môi trường được kiểm soát, và kết quả này chắc chắn sẽ khiến bất kỳ CISO nào cũng phải lo ngại.
Chỉ bằng một chỉ thị duy nhất được gửi qua trình kích hoạt email (email trigger), các nhà nghiên cứu đã yêu cầu agent thực hiện sơ đồ hóa tổ chức. Nó truy vấn Slack, Microsoft Teams và SharePoint. Nó trả về tên các kênh, cấu trúc nhóm và các kho lưu trữ tệp. Agent đã tìm thấy các tài liệu nhạy cảm, bao gồm các bản điều khoản M&A và dữ liệu lương của nhân viên. Mọi lượt truy cập đều được ghi lại dưới danh tính hợp lệ của nạn nhân, khiến việc phát hiện pháp chứng (forensic detection) trở thành vấn đề phân loại giữa "nhiễu" từ người dùng bình thường và "nhiễu" từ người dùng độc hại.
Tiếp theo là tiềm năng về kỹ thuật thao túng tâm lý (social engineering). Vì agent có thể gửi tin nhắn thông qua tài khoản Slack hoặc Teams chính thức của nạn nhân, nó có thể thực hiện các chiến dịch lừa đảo (phishing) nội bộ, vượt qua hoàn toàn các cổng email bên ngoài và các kiểm tra DMARC. Hãy tưởng tượng một tin nhắn trực tiếp từ một đồng nghiệp đáng tin cậy yêu cầu bạn xác nhận việc triển khai SSO sắp tới hoặc nhấp vào một liên kết để kiểm tra cổng phúc lợi mới. Tin nhắn mang tên, ảnh đại diện và ngữ cảnh lịch sử trò chuyện của người đồng nghiệp đó. Nó xuất hiện ngay trong luồng hội thoại mà bạn đã thảo luận về kế hoạch ăn trưa ngày hôm qua. Mức độ tin cậy đó khác biệt hoàn toàn so với một tên miền giả mạo (spoofed domain) hay một địa chỉ người gửi viết sai chính tả, và AgentForger khai thác điều đó một cách không nương tay.
Bài học rút ra thực sự
AgentForger phơi bày một vấn đề mang tính cấu trúc trong cách các nền tảng AI doanh nghiệp kế thừa sự tin cậy. Chúng ta đang xây dựng các agent tự trị có khả năng tự lập lịch trình, viết mã và truy vấn dữ liệu nhạy cảm, nhưng chúng ta vẫn đang sử dụng các mô hình phân quyền được thiết kế cho các ứng dụng web tĩnh, nơi con người phải nhấp vào mọi nút bấm. Ranh giới giữa “những gì người dùng làm” và “những gì agent của người dùng làm” đã sụp đổ, và những kẻ tấn công hiện đang ở vị thế có thể khai thác sự sụp đổ đó trên quy mô lớn.
Các tổ chức đang sử dụng ChatGPT Workspace cần coi việc tạo agent là một hoạt động đặc quyền (privileged operation), chứ không phải là một quy trình làm việc thông thường. Điều đó có nghĩa là phải kiểm tra (auditing) các phạm vi OAuth (OAuth scopes) để đảm bảo các agent không thể âm thầm truy cập toàn bộ hộp thư đến hoặc kho lưu trữ tệp. Điều đó có nghĩa là phải giám sát các đợt hoạt động bùng phát trông giống như các vòng lặp được lập lịch thay vì nhịp độ của con người. Và điều đó có nghĩa là phải nhận ra rằng vụ vi phạm lớn tiếp theo có thể không bắt đầu bằng một mật khẩu bị lộ hay một email lừa đảo. Nó có thể bắt đầu chỉ bằng một cú nhấp chuột mất tập trung, âm thầm ủy quyền danh tính của bạn cho một cỗ máy không bao giờ ngủ.
