My first two write-ups circled the same quiet idea without ever naming it: security lives on a seam between data and instructions, and most of defence is just deciding what you are willing to trust. One post was about an AI agent walking an attacker across localhost — a boundary that turned out to describe geography rather than trust. The other was about reading a third of a million failed logins and learning that defence is structural: reduce what is reachable, then watch. This month a new story made me realise those two posts were about the same crack, and that I had still drawn my trust line in the wrong place.
In mid-June 2026, a startup called Tenet Security disclosed an attack it named Agentjacking (Tenet Security; covered by The Hacker News and The New Stack). The setup is almost insultingly simple. Sentry, the error-tracking tool that sits inside a huge share of production apps, accepts error events from anyone who holds the project’s DSN, and that DSN is a public key that ships in the front-end source of the very sites it monitors. Separately, teams now wire Sentry into their AI coding agents through an MCP server, so the agent can read recent errors and help fix them. Put those two facts together and an attacker can post a fake error whose text is written to look exactly like Sentry’s own “suggested resolution steps.” The agent, whether that is Claude Code, Cursor or Codex, reads it as trusted diagnostic output and runs the attacker’s commands. No exploit, no memory corruption. Just text the machine was told to believe.
I want to be precise about scale, because the numbers are the unsettling part, not the mechanism. Tenet found 2,388 organisations exposing injectable DSNs through nothing but passive reconnaissance, 71 of them in the Tranco top-1M list of busiest sites. Across controlled tests, more than 100 real AI coding agents acted on the injected errors, with a reported 85% exploitation success rate. That number is Tenet’s own, from a vendor selling the fix for it, and the methodology is not published. I use it as a direction, not a measurement. A successful run hands over environment variables, Git credentials, private repository URLs, which is the developer’s keys to everything. Tenet disclosed to Sentry on 3 June 2026; Sentry added a content filter for the proof-of-concept payload, and Tenet shipped hardening configs it calls agent-jackstop. As a single bug, it is being closed. As a pattern, it has barely started.
Here is what I had wrong, and I suspect I am not alone. After AutoJack I had updated my mental model to distrust what the agent browses: web pages, the untrusted internet. Agentjacking distrusts something I would never have put on the list: my own error log. For decades a stack trace was the most trusted text in engineering, because it had one defining property. We wrote it for ourselves. We did name this once. Log injection is CWE-117, and it has been in OWASP guidance since the 2000s. But CWE-117 assumes the victim is a person reading a log viewer, misled about what happened. What is new is a reader that does not merely display the line. It runs it. An AI agent breaks that property twice over: it reads that “internal” data on our behalf, and it can act on it. The seam from my first post is back, but the dangerous input is not the obvious hostile web page. It is the operational exhaust we generate and trust by reflex: logs, errors, tickets, CI output, monitoring. The real attack surface of an agent is everything it reads in order to be helpful.
What makes this land for me is the symmetry with my own server logs. In that earlier post a log line was a signal, something I read to notice an attack and harden the door. Agentjacking is the same artifact, a log, read instead by a machine that can pull a trigger. The exact same data is a defence when a careful human reads it and a weapon when an eager deputy can execute it. The thing that changed is not the data; it is that we handed the reading, and the hands, to something that cannot tell a description of an action from an instruction to take it.
This is not an isolated research toy, either. A month earlier Google’s threat team described the first cybercrime case where an actor used an AI model to help discover and weaponise a zero-day — a 2FA-bypass whose code carried the tells of machine authorship, down to a hallucinated CVSS score (Google Threat Intelligence, via SecurityWeek). Offence is automating the expensive half of an attack, while defence is wiring fast, trusting agents into the centre of its own toolchain. The two trends point straight at each other.
Where does that leave the person on the other side? You cannot patch the fact that the agent believed its inputs. That is the design. You move the boundary outward to something that does not read English: least privilege and strict allowlists on every command an agent can run; treat all ingested data as untrusted even when you generated it; require real authentication on the channels that feed an agent. Sentry’s filter and Tenet’s agent-jackstop help, but they bound the blast radius rather than removing it. The cleanest rule I have found is still Meta’s Agents Rule of Two: never let one agent session hold untrusted input, private data, and the ability to act, all at once (Meta). Agentjacking is precisely what happens when it holds all three.
Nothing new is required to think about this. The oldest mistake in security, confusing input with instructions, came back wearing the friendliest possible disguise: a helpful assistant reading my own bug reports. My first post trusted a door because of where it stood; my second learned to watch the door; this one is about realising the agent will happily read a note slipped under it and do what the note says. The walls are the ones we have always had to build by hand. What is new is that we finally have to build them around the data we wrote for ourselves.
Sources / Nguồn:
- Tenet Security — Agentjacking disclosure (primary): tenetsecurity.ai
- The Hacker News — “Agentjacking Attack Tricks AI Coding Agents Into Running Malicious Code” (June 2026): thehackernews.com
- The New Stack — “A public Sentry key is all it takes to hijack Claude Code, Cursor, and Codex”: thenewstack.io
- Google Threat Intelligence Group — AI-assisted vulnerability exploitation (12 May 2026): cloud.google.com
- SecurityWeek — “Google Detects First AI-Generated Zero-Day Exploit” (11 May 2026): securityweek.com
- Meta — “Agents Rule of Two”, practical AI agent security: ai.meta.com
Hai bài viết đầu của tôi cứ xoay quanh một ý tưởng âm thầm mà chưa từng gọi tên: an ninh sống trên đường nối giữa dữ liệu và lệnh, và phần lớn việc phòng thủ chỉ là quyết định xem mình sẵn lòng tin vào điều gì. Một bài nói về một AI agent dẫn kẻ tấn công vượt qua localhost, một ranh giới hóa ra chỉ mang tính vị trí mạng, chứ không đáng tin. Bài kia nói về việc đọc một phần ba triệu lần đăng nhập thất bại và rút ra rằng phòng thủ mang tính nền tảng: thu hẹp những gì có thể với tới, rồi quan sát. Tháng này, một câu chuyện mới khiến tôi nhận ra hai bài đó nói về cùng một vết nứt, và rằng tôi vẫn vạch lằn ranh tin tưởng sai chỗ.
Giữa tháng 6 năm 2026, một startup tên Tenet Security công bố một đòn tấn công họ gọi là Agentjacking (Tenet Security; được The Hacker News và The New Stack đưa tin). Bối cảnh đơn giản đến mức khó tin. Sentry, công cụ theo dõi lỗi nằm trong vô số ứng dụng đang chạy thật, nhận sự kiện lỗi từ bất kỳ ai có DSN của dự án, mà DSN ấy lại là một khóa công khai nằm ngay trong mã front-end của chính những trang nó giám sát. Mặt khác, các nhóm giờ nối Sentry vào AI coding agent qua một máy chủ MCP, để agent đọc các lỗi gần đây và giúp sửa. Ghép hai điều đó lại, kẻ tấn công có thể gửi một báo lỗi giả mà nội dung được viết trông y hệt “các bước khắc phục gợi ý” của chính Sentry. Agent, Claude Code, Cursor, Codex, đọc nó như đầu ra chẩn đoán đáng tin và chạy lệnh của kẻ tấn công. Không khai thác lỗ hổng bộ nhớ, không gì cả. Chỉ là văn bản mà cỗ máy được bảo phải tin.
Tôi muốn nói chính xác về quy mô, vì chính các con số mới đáng sợ chứ không phải cơ chế. Tenet phát hiện 2.388 tổ chức để lộ DSN có thể bị tiêm, chỉ bằng dò quét thụ động, 71 trong số đó nằm trong danh sách Tranco top-1M các trang đông khách nhất. Qua các đợt kiểm thử có kiểm soát, hơn 100 AI coding agent thật đã hành động theo các lỗi bị tiêm, với tỉ lệ khai thác thành công được báo là 85%. Con số ấy do chính Tenet công bố, mà Tenet thì bán bản vá cho đúng thứ họ đo, và cách đo không được công khai. Tôi đọc nó như một hướng, không phải một phép đo. Một lần thành công là trao đi biến môi trường, thông tin đăng nhập Git, đường dẫn kho mã riêng tư, chìa khóa vào mọi thứ của lập trình viên. Tenet báo cho Sentry ngày 3 tháng 6 năm 2026; Sentry thêm bộ lọc nội dung cho payload mẫu, còn Tenet phát hành cấu hình gia cố tên agent-jackstop. Như một lỗi đơn lẻ, nó đang được vá. Như một quy luật, nó mới chỉ bắt đầu.
Đây là điều tôi đã hiểu sai, và tôi ngờ rằng không chỉ mình tôi. Sau AutoJack, tôi đã cập nhật mô hình tư duy thành “nghi ngờ những gì agent duyệt web”, các trang web, cái Internet không đáng tin. Agentjacking nghi ngờ một thứ tôi sẽ không bao giờ đưa vào danh sách: log lỗi của chính tôi. Suốt nhiều thập kỷ, một stack trace là đoạn văn bản đáng tin nhất trong nghề kỹ thuật, vì nó có một đặc tính cốt lõi: chính ta viết ra nó, cho ta đọc. Ngành này từng đặt tên cho chuyện đó rồi. Log injection, tức CWE-117, đã nằm trong hướng dẫn của OWASP từ những năm 2000. Nhưng CWE-117 giả định nạn nhân là một con người đang đọc log và bị đánh lừa về chuyện đã xảy ra. Cái mới ở đây là một người đọc không chỉ hiển thị dòng log. Nó chạy dòng log. Một AI agent phá vỡ đặc tính ấy theo hai cách: nó đọc thay ta cái dữ liệu “nội bộ” đó, và nó có thể hành động dựa trên đó. Đường nối ở bài đầu của tôi quay lại, nhưng đầu vào nguy hiểm không phải trang web độc hại lộ liễu. Mà là “khí thải vận hành” ta tạo ra và tin theo phản xạ: log, lỗi, ticket, đầu ra CI, giám sát. Bề mặt tấn công thật sự của một agent là tất cả những gì nó đọc để tỏ ra hữu ích.
Điều khiến tôi thấm nhất là sự đối xứng với chính log máy chủ của tôi. Trong bài trước, một dòng log là tín hiệu, thứ tôi đọc để nhận ra đòn tấn công và gia cố cánh cửa. Agentjacking cũng là vật đó, một dòng log, nhưng được một cỗ máy có thể bóp cò đọc thay. Cùng một dữ liệu y hệt: là phòng thủ khi một con người cẩn thận đọc nó, và là vũ khí khi một “kẻ được ủy quyền” hăng hái có thể thực thi nó. Thứ thay đổi không phải dữ liệu; mà là ta đã giao việc đọc, và đôi tay, cho một thứ không phân biệt được “mô tả một hành động” với “lệnh thực hiện hành động đó”.
Đây cũng không phải một món đồ chơi nghiên cứu đơn lẻ. Một tháng trước đó, đội tình báo mối đe dọa của Google mô tả vụ tội phạm mạng đầu tiên mà kẻ tấn công dùng một mô hình AI để hỗ trợ phát hiện và vũ khí hóa một zero-day, một lỗ hổng vượt 2FA, mã của nó mang đậm dấu vết do máy viết, tới mức có cả điểm CVSS bịa ra (Google Threat Intelligence, qua SecurityWeek). Bên tấn công đang tự động hóa nửa tốn kém nhất của một đòn đánh, trong khi bên phòng thủ thì cắm những agent nhanh và cả tin vào ngay giữa bộ công cụ của mình. Hai xu hướng đang chĩa thẳng vào nhau.
Vậy người phòng thủ thực sự làm gì? Bạn không thể vá “agent đã tin đầu vào của nó”, đó là thiết kế. Bạn đẩy ranh giới ra ngoài, tới chỗ một thứ không đọc tiếng Anh: đặc quyền tối thiểu và allowlist nghiêm ngặt cho mọi lệnh agent có thể chạy; coi mọi dữ liệu được nạp vào là không đáng tin kể cả khi chính bạn tạo ra; bắt buộc xác thực thật trên các kênh đưa dữ liệu cho agent. Bộ lọc của Sentry và agent-jackstop của Tenet có ích, nhưng chỉ giới hạn “bán kính sát thương”, không xóa bỏ nó. Nguyên tắc gọn gàng nhất tôi tìm thấy vẫn là Agents Rule of Two của Meta: đừng bao giờ để một phiên agent nắm cùng lúc đầu vào không tin cậy, dữ liệu riêng tư, và khả năng hành động (Meta). Agentjacking chính là điều xảy ra khi nó nắm cả ba.
Phần khiêm nhường, đúng phần mà AutoJack để lại cho tôi, là chẳng cần gì mới để tư duy về nó. Sai lầm cổ xưa nhất trong an ninh, lẫn lộn đầu vào với lệnh, quay lại trong lớp ngụy trang thân thiện nhất có thể: một trợ lý hữu ích đang đọc chính các báo lỗi của tôi. Bài đầu của tôi tin một cánh cửa vì nó đứng ở đâu; bài thứ hai học cách canh cửa; bài này là khoảnh khắc nhận ra agent sẽ vui vẻ đọc mẩu giấy ai đó nhét qua khe cửa và làm đúng những gì mẩu giấy bảo. Những bức tường vẫn là thứ ta luôn phải tự tay xây. Cái mới là cuối cùng ta phải xây chúng quanh chính dữ liệu mà ta viết ra cho mình.
Sources / Nguồn:
- Tenet Security — công bố Agentjacking (nguồn gốc): tenetsecurity.ai
- The Hacker News — “Agentjacking Attack Tricks AI Coding Agents Into Running Malicious Code” (6/2026): thehackernews.com
- The New Stack — “A public Sentry key is all it takes to hijack Claude Code, Cursor, and Codex”: thenewstack.io
- Google Threat Intelligence Group — khai thác lỗ hổng có AI hỗ trợ (12/5/2026): cloud.google.com
- SecurityWeek — “Google Detects First AI-Generated Zero-Day Exploit” (11/5/2026): securityweek.com
- Meta — “Agents Rule of Two”, an ninh agent thực dụng: ai.meta.com