When AI Becomes an Attacker’s Ally
When we talk about cybersecurity, we almost always imagine the same scenario—someone will try to break into our information system.

It has been proven that a compromised AI agent can use legitimate tools to steal data without being detected. This could mark the beginning of a new era of cyberattacks.
For some twenty years, we have protected our IT systems from cyberattacks by investing in firewalls, antivirus systems, multi-factor authentication, intrusion detection systems, network segmentation, and penetration testing. When we talk about cybersecurity, we almost always imagine the same scenario—someone will try to break into our information system.
The artificial intelligence that we ourselves have given access to email, documents, databases, and business applications could switch sides.
I am not referring here to science-fiction scenarios or sensationalist claims about “AI taking over the world”. I am speaking about very concrete security risks that the scientific community has only just begun to investigate seriously. One such paper should concern anyone now considering the broader adoption of artificial intelligence in business.
On 7 April 2026, the research paper “Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use” was published on arXiv and subsequently accepted at one of the world’s most prestigious conferences in artificial intelligence and natural language processing—ACL 2026 (Association for Computational Linguistics).
The authors experimentally demonstrated that a compromised AI agent can use entirely legitimate tools to extract confidential information from conversation memory and send it to an attacker, while its behaviour appears entirely consistent with normal use of an AI system—could you detect such an event?
This is especially important because the attack did not exploit an operating-system vulnerability. It did not use SQL injection. It did not bypass the firewall. It did not exploit flaws in authentication.
The attack exploited the fact that the artificial intelligence was authorised to use certain tools.
In other words, the AI agent did nothing that the organisation had not formally permitted. It merely used legitimate privileges in a way no one expected.
I believe that this is precisely the greatest value of this research. It shows that we are entering an era in which protecting networks and applications will no longer be enough. We will have to protect the way artificial intelligence makes decisions.
This is an entirely new paradigm.
When an employee accesses a business application, we know their identity, assigned privileges, and the business reason for the access. When an AI agent accesses the same application, the situation becomes considerably more complex. The agent decides which tool to use, which documents to access, and in what order to perform individual tasks. It is precisely this autonomy that makes it extraordinarily useful, while at the same time creating security risks we have never faced before.
This is why I find it somewhat surprising that most of the debate today still revolves around which AI model is the “smartest”. Will it be GPT, Gemini, Claude, or an open-source model?
In my opinion, that is the wrong question.
A model is only one part of an AI system. This means that assessing the security of the AI model is no longer enough. We must assess the security of the entire AI ecosystem.
The attack surface is no longer a single piece of software. It is a chain made up of models, agents, memory, RAG systems, tools, APIs, vector databases, GPU infrastructure, and the supply chain. Compromising just one link is enough.
An enterprise AI solution today consists of a large language model (LLM), AI agents, conversation memory, RAG systems, vector databases, various business tools, APIs, a web browser, integrations with business applications, GPU infrastructure, and the entire supply chain on which all of this runs. Each of these components represents a new potential attack surface.
I am therefore convinced that an entirely new cybersecurity discipline will emerge over the next few years.
Today, we conduct information-system risk assessments.
Tomorrow, we will conduct artificial-intelligence risk assessments. In fact, we already conduct them today, but the vast majority of organisations do so at a rudimentary level. The panic over losing the race to adopt artificial intelligence further marginalises serious analysis of the dangers it brings.
Can AI agents be manipulated into making incorrect decisions or disclosing confidential information? Could a backdoor embedded in GPU firmware or drivers, an LLM, or a RAG system be exploited in our environment? We must start from the assumption that such a backdoor exists.
Today, we verify who has access to documents.
Tomorrow, we will have to prove which documents artificial intelligence can access, which tools it may use, under what conditions, and whether someone can maliciously alter its behaviour.
Our understanding and analysis of AI risks must not lag behind the development and adoption of AI.
AI agents already independently search documentation, read email, call APIs, generate reports, write code and decide which tool to use to complete a task.
With every new privilege we grant artificial intelligence, we also increase its value to a potential attacker.
This does not mean that we should slow the adoption of artificial intelligence. On the contrary, organisations that fail to use it will very quickly lose competitiveness.
It means that we can no longer view AI solely as a technology for increasing productivity.
An AI system is no longer merely an application that needs to be protected. It is becoming an active participant in business processes: we entrust it with data, base decisions on its recommendations, and grant it privileges that no software has ever had before. That is precisely why it is also becoming a new category of security risk.
We entrust it with our most valuable data and ask it to advise us on our most important business decisions.
Perhaps the real question is not whether we can trust artificial intelligence, but whether we can prove that it acted exactly as we expected—and in no other way.
The answer to that question could well define the next decade of cybersecurity.
About the Author:
Stanko Cerin has advised some of the largest companies in Croatia and around the world on cybersecurity, IT risk management and regulatory compliance for more than 20 years. He created ITrevizija.hr, Croatia’s first AI platform for NIS2 compliance, and co-founded IVERIOS, an AI-native platform for continuous cybersecurity risk and compliance management. In developing products, he continually monitors new research, attack techniques, and AI-related regulatory requirements. He is a Certified Information Systems Auditor and Certified Information Security Manager.