
Posted on
by AAHB
Prompt injection is when someone hides instructions inside ordinary content, and your AI assistant follows them as if they came from you. The content can be a message, a calendar invite, a document, a web page. Anything your assistant reads, it can act on. And it can’t reliably tell the difference between information you want it to use and a command an attacker has buried in that information.
That gap is not theoretical. In June 2026, security research firm SafeBreach published a study showing exactly how far it can go.
What the researchers did
SafeBreach Labs, led by researcher Or Yair, hijacked Google’s Gemini assistant with nothing more than a message. Gemini can read your notifications aloud, so the researchers hid instructions inside an ordinary WhatsApp, Slack or SMS message. When you ask Gemini to catch you up, it reads those hidden commands along with everything else and treats them as if they came from you. The attacker never touches your phone.
That let them make Gemini open smart-home windows and lights, launch a Zoom call and stream the victim’s video, and impersonate a trusted contact by announcing a fake message from the victim’s manager asking them to hand over files. They could even write false information into Gemini’s long-term memory, so the compromise followed the victim across their phone, tablet and computer.
Google had already patched an earlier version of this attack, adding a check that the user had agreed before the assistant acted. The researchers got around it. They had Gemini ask the real permission question in Chinese, or hide it in a silent on-screen link a hands-free driver never sees, then ask a harmless question aloud in English. Hearing only the English, the victim says “yes” and unknowingly authorises the hidden command. SafeBreach reported all of this privately in August 2025, and Google confirmed it was fixed in November 2025, before any of it was made public.
Why this matters beyond Gemini
It’s tempting to file this under “Google’s problem, now solved.” That would be the wrong lesson.
Google fixed these specific attacks, but the weakness underneath them is still there, because it isn’t a flaw in Gemini that a patch can remove for good. AI assistants work by reading content and acting on it, and they can’t reliably separate the information you give them from instructions a stranger has hidden inside it. Any assistant that reads your email, calendar, messages or documents has the same gap. It was Gemini this time, but it could just as easily be a tool you’ve connected to your own business.
In April 2026, Google’s own threat-intelligence team scanned billions of public web pages and found prompt injections already planted in real content. Most are still crude: pranks, SEO tricks, and a few clumsy attempts at data theft. But the malicious ones rose 32% in three months, and Google expects both the scale and the sophistication to grow as attackers automate their own attacks. What SafeBreach showed was possible, Google is already finding in the wild.
You may already have staff using AI assistants connected to a shared inbox, a CRM, a document store or a booking system. The convenience is real, and so is the risk: any content flowing into those tools is a channel for instructions you never wrote. A booking enquiry, a supplier email, an attached PDF. None of it is safe just because it looks routine.
This is not a reason to avoid AI. It’s a reason to adopt it deliberately.
What deliberate adoption looks like
The defence is the discipline of how you bring AI into the business in the first place. At AAHB Rewired, Responsible AI is one of The Foundations, the pillars a business needs in place before AI can deliver value safely. A few practical principles follow directly from what SafeBreach found.
Keep a human in the loop for anything consequential. An assistant that drafts a reply for your review is in a very different risk category from one that sends emails, moves money, or changes records on its own. The more autonomy you grant, the more an injected instruction can do unsupervised.
Be deliberate about what you connect. Every system you plug an assistant into widens the surface an attacker can reach through. Connect what earns its place, and understand what each connection lets the assistant do.
Know where your assistant’s information comes from. Content from outside your business, messages, web pages, attachments from people you don’t know, deserves more caution than content you created yourself. That distinction should shape what your tools are allowed to act on.
Set the rules before you scale, not after. An acceptable use policy that defines what staff can connect, what AI is allowed to do unsupervised, and where a human sign-off is required turns these principles into something your team actually follows.
None of this requires a security team or a technical background. It requires treating AI adoption as a process to be managed rather than a tool to be switched on. The businesses that get hurt will be the ones that connected everything to everything because it was convenient, without ever asking what their assistant was actually allowed to do.
The technology is moving quickly. Your adoption of it doesn’t have to.
Source: “Gemini’s Secret Affair: Exploiting Gemini Voice Assistant Through Instant Messaging Apps,” SafeBreach Labs, 3 June 2026.
Google: “AI threats in the wild: The current state of prompt injections on the web”
Disclosure: This article was created with the assistance of Artificial Intelligence tools to support research and outlining. While AI helped structure the piece, the final writing, views, and insights are entirely those of the author. All content has been reviewed and fact-checked by the author to ensure accuracy. The images featured in this post were generated using AI.

