Manus Prompt Injection: Warning After Code Execution
A security warning can prevent an unwanted agent action only while there is still time to stop execution. The Manus case puts that ordering problem into focus: detecting a threat after the tool has run leaves its effects in place. When designing your agent, identify the component that can actually hold an action until its checks finish.
In short
- External content needs a different trust level from the task you gave the agent.
- Required checks must finish before the corresponding action begins.
- Approval belongs to a specific tool call with specific parameters.
- Restricted permissions and isolation still matter when detection misses something.
An email becomes the entry point
On 1 October 2026, Salt Labs published its technical analysis of the Manus attack. Manus executed JSFuck-obfuscated JavaScript from an email through Node.js. The warning about the malicious content came afterwards.
According to Salt Security's disclosure, the research took place earlier in the year. The vulnerability has been fixed. Researchers found credentials for connected services in the execution environment, making the potential reach dependent on the user's integrations. This was a research demonstration; these sources do not establish an attack on actual customers.
The important boundary is between your task and somebody else's content. Asking an agent to summarize a message does not authorize its sender to direct your tools. Delivery through your own inbox does not change who supplied the text. Our prompt injection explainer covers this distinction in more detail.
Decoding needs its own execution boundary
An agent reading a document may encounter an unfamiliar representation. Opening it with a helper program can look like a reasonable step toward understanding it. If that helper evaluates untrusted code, the step becomes a separate operation with its own permissions.
Review two questions independently: what information should the agent understand, and what operation is it allowed to perform to understand it? Permission to read a document does not automatically answer the second question.
Prefer fixed parsers and decoders that handle their input as data. The text being inspected should not grant permission to launch an interpreter. If a representation genuinely requires execution to investigate, route it through a separate, constrained workflow. That boundary belongs in the application architecture, rather than depending on another sentence in a system prompt.
Approval must take effect before the tool runs
Consider an agent that summarizes incoming invoices. Reading an invoice fits the task. Sending its contents to a new recipient is a different operation. That step needs a fresh check of its destination, data and authorization.
The useful sequence is:
Receive content → inspect it → authorize the exact action → execute.
The execution component must enforce the decision. Bind approval to the tool, its parameters and the user, and make it valid for that action. If the recipient changes, the earlier approval should no longer authorize the request. Instructions embedded in an email must not substitute for an internal approval signal. The OWASP AI Agent Security Cheat Sheet describes controls for these boundaries.
Put checks where actions can be stopped
The Patronus Scanner API returns assessments for text, including tool responses. Your application decides whether to allow, block, redact or request approval. Patronus Ark supports local inference when you run the check in your own environment. Product information checked on 6 October 2026.
For an email agent, we would plan checks at two points: before message content enters the working context, and before a sensitive action derived from it is dispatched. A completed scan of a message should not grant blanket permission to use every subsequent tool.
One practical integration test is to delay the assessment deliberately and observe whether execution starts anyway. Repeat with a failed scan. If an action requires a completed check, a missing result must not silently become permission. Record the ordering of assessment completion, authorization and execution without copying unnecessary sensitive content into logs.
Classifiers need restricted permissions around them
Wolf Defender can provide a text assessment within that workflow. Its Hugging Face model card documents possible false decisions and limitations with unfamiliar obfuscation. The model evaluates text; it does not replace identity or tool-permission checks. Information checked on 6 October 2026.
We have not tested the reported Manus attack against Wolf Defender and do not claim demonstrated protection against this specific case.
Also restrict the files a tool can read and the services it can reach. Credentials for unrelated integrations should not be available together in every execution environment. Sensitive actions need an authorization decision outside the content the agent reads. The OWASP prompt injection guidance likewise recommends multiple defensive layers.
Start with one workflow: identify the external source your agent reads, the action it can trigger and the control that holds that action before execution. That gives you a concrete place to verify whether the protection works.
FAQ
Frequently asked questions
What was the Manus email prompt injection attack?
Salt Labs described a now-fixed attack in which Manus executed obfuscated code from an email. A security warning appeared after execution.
Are security warnings enough to protect AI agents?
A warning cannot prevent an action that has already happened. Sensitive tool calls need checks and, where required, approval before execution.
Would Wolf Defender have stopped the Manus attack?
That has not been demonstrated. Wolf Defender assesses text for prompt injection. Preventing an attack also depends on tool permissions, isolation and how the application enforces its decisions.