Lakera vs. Patronus: Comparing prompt injection detection

If you are searching for Lakera vs. Patronus or a Lakera alternative for prompt injection detection, two questions matter: How many attacks does a service catch? How often does it block legitimate content? Production AI applications need both measures.
We compared the Patronus Control API with the APIs from Lakera and PromptGuard.co on 825 identical single-turn inputs. On 412 deliberately difficult benign inputs, Patronus produced 5 false positives, Lakera 48, and PromptGuard 82. On a selected group of AgentPIMA and ClawTrojan texts, Lakera detected more attacks. The result shows a clear difference in false positives, alongside a limitation in Patronus attack recall on this selection.
Lakera vs. Patronus: Results at a glance
| Measure | Patronus | Lakera | PromptGuard.co |
|---|---|---|---|
| Precision, internal validation, n = 134 | 100.0% | 86.6% | 89.2% |
| Precision, selected benchmarks, n = 279 | 70.6% | 60.0% | 60.9% |
| Specificity, hard benign, n = 412 | 98.8% | 88.3% | 80.1% |
| Specificity, real benign, n = 178 | 97.2% | 88.8% | 94.9% |
Real benign is a subset of the 279 benchmark inputs. It is shown separately in the table and must not be added to the total again. The 825 unique inputs comprise 412 hard-benign cases, 279 benchmark cases, and 134 internal validation cases.
On hard benign, the specificity figures correspond to 5 false positives for Patronus, 48 for Lakera, and 82 for PromptGuard. In this sample, Patronus produced about 90 percent fewer false positives than Lakera. The Real-World Benign dataset from Rogue Security contains 178 legitimate inputs from AI coding workflows. In that subset, the APIs incorrectly flagged 5, 20, and 9 inputs, respectively.
Specificity is the share of benign inputs correctly allowed. On its own, it says nothing about attack detection. Recall must be examined separately.
Attack detection: Where Lakera leads
On our internal validation data, Patronus detected 89 of 94 attacks and correctly allowed all 40 benign cases. That gives 94.7 percent recall and 100 percent precision. Because these data are ours, we keep them separate from the external benchmark texts.
The selected benchmark group combines raw texts from AgentPIMA, ClawTrojan, and the 178 real-benign inputs. Of its 279 cases, 89 are attacks. Lakera detected 30 of those attacks (33.7 percent), PromptGuard 14 (15.7 percent), and Patronus 12 (13.5 percent). Lakera leads on recall in this selection. Patronus had the lowest false-positive rate and highest precision across the full group.
AgentPIMA and ClawTrojan were designed for multi-step agent scenarios. In our test, selected texts were sent to each API individually. This is not an evaluation of complete agent trajectories. Wolf was not trained on either dataset and did not detect any of the 36 positive ClawTrojan cases in the individual-text replay. That result is part of a fair reading of this comparison.
PromptGuard vs. Patronus: What was tested?
The figures above refer to the PromptGuard.co API, not Meta's separately named Prompt Guard model. On the same 412 hard-benign inputs, PromptGuard.co produced 82 false positives, compared with 5 for Patronus. In the selected benchmark group, PromptGuard.co detected 14 of 89 attacks, compared with 12 for Patronus.
PromptGuard.co publishes its own benchmark results on different datasets and for its full pipeline. Those numbers cannot be compared directly with this test. Nor can a sample establish how either API will perform on every production workload.
How the comparison was run
All three APIs used their default thresholds and decision rules. No threshold was tuned to this sample. For Patronus, we forced L3 Only: every input reached the Wolf model without an early L1 or L2 decision. We compared injection decisions only.
The external benchmarks were tested as selected samples, not complete suites. The practical volume was limited by the available free access. Measurements took place from September 24 to 27, 2026. Model versions, configurations, and vendor APIs can change. This is a comparison measured by Patronus, not independent certification.
Latency: Separate backend time from HTTP time
Across 1,057 distinct single-turn inputs, the Patronus backend reported 12.0 ms p50 and 57.3 ms p95 scan time (total_ms) on the forced L3 path. Complete HTTP requests on that set took 167.8 ms p50 and 427.7 ms p95. The 12 ms figure is therefore not end-to-end client latency.
For the 825 shared inputs, median HTTP latency was 93.3 ms for Lakera, 167.4 ms for Patronus, and 21,923.6 ms for PromptGuard.co. Vendors were measured at different times. Lakera responded faster than Patronus in this run. The observed PromptGuard response times do not establish its typical production latency or prove active throttling.
Patronus as a Lakera alternative
According to its product site, Lakera offers a broader AI security platform for applications, agents, and workplace use. Patronus is a Lakera alternative when the priority is prompt injection detection with few false positives and the option to operate the detector yourself.
Wolf Defender Small is an open Apache 2.0 model with ONNX variants. Local checks need no external scan API call or monthly API allowance. The API results above, however, apply to the hosted Patronus configuration. They are not a fresh evaluation of every published quantized model variant.
The graphic's "No account rate limit" note refers only to the Control test account used here: it processed 1,057 examples without an account-specific quota stop. The test runner itself used two parallel requests and no more than two starts per second. This is not a promise of unlimited throughput for other accounts. Lakera reached its free allowance during this run; PromptGuard.co returned no recorded HTTP 429 response.
Conclusion
In this measurement, Lakera vs. Patronus is largely a trade-off between attack recall and false positives. Patronus allowed more legitimate inputs in both benign groups and reached high precision on its internal validation set. Lakera detected more attacks in the selected agent benchmark texts and had the shorter HTTP response time on the shared inputs. PromptGuard vs. Patronus likewise shows a substantial false-positive difference in this sample, not a universal ranking for every workload.
To evaluate the model directly, try Wolf Defender Small or explore the Patronus AI models. A test on your own inputs, with matching decision rules and operating conditions, remains the most useful comparison.
FAQ
Frequently asked questions
Is Patronus an alternative to Lakera?
Yes, if you need prompt injection detection for your AI applications. Patronus produced fewer false positives on benign content in the test described here. Lakera also offers a broader AI security platform. Wolf Defender is additionally available as an open model for local deployment.
Which detects more attacks, Lakera or Patronus?
It depends on the dataset. Patronus detected 89 of 94 attacks in its own validation data. In the selected AgentPIMA and ClawTrojan benchmark group, Lakera detected 30 of 89 attacks and Patronus detected 12. These groups should be evaluated separately.
How does PromptGuard vs. Patronus compare?
On 412 shared hard-benign inputs, the tested PromptGuard.co API flagged 82 false positives and Patronus flagged 5. In the selected benchmark group, PromptGuard detected 14 of 89 attacks and Patronus detected 12. This test concerns PromptGuard.co, not Meta's separate Prompt Guard model.
Is 12 ms the complete Patronus API latency?
No. The 12.0 ms median is the scan time reported by the Patronus backend. The complete HTTP request had a 167.8 ms median in the same run. The test forced the Wolf model path without an early L1 or L2 decision.