
Can a decision model replace your guardrail?
Overall ranking
12 systems · version 1.0.0 · 7 October 2026 · changelog
The overall score is the plain average of 6 guardrail suites, which together cover the 8 jobs. The leaders change a lot from job to job. Select your job above.

| # | System | Score ±95% CI | Tier | Catch rate | False-block rate | $ per 1,000 checks | Hosting |
|---|---|---|---|---|---|---|---|
90.4±0.7 | 1 | 7.9% | $0.056 | Managed API | |||
90.0±0.7 | 1 | 12.3% | $0.110 | Managed API | |||
89.5±0.7 | 2 | 12.5% | $0.202 | Managed API | |||
88.5±0.7 | 3 | 11.8% | $0.039 | Managed API | |||
81.9±0.9 | 4 | 16.8% | $0.076 | Managed API | |||
81.4±0.8 | 4 | 8.6% | $0.230 | Self-hosted | |||
80.3±0.8 | 5 | 9.4% | $0.124 | Self-hosted | |||
79.0±0.9 | 5 | 10.7% | $0.113 | Managed API | |||
75.4±0.9 | 6 | 12.1% | $0.124 | Self-hosted | |||
66.6±0.7 | 7 | 4.5% | $0.277 | Self-hosted | |||
66.2±0.9 | 7 | 20.3% | $0.142 | Self-hosted | |||
61.1±1.0 | 8 | 23.0% | $0.150 | Self-hosted |
Every system answers the same 9,975 checks under one fixed rule: a probability of 0.5 or more blocks. Read why. Our statistical tests cannot tell the systems in one tier apart from that tier's leader. Self-hosted cost is our shared GPU time.
- Overall score
- 90.4
- 95% interval
- 89.7 to 91.0
- Catches harmful rows
- 88.7%
- Blocks safe rows
- 7.9%
- Cost per 1,000 checks
- $0.056
- Statistical tier
- 1 of 8
pplx-decider v1 27B by job
- RAG and agentsProtect RAG and agents85.9
- Chat attacksStop jailbreaks and injection72.5
- User inputScreen what users type86.5
- Model repliesCheck what the model says89.8
- Off-topicKeep the bot on topic98.6
- Personal dataCatch personal data98.2
- GroundingCatch unsupported answers88.1
- ProfanityFilter profanity90.2
Bars start at 50, which is a coin flip. The thin mark shows the best score on that job. Select a system in the chart or the table to see its details.
MethodHow we test every system the same wayShow the methodHide the method
- 1Clean test rows9,975 checks8 guardrail jobs. We remove rows that match a model's published training data. We keep 2,269 rows private as a held-back slice.
- 2Same checks for all12 systemsEvery system answers every row. A failed call counts as wrong. We do not rank a run with more than 2% failures.
- 3One fixed rule≥ 0.5 blocksWe tune no threshold on the test rows, so you see each system's default. Verdict APIs use their own flag. Amazon Bedrock Guardrails runs at one documented setting.
- 4Score and costaccuracy and $Balanced accuracy with 95% intervals and tiers, catch rate, false blocks, and dollars per 1,000 checks.
Published guardrail numbers rarely compare. Each vendor picks its own threshold and its own test set. Here every system gets the same rows and the same rule, so the differences come from the systems. Each prompt attack has safe rows in the same style. A keyword classifier scores 50.1 to 67.7 on them. Word and character n-gram classifiers that we trained and tested on these rows score 80.1 to 87.7. Read the full method.
The leader changes with the job
- 01
Protect RAG and agents
Instructions hidden in emails, web pages and tool output
Top tier: Clef and GPT-6 Luna
Top tier blocks 4% to 8% of safe rows
1,037 rows for Protect RAG and agents89.5best score - 02
Stop jailbreaks and injection
Attacks typed directly into the chat
Top tier: Clef, Kev 4B, Clef Flash and Kev 9B
Top tier blocks 19% to 43% of safe rows
2,136 rows for Stop jailbreaks and injection74.4best score - 03
Screen what users type
Harmful requests about violence, hate, sex and crime
Top tier: Amazon Bedrock Guardrails, pplx-decider v1 27B, GPT-6 Luna and Kev 9B
Top tier blocks 14% to 19% of safe rows
1,436 rows for Screen what users type86.5best score - 04
Check what the model says
Harmful content in assistant replies
Top tier: pplx-decider v1 27B, Clef, GPT-6 Luna and Clef Flash
Top tier blocks 9% to 17% of safe rows
655 rows for Check what the model says89.8best score - 05
Keep the bot on topic
Off-limits subjects, such as investment or legal advice
Top tier: Clef, pplx-decider v1 27B, GPT-6 Luna, Kev 4B and Jev 1.13.0
Top tier blocks 2% to 6% of safe rows
601 rows for Keep the bot on topic98.8best score - 06
Catch personal data
Names, email addresses, card numbers and ID numbers
Top tier: Jev 1.13.0, Clef, pplx-decider v1 27B, Amazon Bedrock Guardrails and Clef Flash
Top tier blocks 2% to 3% of safe rows
600 rows for Catch personal data98.2best score - 07
Catch unsupported answers
Claims that the source document does not support
Top tier: Jev 1.13.0, GPT-6 Luna and pplx-decider v1 27B
Top tier blocks 1% to 18% of safe rows
556 rows for Catch unsupported answers91.2best score - 08
Filter profanity
Offensive and obscene language
Top tier: GPT-6 Luna
Top tier blocks 5% of safe rows
585 rows for Filter profanity93.6best score
Can you trust these numbers?
Who funded this, and are you tied to any vendor?
No one funded it. raxIT Labs has no commercial relationship with Amazon Web Services, Cloudflare, OpenAI, Perplexity and TypeSafe or any other company on this page. No vendor saw the data, the questions or the results before we published them.
Doesn't the question format favour TypeSafe's Jev?
It may. Every decision model gets the same yes/no questions in the format that Jev's API uses. Jev was trained on that format. The other models get the questions through our adapters. We disclose this home advantage and do not remove it, because any other format would favour a different system. Jev 1.13.0 finishes in tier 3 of 8 overall.
Did you tune thresholds for anyone?
No. Every system uses the same fixed rule: a probability of 0.5 or more blocks. This shows how each system behaves by default. Some models rank risk well but have a poor default threshold. If you tune on your own labelled data, you can gain several points. The smaller models gain the most. The results files report AUROC, so you can see how much room each model has.
Who labelled the data, and how good are the labels?
Labels come from each source and from our written labelling rules. Our lead, with an AI assistant, labelled a 400-row content sample a second time without seeing the first label. The two labels agreed on 86% of rows. An AI model with no access to the answers labelled a 400-row prompt-attack sample a second time. It agreed on 95%. After the runs, we checked again every row that at least 11 of the 12 systems got wrong. We corrected 117 labels and removed 78 ambiguous rows.
Could a model have trained on the test rows?
We compared every row with the published training data of the models on the board. We removed the matches. We also keep a held-back slice of 2,269 rows that we do not publish. Overall scores on that slice are within 2.6 points of the public rows. Some content rows come from datasets that the vendors published themselves. We also score content without those rows. GPT-6 Luna scores 85.8 on all content rows and 86.5 without OpenAI's own rows. So it did worse, not better, on its vendor's data. We cannot rule out training data that a vendor has not disclosed.
Can I reproduce this?
Yes. The dataset is on Hugging Face. The code, the scoring rules and the per-row results are on GitHub. Because of their licences, some sources ship only the row ids. A script rebuilds that text from the original publishers. The Reproduce page has the commands.
What does a tier mean?
Our statistical tests cannot tell the systems in one tier apart from that tier's leader. We resample the rows 2,000 times. We test each system against its tier's leader. Then we correct for the number of comparisons. Two systems in one tier can still differ a lot in what they block and what they cost. Use those two points to choose between them.
Does it cover images, multi-turn chats or my own policies?
Not yet. The benchmark tests text only: one message and its context. Only the off-topic job and an exact-word check cover custom policies. Multimodal and multi-turn tests are not part of this release.
What happens to the data I send these APIs?
We did not evaluate data retention or compliance. These depend on each vendor's terms and on your contract. Some vendors offer zero data retention or HIPAA support to eligible customers. Check the terms before you send regulated data.

Run it yourself
Clone the repository. Get the pinned dataset. Then score your own guardrail on the same checks. We publish new results as GitHub releases. Watch the repository to get an email for each release.