Skip to content
raxIT
raxIT Labs · Independent benchmark · v1.0.0

Can a decision model replace your guardrail?

Leaderboard

Overall ranking

12 systems · version 1.0.0 · 7 October 2026 · changelog

The overall score is the plain average of 6 guardrail suites, which together cover the 8 jobs. The leaders change a lot from job to job. Select your job above.

Score is balanced accuracy, where 50 is a coin flip. The horizontal axis is USD per 1,000 checks. Lines show the 95% interval. Squares that would overlap are moved apart slightly; the table below gives every number.Tier 1Tier 2Tier 3Tier 4Tier 5+of 8
Overall: rank, score with 95% interval, tier, catch rate, false-block rate, cost and hosting for every system.
SystemScore ±95% CIFalse-block rate$ per 1,000 checks
90.4±0.7
7.9%$0.056
90.0±0.7
12.3%$0.110
89.5±0.7
12.5%$0.202
88.5±0.7
11.8%$0.039
81.9±0.9
16.8%$0.076
81.4±0.8
8.6%$0.230
80.3±0.8
9.4%$0.124
79.0±0.9
10.7%$0.113
75.4±0.9
12.1%$0.124
66.6±0.7
4.5%$0.277
66.2±0.9
20.3%$0.142
61.1±1.0
23.0%$0.150

Every system answers the same 9,975 checks under one fixed rule: a probability of 0.5 or more blocks. Read why. Our statistical tests cannot tell the systems in one tier apart from that tier's leader. Self-hosted cost is our shared GPU time.

pplx-decider v1 27BPerplexity · managed API
Overall score
90.4
95% interval
89.7 to 91.0
Catches harmful rows
88.7%
Blocks safe rows
7.9%
Cost per 1,000 checks
$0.056
Statistical tier
1 of 8
See the rows it got wrong

pplx-decider v1 27B by job

  • RAG and agents85.9
  • Chat attacks72.5
  • User input86.5
  • Model replies89.8
  • Off-topic98.6
  • Personal data98.2
  • Grounding88.1
  • Profanity90.2

Bars start at 50, which is a coin flip. The thin mark shows the best score on that job. Select a system in the chart or the table to see its details.

MethodHow we test every system the same wayShow the method
  1. 1Clean test rows9,975 checks8 guardrail jobs. We remove rows that match a model's published training data. We keep 2,269 rows private as a held-back slice.
  2. 2Same checks for all12 systemsEvery system answers every row. A failed call counts as wrong. We do not rank a run with more than 2% failures.
  3. 3One fixed rule≥ 0.5 blocksWe tune no threshold on the test rows, so you see each system's default. Verdict APIs use their own flag. Amazon Bedrock Guardrails runs at one documented setting.
  4. 4Score and costaccuracy and $Balanced accuracy with 95% intervals and tiers, catch rate, false blocks, and dollars per 1,000 checks.

Published guardrail numbers rarely compare. Each vendor picks its own threshold and its own test set. Here every system gets the same rows and the same rule, so the differences come from the systems. Each prompt attack has safe rows in the same style. A keyword classifier scores 50.1 to 67.7 on them. Word and character n-gram classifiers that we trained and tested on these rows score 80.1 to 87.7. Read the full method.

By job

The leader changes with the job

See all 7,606 public rows
  1. 01

    Protect RAG and agents

    Instructions hidden in emails, web pages and tool output

    Top tier: Clef and GPT-6 Luna

    Top tier blocks 4% to 8% of safe rows

    1,037 rows for Protect RAG and agents
    89.5best score
  2. 02

    Stop jailbreaks and injection

    Attacks typed directly into the chat

    Top tier: Clef, Kev 4B, Clef Flash and Kev 9B

    Top tier blocks 19% to 43% of safe rows

    2,136 rows for Stop jailbreaks and injection
    74.4best score
  3. 03

    Screen what users type

    Harmful requests about violence, hate, sex and crime

    Top tier: Amazon Bedrock Guardrails, pplx-decider v1 27B, GPT-6 Luna and Kev 9B

    Top tier blocks 14% to 19% of safe rows

    1,436 rows for Screen what users type
    86.5best score
  4. 04

    Check what the model says

    Harmful content in assistant replies

    Top tier: pplx-decider v1 27B, Clef, GPT-6 Luna and Clef Flash

    Top tier blocks 9% to 17% of safe rows

    655 rows for Check what the model says
    89.8best score
  5. 05

    Keep the bot on topic

    Off-limits subjects, such as investment or legal advice

    Top tier: Clef, pplx-decider v1 27B, GPT-6 Luna, Kev 4B and Jev 1.13.0

    Top tier blocks 2% to 6% of safe rows

    601 rows for Keep the bot on topic
    98.8best score
  6. 06

    Catch personal data

    Names, email addresses, card numbers and ID numbers

    Top tier: Jev 1.13.0, Clef, pplx-decider v1 27B, Amazon Bedrock Guardrails and Clef Flash

    Top tier blocks 2% to 3% of safe rows

    600 rows for Catch personal data
    98.2best score
  7. 07

    Catch unsupported answers

    Claims that the source document does not support

    Top tier: Jev 1.13.0, GPT-6 Luna and pplx-decider v1 27B

    Top tier blocks 1% to 18% of safe rows

    556 rows for Catch unsupported answers
    91.2best score
  8. 08

    Filter profanity

    Offensive and obscene language

    Top tier: GPT-6 Luna

    Top tier blocks 5% of safe rows

    585 rows for Filter profanity
    93.6best score
Questions

Can you trust these numbers?

Who funded this, and are you tied to any vendor?

No one funded it. raxIT Labs has no commercial relationship with Amazon Web Services, Cloudflare, OpenAI, Perplexity and TypeSafe or any other company on this page. No vendor saw the data, the questions or the results before we published them.

Doesn't the question format favour TypeSafe's Jev?

It may. Every decision model gets the same yes/no questions in the format that Jev's API uses. Jev was trained on that format. The other models get the questions through our adapters. We disclose this home advantage and do not remove it, because any other format would favour a different system. Jev 1.13.0 finishes in tier 3 of 8 overall.

Did you tune thresholds for anyone?

No. Every system uses the same fixed rule: a probability of 0.5 or more blocks. This shows how each system behaves by default. Some models rank risk well but have a poor default threshold. If you tune on your own labelled data, you can gain several points. The smaller models gain the most. The results files report AUROC, so you can see how much room each model has.

Who labelled the data, and how good are the labels?

Labels come from each source and from our written labelling rules. Our lead, with an AI assistant, labelled a 400-row content sample a second time without seeing the first label. The two labels agreed on 86% of rows. An AI model with no access to the answers labelled a 400-row prompt-attack sample a second time. It agreed on 95%. After the runs, we checked again every row that at least 11 of the 12 systems got wrong. We corrected 117 labels and removed 78 ambiguous rows.

Could a model have trained on the test rows?

We compared every row with the published training data of the models on the board. We removed the matches. We also keep a held-back slice of 2,269 rows that we do not publish. Overall scores on that slice are within 2.6 points of the public rows. Some content rows come from datasets that the vendors published themselves. We also score content without those rows. GPT-6 Luna scores 85.8 on all content rows and 86.5 without OpenAI's own rows. So it did worse, not better, on its vendor's data. We cannot rule out training data that a vendor has not disclosed.

Can I reproduce this?

Yes. The dataset is on Hugging Face. The code, the scoring rules and the per-row results are on GitHub. Because of their licences, some sources ship only the row ids. A script rebuilds that text from the original publishers. The Reproduce page has the commands.

What does a tier mean?

Our statistical tests cannot tell the systems in one tier apart from that tier's leader. We resample the rows 2,000 times. We test each system against its tier's leader. Then we correct for the number of comparisons. Two systems in one tier can still differ a lot in what they block and what they cost. Use those two points to choose between them.

Does it cover images, multi-turn chats or my own policies?

Not yet. The benchmark tests text only: one message and its context. Only the off-topic job and an exact-word check cover custom policies. Multimodal and multi-turn tests are not part of this release.

What happens to the data I send these APIs?

We did not evaluate data retention or compliance. These depend on each vendor's terms and on your contract. Some vendors offer zero data retention or HIPAA support to eligible customers. Check the terms before you send regulated data.

Run it yourself

Clone the repository. Get the pinned dataset. Then score your own guardrail on the same checks. We publish new results as GitHub releases. Watch the repository to get an email for each release.