Back to feed
arXiv cs.LG
arXiv cs.LG
7/28/2026
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

Short summary

Semalith v1.4 is a 184M-parameter DeBERTa-v3 classifier that performs three-axis safety classification—prompt injection, general harm, and BFSI regulatory compliance—in a single forward pass. It outperforms Llama-Guard-3-8B on all seven prompt-injection benchmarks at 44x fewer parameters, with zero false positives on benign agentic prompts. Llama-Guard-3 still leads on general-harm benchmarks, making the two complementary rather than substitutes.

  • 184M-param classifier beats Llama-Guard-3-8B on prompt-injection detection at 44x fewer parameters
  • Single-pass three-axis safety: prompt injection, general harm, BFSI compliance across 22 classes
  • Zero false-positive rate on 208 benign agentic prompts vs 0.063 for Llama-Guard-3-8B

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more