Back to feed
arXiv cs.CL
arXiv cs.CL
8/5/2026
Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

Short summary

The first independent evaluation of OpenAI's 1.5B-parameter Privacy Filter tests PII detection across 42 benchmarks in 22 languages and 5 domains. OPF outperforms Presidio and XLM-RoBERTa on structured synthetic PII but collapses on non-Latin scripts (Arabic F1=0.04, Cyrillic F1=0.03) and degrades on narrative prose. GPT-4o leads on medical, legal, and financial PII, while OPF is strongest on structurally regular types like emails and phone numbers.

  • OpenAI's Privacy Filter achieves F1=0.855 on AI4Privacy but collapses on non-Latin scripts and narrative prose
  • GPT-4o outperforms OPF on medical, legal, and financial PII detection
  • OPF is strongest on structurally regular PII (email: 0.78, phone: 0.76) and weakest on culturally variable types (person: 0.40, address: 0.49)

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more