Papers/2608.16928
🧪 Test?View on arXiv

Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification

Not specified in the provided content

sensitivity classificationleakage controltransformer modelsbenchmarking
2608.16928
Builder Relevance
80%
2h ago

Abstract

This paper addresses the problem of automatic sensitivity classification of organizational documents, highlighting the issue of label leakage and presenting a new benchmark for evaluating model performance.

Reality Card

Core Claim

The introduction of the Strategic 16K corpus and a systematic benchmark demonstrates that BERT achieves the highest performance in sensitivity classification under leakage-controlled conditions.

Method / Result

BERT achieved an accuracy of 89.14% and an F1 score of 89.33%.

Limitations

The reliance on a specific dataset (WikiLeaks PlusD) may limit generalizability to other document types.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers