Papers/2610.10592
🧪 Test?View on arXiv

Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch

Not specified in the provided content

language modelinghistorical linguisticsdata augmentationmachine learning
2610.10592
Builder Relevance
70%
2h ago

Abstract

The paper presents Wieszcz-XIX, a large corpus of historical Polish text, significantly expanding the available machine-readable data for this language period.

Reality Card

Core Claim

The authors created a 6.75 billion token corpus from 294,369 documents, which is over three orders of magnitude larger than existing annotated datasets for pre-1918 Polish.

Method / Result

The character error rate on a hand-corrected sample is 0.68%, with 45% of sampled passages being uncorrectable.

Limitations

The models trained on this corpus reproduce period-specific prejudices, including antisemitic statements.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers