🧪 Test?View on arXiv
Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch
Not specified in the provided content
language modelinghistorical linguisticsdata augmentationmachine learning
2610.10592
Builder Relevance
2h ago70%
Abstract
The paper presents Wieszcz-XIX, a large corpus of historical Polish text, significantly expanding the available machine-readable data for this language period.
Reality Card
Core Claim
The authors created a 6.75 billion token corpus from 294,369 documents, which is over three orders of magnitude larger than existing annotated datasets for pre-1918 Polish.
Method / Result
The character error rate on a hand-corrected sample is 0.68%, with 45% of sampled passages being uncorrectable.
Limitations
The models trained on this corpus reproduce period-specific prejudices, including antisemitic statements.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.