This lecture outlines the complete lifecycle of creating digital linguistic resources for low-resource South Slavic languages, mapping the journey from physical archives to advanced semantic networks. The presentation will walk through the "full circle" of corpus building: beginning with primary, from-scratch digitization and the crucial foundational step of securing institutional Creative Commons agreements between the Drama Theater in Skopje and the Faculty of Philology "Blaze Koneski".
Following primary preservation via Archive.org, the workflow focuses on transforming raw documents into a robust, 4.75-million-word linguistic corpus (Teatarski Glasnik zenodo.org/records/22917991). A key focus will be digital data sovereignty—specifically, enforcing open-access licenses while protecting the dataset from unauthorized commercial AI scraping using novel forensic watermarking techniques.
Crucially, the lecture will demonstrate how his previous and ongoing digitization projects are actively gathering the missing corpus data for the Macedonian language. To illustrate this scale, he will present newly developed corpora of 81 million tokens, built by scraping baseline literary journals i.e. Macedonian language periodicals like Stremez, Sovremenost, Razgledi (partially digitized so far), contemporary postmodern magazines like Margina, and the Macedonian Anarchist Library among others. However, the presentation will argue that while large-scale scraping proves data availability, the future of Macedonian digital humanities requires a shift from raw extraction to legal and structural curation. The focus must now be on building corpora grounded in explicit Creative Commons agreements and ethical data sovereignty. Ultimately, this workflow serves as the rationale for future doctoral research: taking these massive, legally cleared Macedonian datasets and executing rigorous TEI P5 XML integration to seamlessly embed them into Austrian digital infrastructures like ARCHE and CLARIAH-AT.