Digital Humanities
Training a language model to read a language that almost no one still writes.
Image: Carol M. Highsmith, Library of Congress
From Archive to Open Model
Training a large language model on a low-resource or historical language means confronting a hard data problem: there simply isn't enough written material to match the scale of English-language training corpora. Digital humanities researchers compensate by mining national library archives, broadcast transcripts, and digitized manuscripts, then pairing that text with cross-lingual signals from high-resource languages. Computer vision models trained the same way can trace stylistic influence across centuries of paintings and manuscripts. Malgukke's GPU partitions are built to make this kind of large-scale training accessible outside a handful of major AI labs.
HPC Solution Architecture: Archive to Open Language Model
A low-resource language model training pipeline, as run on the LUMI supercomputer for the Finnish Poro model:
Running on the TOP500: LUMI
LUMI, hosted by CSC in Kajaani, Finland, is Europe's leading supercomputer, ranked 5th on the TOP500 list with a measured 379.7 petaflops. Its GPU partition was used by TurkuNLP and Silo AI to train Poro, the first multilingual LLM trained on a EuroHPC supercomputer, pairing Finnish with English text to overcome the scarcity of written Finnish source material — the exact archival-to-model pipeline described on this page.
Yes — this is a real, currently operating top-5 TOP500 system, directly used for the low-resource language model training described here.
Voices from the Field
Sampo Pyysalo
TurkuNLP, University of Turku
On why language technology is essential to the survival of small languages, and on mining Finland's national broadcast and library archives for training data.
Read articleFilip Ginter
Professor, University of Turku (TurkuNLP)
On the LUMI pilot project that produced the largest Finnish language model to date, published for open access.
Read article