Main Reading Room, Library of Congress (Credit: Carol M. Highsmith, Public Domain)
SUBTOPIC

Digital Humanities

Training a language model to read a language that almost no one still writes.

Image: Carol M. Highsmith, Library of Congress

From Archive to Open Model

Training a large language model on a low-resource or historical language means confronting a hard data problem: there simply isn't enough written material to match the scale of English-language training corpora. Digital humanities researchers compensate by mining national library archives, broadcast transcripts, and digitized manuscripts, then pairing that text with cross-lingual signals from high-resource languages. Computer vision models trained the same way can trace stylistic influence across centuries of paintings and manuscripts. Malgukke's GPU partitions are built to make this kind of large-scale training accessible outside a handful of major AI labs.

HPC Solution Architecture: Archive to Open Language Model

A low-resource language model training pipeline, as run on the LUMI supercomputer for the Finnish Poro model:

Archival Text National library, broadcast archives Training Core CPU Nodes Corpus cleaning & tokenization GPU Nodes Cross-lingual transformer training Base Model Low-resource language LLM Evaluation Benchmarking vs. high-resource models Open Release Public model, research checkpoints Community feedback on the open release guides the next archival mining pass

Running on the TOP500: LUMI

TOP500 — 5TH PLACE

LUMI, hosted by CSC in Kajaani, Finland, is Europe's leading supercomputer, ranked 5th on the TOP500 list with a measured 379.7 petaflops. Its GPU partition was used by TurkuNLP and Silo AI to train Poro, the first multilingual LLM trained on a EuroHPC supercomputer, pairing Finnish with English text to overcome the scarcity of written Finnish source material — the exact archival-to-model pipeline described on this page.

Yes — this is a real, currently operating top-5 TOP500 system, directly used for the low-resource language model training described here.

#5
Global TOP500 Rank
379.7
Petaflop/s (HPL)

Voices from the Field

Sampo Pyysalo

TurkuNLP, University of Turku

On why language technology is essential to the survival of small languages, and on mining Finland's national broadcast and library archives for training data.

Read article
Filip Ginter

Professor, University of Turku (TurkuNLP)

On the LUMI pilot project that produced the largest Finnish language model to date, published for open access.

Read article