Opening Historical Archives

Many libraries and museums have digitized their historical collections, but these documents can be searched most of the time only by catalog details, not by the actual page content. We are working to change that, so researchers and the general public can explore this vast data.

We use frontier OCR models to recognize both the text and structure of each page. We deliver the results in ALTO XML, a format already used by many library and archive systems. We have processed more than 800,000 pages and continue this work with collections in Spain and other countries.

Collage of historical newspaper front pages

How we work

01

Leading models, one toolkit

Use leading OCR and vision-language models, including OlmOCR-2, PaddleOCR and GLM-OCR, from a single toolkit. They capture the text, layout and structure of each page.

02

Formats archives already use

Every page is delivered as ALTO XML v3, a format already used by library and museum search systems. The output works directly with search and highlighting, without custom integration.

03

Built for large collections

Whether you are processing thousands or millions of pages, our workflows scale with the collection and provide clear reports on progress and quality.

See it in action

Hover over the page to explore the regions the model detected, and toggle the ALTO XML view to see the standard format libraries and archives consume.

Loading example…

Image from Biblioteca Virtual del Patrimonio Bibliográfico

Work with us

Are you a library, archive, or museum looking to unlock your collections? Get in touch.