Exporting Finnish digitized historical newspaper contents for offline use

Research output: Contribution to journalArticleScientificpeer-review

Researchers

Research units

  • National Library of Finland
  • University of Turku

Abstract

Digital collections of the National Library of Finland (NLF) contain over 10 million pages of historical newspapers, journals and some technical ephemera. The material ranges from the early Finnish newspapers from 1771 until the present day. The material up to 1910 can be viewed in the public web service, where as anything later is available at the six legal deposit libraries in Finland. A recent user study noticed that a different type of researcher use is one of the key uses of the collection. National Library of Finland has gotten several requests to provide the content of the digital collections as one offline bundle, where all the needed content is included. For this purpose we introduced a new format, which contains three different information sets: the full metadata of a publication page, the actual page content as ALTO XML, and the raw text content. We consider these formats most useful to be provided as raw data for the researchers. In this paper we will describe how the export format was created, how other parties have packaged the same data and what the benefits are of the current approach. We shall also briefly discuss word level quality of the content and show a real research scenario for the data.

Details

Original languageEnglish
Number of pages1
JournalD-LIB MAGAZINE
Volume22
Issue number7-8
Publication statusPublished - 1 Jul 2016
MoE publication typeA1 Journal article-refereed

ID: 9630490