Sie befinden Sich nicht im Netzwerk der Universität Paderborn. Der Zugriff auf elektronische Ressourcen ist gegebenenfalls nur via VPN oder Shibboleth (DFN-AAI) möglich. mehr Informationen...
Ergebnis 6 von 56

Details

Autor(en) / Beteiligte
Titel
Searchable Turkish OCRed historical newspaper collection 1928–1942
Ist Teil von
  • Journal of information science, 2023-04, Vol.49 (2), p.335-347
Ort / Verlag
London, England: SAGE Publications
Erscheinungsjahr
2023
Quelle
Alma/SFX Local Collection
Beschreibungen/Notizen
  • The newspaper emerged as a distinct cultural form in early 17th-century Europe. It is bound up with the early modern period of history. Historical newspapers are of utmost importance to nations and its people, and researchers from different disciplines rely on these papers to improve our understanding of the past. In pursuit of satisfying this need, Istanbul University Head Office of Library and Documentation provides access to a big database of scanned historical newspapers. To take it another step further and make the documents more accessible, we need to run optical character recognition (OCR) and named entity recognition (NER) tasks on the whole database and index the results to allow for full-text search mechanism. We design and implement a system encompassing the whole pipeline starting from scrapping the dataset from the original website to providing a graphical user interface to run search queries, and it manages to do that successfully. Proposed system provides to search people, culture and security-related keywords and to visualise them.
Sprache
Englisch
Identifikatoren
ISSN: 0165-5515
eISSN: 1741-6485
DOI: 10.1177/01655515211000642
Titel-ID: cdi_proquest_journals_2787856710

Weiterführende Literatur

Empfehlungen zum selben Thema automatisch vorgeschlagen von bX