Katakura Shumpei
Tohoku University Archives. Specially Appointed Lecturer

Integrated Access to Japanese Materials and AI Applications in the Tohoku University Digital Archives (ToUDA)

Authors: Shumpei Katakura (Tohoku University Archives), Satoshi Kato (Tohoku University Archives), Tomoe Hanzawa (Tohoku University Library)

Tohoku University Digital Archives (ToUDA) is a platform that integrates Japanese resources held across Tohoku University—including its library, archives, museum, and archaeological unit—and publishes images and metadata via a bilingual Japanese/English interface. Launched in April 2024 with the guiding concepts of "integration, visibility, and research use," ToUDA aims to excavate dormant content, visualize it centrally, and expand continuously based on a versatile data design. It is operated through a faculty–staff collaborative framework centered on the University Library, the Center for Academic Resources and Archives, and the Center for Integrated Japanese Studies.

As of October 8, 2025, ToUDA provides 3,705,207 TIFF image frames and 156,840 metadata records across 14 collections. Items with images include 19,309 color titles and 29,460 monochrome titles.

This presentation first outlines the platform’s policies and participating units, then introduces the scope and distinctive features of the materials currently available. These include postwar student movement resources, university history and administrative records, campus-related photographs, rare books, archaeological survey materials, and natural science specimens. It also demonstrates practical workflows for research and teaching, ranging from basic search and faceted filtering (period, holding unit, material type, and variant characters) to high-resolution image access via an IIIF viewer. Finally, the presentation discusses future directions, including open-data initiatives accompanied by the assignment of persistent identifiers (PIDs) to metadata and content, external API integration, and AI-enabled services such as OCR, LLM-assisted transcription support, automatic summarization, and keyword extraction.

東北大学総合知デジタルアーカイブ(ToUDA)による日本資料の統合公開とAI活用

東北大学総合知デジタルアーカイブToUDAは、附属図書館、史料館、博物館、埋蔵文化財調査室、など学内に分散する日本資料を統合し、画像とメタデータを日英2言語で公開するプラットフォームとして2024年4月に公開された。「統合・可視性・研究利用」を掲げ、休眠コンテンツの発掘と一元可視化、汎用性の高いデータ設計にもとづく継続的な拡張を進めている。附属図書館・学術資源研究公開センター・統合日本学センターを中心とする教職協働体制で運用し、国内外のデジタルアーカイブとの連携も視野に入れている。収録状況は2025年10月8日現在、TIFF3,705,207コマ、メタデータ156,840件、画像有タイトル数はカラー19,309点・モノクロ29,460点、総コレクション14点である。

本発表では、まずToUDAの方針や連携機関等の概観を説明し、次いで戦後学生運動資料、大学史・学内行政資料、学内関係写真、貴重書、発掘調査資料、自然科学標本等、ToUDAが公開する資料の範囲と特徴を示し、基本的な検索・絞り込み(年代、所蔵、資料種別、異体字等)から、IIIFビューワでの画像公開システムなど、研究・教育での実践的な使い方を紹介する。最後に、「メタデータ・コンテンツの永続的識別子付与に伴うオープンデータ化とデータセット公開」、「外部API連携」、「OCRとLLMによる翻刻支援や自動要約・キーワード抽出等のAI活用」といった将来への展望を述べる。