Word Segmentation in Japanese and Romanisation
日本語の分かち書きとローマ字翻字
This Cross Talk series was conceived following its adoption at the 2025 annual conference of the European Association of Japanese Resource Specialists, held in Heidelberg. It takes as its central theme the development and use of metadata within the GLAM (Galleries, Libraries, Archives and Museums) sector and sets out to examine the challenges surrounding Japanese word segmentation (wakachigaki) and romanised transcription. At its core lies a concern with how Japanese-language materials can be described in ways that ensure both discoverability and readability across diverse user environments, including those involving non-Japanese speakers.
Japanese differs markedly from alphabet-based Western languages, where words are typically separated clearly in writing. It is primarily written using a mixture of kanji and kana, and although readings may occasionally be indicated through ruby annotations, boundaries between words are not explicitly marked. As a result, determining “where to segment” is far from straightforward.
Such differences in segmentation have a direct impact on the outcomes of romanised transcription. At the same time, there is no single, universally accepted standard for the romanisation of Japanese. Instead, multiple systems coexist, reflecting differing purposes such as bibliographic description, general usage and international dissemination.
While word segmentation and romanisation might appear to be of limited relevance to domestic metadata creation by Japanese speakers, this is not in fact the case. Even in Japanese-language cataloguing contexts, there are many instances in which practitioners must make difficult judgements about how to segment words, particularly when assigning readings. Moreover, with the increasing global circulation of information about Japan, opportunities — and indeed requirements — to provide romanised forms within metadata are becoming more frequent.
There is a growing expectation that such issues might be resolved wholesale through AI; however, this is not yet a realistic prospect. Simply assigning alphabetic equivalents to the basic syllabary of hiragana is insufficient. Decisions about word segmentation depend heavily on context, convention and exceptional usage, and romanisation likewise involves a complex interplay of rules and exceptions. In practice, therefore, careful rule design, the management of exceptions and ongoing operational calibration are all indispensable.
In addition, requirements within the GLAM sector are far from uniform. For example, the LC Romanization System used in bibliographic data is optimised for cataloguing purposes and does not necessarily suit all contexts of use. Requirements differ according to function — whether for search, data linking or display — and accordingly, what constitutes an “optimal” romanised form also varies.
Rather than proposing a single solution, this Cross Talk series aims to share conceptual frameworks for negotiating these diverse requirements, while fostering deeper cross-disciplinary discussion and contributing to practical implementation.
The kick-off event will be held at the end of September, 2026.


