Ogiso Toshinobu
National Institute for Japanese Language and Linguistics. Professor

オンラインツール「Web茶まめ」を用いた日本語表記の自動ローマ字変換

漢字仮名混じりの日本語表記をローマ字化する際、カナとアルファベットの対応ルールに加えて、語の分かち書きのルールが最低限必要となる。そして語の分かち書きには、助詞助動詞、接頭辞・接尾辞の判別などの品詞分類を前提としたルールがつきまとうが、この区別は一般には容易ではない。一方、国立国語研究所では日本語コーパスの構築のために自動でテキストを単語に分割して品詞や読みの情報を付ける形態素解析用辞書と、それを利用するためのオンラインツール「Web茶まめ」を開発し、公開してきた(https://chamame.ninjal.ac.jp/)。今回、この「Web茶まめ」の読み・品詞付与を応用して、入力したテキストを自動てローマ字に変換する機能を用意した。このツールでの変換に用いるローマ字は2025年12月の閣議告示による新しい「ローマ字のつづり方」(おおむね修正ヘボン式にもとづく)を採用し、分かち書きの基準はアメリカ議会図書館のローマ字化方式(Japanese Romanization Table 2022 version)を参考にしている。現時点では、慣習的な固有名詞のローマ字表記、誤った単語に解析する問題など対応すべき課題は多いが、専門的な知識がなくとも品詞などの文法に関連した分かち書きルールを意識することなくローマ字化が可能な点で一定の利用価値がある。本発表ではこのローマ字変換機能について、特に日本語資料の目録作成・検索・国際的な情報共有における利用可能性と課題の観点から論じる。

Automatic Romanisation of Japanese Orthography Using WebChamame

When romanising Japanese text written in mixed kanji and kana, it is necessary not only to apply rules for converting kana into Roman letters, but also, at a minimum, to determine how words should be separated. Word separation, in turn, inevitably involves rules based on part-of-speech classification, such as distinguishing particles and auxiliary verbs from other elements, and identifying prefixes and suffixes. These distinctions are generally not easy to make.

The National Institute for Japanese Language and Linguistics has developed and made publicly available morphological dictionaries for automatically segmenting Japanese texts into words and assigning information such as part of speech and readings, together with the online tool WebChamame, which makes use of these dictionaries for the construction of Japanese corpora (https://chamame.ninjal.ac.jp/). We have now developed a function that applies the reading and part-of-speech annotation provided by WebChamame to automatically convert input text into romanised form.

The romanisation used in this tool follows the new official Romanization System for Japanese announced by the Japanese Cabinet in December 2025, which is broadly based on modified Hepburn romanisation. The criteria for word separation are based on the Library of Congress romanisation system, specifically the Japanese Romanization Table, 2022 version.

At present, a number of issues remain to be addressed, including the treatment of conventional romanisations of proper nouns and errors caused by incorrect word segmentation or analysis. Nevertheless, the tool has practical value in that it enables users without specialist knowledge to romanise Japanese text without having to be aware of word-separation rules related to grammatical information such as part of speech. This presentation discusses the romanisation function of WebChamame, focusing in particular on its potential applications and remaining challenges in the cataloguing, searching, and international sharing of information on Japanese-language materials.