50 releases (26 breaking)

new 0.32.3 Mar 18, 2025
0.32.2 Jun 29, 2024
0.31.0 May 28, 2024
0.29.0 Mar 18, 2024
0.3.2 Feb 20, 2020

#410 in Text processing

Download history 4682/week @ 2024-11-27 3362/week @ 2024-12-04 3517/week @ 2024-12-11 3030/week @ 2024-12-18 2070/week @ 2024-12-25 2862/week @ 2025-01-01 4504/week @ 2025-01-08 3582/week @ 2025-01-15 2847/week @ 2025-01-22 3923/week @ 2025-01-29 3706/week @ 2025-02-05 3487/week @ 2025-02-12 3608/week @ 2025-02-19 2762/week @ 2025-02-26 3210/week @ 2025-03-05 2983/week @ 2025-03-12

13,167 downloads per month
Used in 14 crates (2 directly)

MIT license

155KB
3K SLoC

Lindera UniDic Builder

License: MIT Join the chat at https://gitter.im/lindera-morphology/lindera Crates.io

UniDic builder for Lindera.

Dictionary version

This repository contains unidic-mecab.

Dictionary format

Refer to the manual for details on the unidic-mecab dictionary format and part-of-speech tags.

Index Name (Japanese) Name (English) Notes
0 表層形 Surface
1 左文脈ID Left context ID
2 右文脈ID Right context ID
3 コスト Cost
4 品詞大分類 Major POS classification
5 品詞中分類 Middle POS classification
6 品詞小分類 Small POS classification
7 品詞細分類 Fine POS classification
8 活用型 Conjugation form
9 活用形 Conjugation type
10 語彙素読み Lexeme reading
11 語彙素(語彙素表記 + 語彙素細分類) Lexeme
12 書字形出現形 Orthography appearance type
13 発音形出現形 Pronunciation appearance type
14 書字形基本形 Orthography basic type
15 発音形基本形 Pronunciation basic type
16 語種 Word type
17 語頭変化型 Prefix of a word form
18 語頭変化形 Prefix of a word type
19 語末変化型 Suffix of a word form
20 語末変化形 Suffix of a word type

User dictionary format (CSV)

Simple version

Index Name (Japanese) Name (English) Notes
0 表層形 Surface
1 品詞大分類 Major POS classification
2 語彙素読み Lexeme reading

Detailed version

Index Name (Japanese) Name (English) Notes
0 表層形 Surface
1 左文脈ID Left context ID
2 右文脈ID Right context ID
3 コスト Cost
4 品詞大分類 Major POS classification
5 品詞中分類 Middle POS classification
6 品詞小分類 Small POS classification
7 品詞細分類 Fine POS classification
8 活用型 Conjugation form
9 活用形 Conjugation type
10 語彙素読み Lexeme reading
11 語彙素(語彙素表記 + 語彙素細分類) Lexeme
12 書字形出現形 Orthography appearance type
13 発音形出現形 Pronunciation appearance type
14 書字形基本形 Orthography basic type
15 発音形基本形 Pronunciation basic type
16 語種 Word type
17 語頭変化型 Prefix of a word form
18 語頭変化形 Prefix of a word type
19 語末変化型 Suffix of a word form
20 語末変化形 Suffix of a word type
21 - - After 21, it can be freely expanded.

How to use IPADIC dictionary

For more details about lindera command, please refer to the following URL:

API reference

The API reference is available. Please see following URL:

Dependencies

~9MB
~211K SLoC