14 releases

✓ Uses Rust 2018 edition

0.3.4 Feb 25, 2020
0.3.3 Feb 25, 2020
0.2.1 Feb 12, 2020
0.1.6 Feb 7, 2020

#168 in Text processing

Download history 126/week @ 2020-02-03 66/week @ 2020-02-10 84/week @ 2020-02-17 120/week @ 2020-02-24 31/week @ 2020-03-02 50/week @ 2020-03-09 38/week @ 2020-03-16 95/week @ 2020-03-23 24/week @ 2020-03-30

102 downloads per month
Used in 3 crates (2 directly)

MIT license

18KB
348 lines

Lindera

License: MIT Join the chat at https://gitter.im/lindera-morphology/lindera

A Japanese morphological analysis library in Rust. This project fork from fulmicoton's kuromoji-rs.

Lindera aims to build a library which is easy to install and provides concise APIs for various Rust applications.

Build

The following products are required to build:

  • Rust >= 1.39.0
  • make >= 3.81
% make build

Usage

Basic example

This example covers the basic usage of Lindera.

It will:

  • Create a tokenizer in normal mode
  • Tokenize the input text
  • Output the tokens
use lindera::tokenizer::Tokenizer;

fn main() -> std::io::Result<()> {
    // create tokenizer
    let mut tokenizer = Tokenizer::default_normal();

    // tokenize the text
    let tokens = tokenizer.tokenize("関西国際空港限定トートバッグ");

    // output the tokens
    for token in tokens {
        println!("{}", token.text);
    }

    Ok(())
}

The above example can be run as follows:

% cargo run --example basic_example

You can see the result as follows:

関西国際空港
限定
トートバッグ

API reference

The API reference is available. Please see following URL:

Dependencies

~17MB
~123K SLoC