nori_readingform token filter

edit

The nori_readingform token filter rewrites tokens written in Hanja to their Hangul form.

PUT nori_sample
{
  "settings": {
    "index": {
      "analysis": {
        "analyzer": {
          "my_analyzer": {
            "tokenizer": "nori_tokenizer",
            "filter": [ "nori_readingform" ]
          }
        }
      }
    }
  }
}

GET nori_sample/_analyze
{
  "analyzer": "my_analyzer",
  "text": "鄕歌"      
}

A token written in Hanja: Hyangga

Which responds with:

{
  "tokens" : [ {
    "token" : "향가",     
    "start_offset" : 0,
    "end_offset" : 2,
    "type" : "word",
    "position" : 0
  }]
}

The Hanja form is replaced by the Hangul translation.