IK中文分词器在Elasticsearch上的使用。原生IK中文分词是从文件系统中读取词典,es-ik本身可扩展成从不同的源读取词典。目前提供从sqlite3数据库中读取。es-ik-plugi-sqlite3使用方法:
1.在elasticsearch.yml中设置你的sqlite3词典的位置:
ik_aalysis_db_path: /opt/ik/dictioary.db我提供了默认的词典:https://github.com/zacker330/es-ik-sqlite3-dictioary
2.安装(目前是1.0.1版本)
./bi/plugi -i ik-aalysis -u https://github.com/zacker330/es-ik-plugi-sqlite3-release/raw/master/es-ik-sqlite3-1.0.1.zip3.现在可以测试了:
1.创建idex
curl -X PUT -H "Cache-Cotrol: o-cache" -d '{ "settigs":{ "idex":{ "umber_of_shards":1, "umber_of_replicas": 1 } }}' 'https://localhost:9200/sogs/'2.创建map:
curl -X PUT -H "Cache-Cotrol: o-cache" -d '{ "sog": { "_source": {"eabled": true}, "_all": { "idexAalyzer": "ik_aalysis", "searchAalyzer": "ik_aalysis", "term_vector": "o", "store": "true" }, "properties":{ "title":{ "type": "strig", "store": "yes", "idexAalyzer": "ik_aalysis", "searchAalyzer": "ik_aalysis", "iclude_i_all": "true" } } }} ' 'https://localhost:9200/sogs/_mappig/sog'3.
curl -X POST -d '林夕为我们作词' 'https://localhost:9200/sogs/_aalyze?aalyzer=ik_aalysis'respose:{"tokes":[{"toke":"林夕","start_offset":0,"ed_offset":2,"type":"CN_WORD","positio":1},{"toke":"作词","start_offset":5,"ed_offset":7,"type":"CN_WORD","positio":2}]}






评论