Jiayan Download - Jiayan Source code download

Jiayan

Other source code

1.0.0

Download

Jiayan

Chinese
English

Introduction

A, which means "oracle bone classical Chinese", is an NLP toolkit focusing on ancient Chinese processing.
Currently, the common Chinese NLP tools mostly use modern Chinese as the core corpus, and the processing effect of ancient Chinese is not satisfactory (see Participle for details). The original intention of this project is to assist in the processing of ancient Chinese information, and help ancient Chinese scholars and enthusiasts who are interested in digging out ancient cultural minerals to better analyze and utilize classical Chinese materials to create "new cultural products" from "cultural heritage".
The current version supports five functions: lexicon construction, automatic word segmentation, part-of-speech annotation, classical Chinese sentence reading and punctuation, and more functions are under development.

Function

Thesaurus construction
- The classical Chinese vocabulary is automatically constructed using unsupervised double dictionary tree, point mutual information, and left and right adjacent entropy.
Participle
- Automatic word segmentation in ancient Chinese is used to use unsupervised, dictionary-free N-metal grammar and hidden Markov model.
- The classical Chinese dictionary generated by the lexicon construction function is used to perform word segmentation based on directed ring-free word graphs, sentence maximum probability paths and dynamic programming algorithms.
Part of speech annotation
- For sequence annotation based on the word conditional random field, please refer to the part-of-speech table for details.
Break sentence
- Based on the sequence annotation of the conditional random field of characters, the introduction of point mutual information and t-test values as characteristics, and automatically breaks sentences for classical Chinese paragraphs.
punctuation
- The sequence annotation of the cascading condition random field based on characters is automatically punctuated on classical Chinese paragraphs based on the sentence breaking.
Translation of Wenbai
- During development, it is currently in the stage of collecting and cleaning parallel corpus of text and white.
- Based on the neural network generation model of bidirectional long and short-term memory recurrent network and attention mechanism, ancient texts are automatically translated.
Note: Due to the influence of corpus, traditional Chinese is not currently supported. If you need to deal with traditional Chinese, you can first use OpenCC to convert the input to simplified Chinese, and then convert the results to the corresponding traditional Chinese (such as Hong Kong, Macao and Taiwan).

Install

 $ pip install jiayan 
$ pip install https://github.com/kpu/kenlm/archive/master.zip

use

The following modules are used from examples.py.

Download the model and decompress: Baidu Netdisk, extract code: p0sc
- jiayan.klm: Language model, mainly used for word segmentation and feature extraction in sentence reading and punctuation tasks;
- pos_model: CRF part-of-speech annotation model;
- cut_model: CRF sentence reading model;
- punc_model: CRF punctuation model;
- Zhuangzi.txt: The full text of Zhuangzi used to test the vocabulary construction.

Thesaurus construction

 from jiayan import PMIEntropyLexiconConstructor

constructor = PMIEntropyLexiconConstructor()
lexicon = constructor.construct_lexicon('庄子.txt')
constructor.save(lexicon, '庄子词库.csv')

result:

 Word,Frequency,PMI,R_Entropy,L_Entropy
之,2999,80,7.944909328101839,8.279435615456894
而,2089,80,7.354575005231323,8.615211168836439
不,1941,80,7.244331150611089,6.362131306822925
...
天下,280,195.23602384978196,5.158574399464853,5.24731990592901
圣人,111,150.0620531154239,4.622606551534004,4.6853474419338585
万物,94,377.59805590304126,4.5959107835319895,4.538837960294887
天地,92,186.73504238078462,3.1492586603863617,4.894533538722486
孔子,80,176.2550051738876,4.284638190120882,2.4056390622295662
庄子,76,169.26227942514097,2.328252899085616,2.1920058354921066
仁义,58,882.3468468468468,3.501609497059026,4.96900162987599
老聃,45,2281.2228260869565,2.384853500510039,2.4331958387289765
...

Participle
1. Character-level hidden Markov model word participle, the effect is in line with the sense of language, it is recommended to use, and the language model jiayan.klm needs to be loaded
```
 from jiayan import load_lm
from jiayan import CharHMMTokenizer

text = '是故内圣外王之道，暗而不明，郁而不发，天下之人各为其所欲焉以自为方。'

lm = load_lm('jiayan.klm')
tokenizer = CharHMMTokenizer(lm)
print(list(tokenizer.tokenize(text)))
```
  result:
  ['是', '故', '内圣外王', '之', '道', '，', '暗', '而', '不', '明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各', '为', '其', '所', '欲', '焉', '以', '自', '为', '方', '。']
  Since ancient Chinese does not have public word segmentation data, it is impossible to evaluate the effect, but we can intuitively feel the advantages of this project through different NLP tools:
  Try to compare the LTP (3.4.0) model participle results:
  ['是', '故内', '圣外王', '之', '道', '，', '暗而不明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各', '为', '其', '所', '欲', '焉以自为方', '。']
  Try comparing HanLP word participle results again:
  ['是故', '内', '圣', '外', '王之道', '，', '暗', '而', '不明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各为其所欲焉', '以', '自为', '方', '。']
  It can be seen that the word participle effect of this tool on ancient Chinese is significantly better than that of the general Chinese NLP tool.
  *Update: Thanks to HanLP's author hankc for letting you know - from early 2021, HanLP released deep learning-driven 2.x. Due to the use of pre-trained language models on large-scale corpus, these corpus have already included almost all ancient and modern Chinese on the Internet, so the effect on ancient Chinese has been qualitatively improved. Not only participle words, but also part-of-shot learning effects and semantic analysis. For the corresponding specific word participle effect, please refer to this Issue.
2. Word-level maximum probability path participle, basically in units of characters, with coarse grain size
```
 from jiayan import WordNgramTokenizer

text = '是故内圣外王之道，暗而不明，郁而不发，天下之人各为其所欲焉以自为方。'
tokenizer = WordNgramTokenizer()
print(list(tokenizer.tokenize(text)))
```
  result:
  ['是', '故', '内', '圣', '外', '王', '之', '道', '，', '暗', '而', '不', '明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各', '为', '其', '所', '欲', '焉', '以', '自', '为', '方', '。']

Part of speech annotation

 from jiayan import CRFPOSTagger

words = ['天下', '大乱', '，', '贤圣', '不', '明', '，', '道德', '不', '一', '，', '天下', '多', '得', '一', '察', '焉', '以', '自', '好', '。']

postagger = CRFPOSTagger()
postagger.load('pos_model')
print(postagger.postag(words))

result:
['n', 'a', 'wp', 'n', 'd', 'a', 'wp', 'n', 'd', 'm', 'wp', 'n', 'a', 'u', 'm', 'v', 'r', 'p', 'r', 'a', 'wp']

Break sentence

 from jiayan import load_lm
from jiayan import CRFSentencizer

text = '天下大乱贤圣不明道德不一天下多得一察焉以自好譬如耳目皆有所明不能相通犹百家众技也皆有所长时有所用虽然不该不遍一之士也判天地之美析万物之理察古人之全寡能备于天地之美称神之容是故内圣外王之道暗而不明郁而不发天下之人各为其所欲焉以自为方悲夫百家往而不反必不合矣后世之学者不幸不见天地之纯古之大体道术将为天下裂'

lm = load_lm('jiayan.klm')
sentencizer = CRFSentencizer(lm)
sentencizer.load('cut_model')
print(sentencizer.sentencize(text))

result:
['天下大乱', '贤圣不明', '道德不一', '天下多得一察焉以自好', '譬如耳目', '皆有所明', '不能相通', '犹百家众技也', '皆有所长', '时有所用', '虽然', '不该不遍', '一之士也', '判天地之美', '析万物之理', '察古人之全', '寡能备于天地之美', '称神之容', '是故内圣外王之道', '暗而不明', '郁而不发', '天下之人各为其所欲焉以自为方', '悲夫', '百家往而不反', '必不合矣', '后世之学者', '不幸不见天地之纯', '古之大体', '道术将为天下裂']

punctuation

 from jiayan import load_lm
from jiayan import CRFPunctuator

text = '天下大乱贤圣不明道德不一天下多得一察焉以自好譬如耳目皆有所明不能相通犹百家众技也皆有所长时有所用虽然不该不遍一之士也判天地之美析万物之理察古人之全寡能备于天地之美称神之容是故内圣外王之道暗而不明郁而不发天下之人各为其所欲焉以自为方悲夫百家往而不反必不合矣后世之学者不幸不见天地之纯古之大体道术将为天下裂'

lm = load_lm('jiayan.klm')
punctuator = CRFPunctuator(lm, 'cut_model')
punctuator.load('punc_model')
print(punctuator.punctuate(text))

result:
天下大乱，贤圣不明，道德不一，天下多得一察焉以自好，譬如耳目，皆有所明，不能相通，犹百家众技也，皆有所长，时有所用，虽然，不该不遍，一之士也，判天地之美，析万物之理，察古人之全，寡能备于天地之美，称神之容，是故内圣外王之道，暗而不明，郁而不发，天下之人各为其所欲焉以自为方，悲夫！百家往而不反，必不合矣，后世之学者，不幸不见天地之纯，古之大体，道术将为天下裂。

Version

v0.0.21
- Divide the installation process into two steps to ensure the latest kenlm version is obtained.
v0.0.2
- Add part-of-speech annotation function.
v0.0.1
- The functions of the vocabulary construction, automatic word segmentation, classical Chinese sentence reading, and punctuation are open.

Introduction

Jiayan, which means Chinese characters engraved on oracle bones, is a professional Python NLP tool for Classical Chinese.
Prevailing Chinese NLP tools are mainly trained on modern Chinese data, which leads to bad performance on Classical Chinese (See Tokenizing ). The purpose of this project is to assist Classical Chinese information processing.
Current version supports lexicon construction, tokenizing, POS tagging, sentence segmentation and automatic punctuation, more features are in development.

Features

Lexicon Construction
- With an unsupervised approach, construct lexicon with Trie -tree, PMI ( point-wise mutual information ) and neighboring entropy of left and right characters.
Tokenizing
- With an unsupervised, no dictionary approach to tokenize a Classical Chinese sentence with N-gram language model and HMM ( Hidden Markov Model ).
- With the dictionary produced from lexicon construction, tokenize a Classical Chinese sentence with Directed Acyclic Word Graph, Max Probability Path and Dynamic Programming.
POS Tagging
- Word level sequence tagging with CRF ( Conditional Random Field ). See POS tag categories here.
Sentence Segmentation
- Character level sequence tagging with CRF, introduces PMI and T-test values as features.
Punctuation
- Character level sequence tagging with layered CRFs, punctuate given Classical Chinese texts based on results of sentence segmentation.
Note: Due to data we used, we don't support traditional Chinese for now. If you have to process traditional one, please use OpenCC to convert traditional input to simplified, then you could convert the results back.

Installation

 $ pip install jiayan 
$ pip install https://github.com/kpu/kenlm/archive/master.zip

Usages

The usage codes below are all from examples.py.

Download the models and unzip them: Google Drive
- jiayan.klm: the language model used for tokenizing and feature extraction for sentence segmentation and punctuation;
- pos_model: the CRF model for POS tagging;
- cut_model: the CRF model for sentence segmentation;
- punc_model: the CRF model for punctuation;
- Zhuangzi.txt: the full text of "Zhuangzi" used for testing lexicon construction.

Lexicon Construction

 from jiayan import PMIEntropyLexiconConstructor

constructor = PMIEntropyLexiconConstructor()
lexicon = constructor.construct_lexicon('庄子.txt')
constructor.save(lexicon, 'Zhuangzi_Lexicon.csv')

Results:

 Word,Frequency,PMI,R_Entropy,L_Entropy
之,2999,80,7.944909328101839,8.279435615456894
而,2089,80,7.354575005231323,8.615211168836439
不,1941,80,7.244331150611089,6.362131306822925
...
天下,280,195.23602384978196,5.158574399464853,5.24731990592901
圣人,111,150.0620531154239,4.622606551534004,4.6853474419338585
万物,94,377.59805590304126,4.5959107835319895,4.538837960294887
天地,92,186.73504238078462,3.1492586603863617,4.894533538722486
孔子,80,176.2550051738876,4.284638190120882,2.4056390622295662
庄子,76,169.26227942514097,2.328252899085616,2.1920058354921066
仁义,58,882.3468468468468,3.501609497059026,4.96900162987599
老聃,45,2281.2228260869565,2.384853500510039,2.4331958387289765
...

Tokenizing
1. The character based HMM, recommended, needs language model: jiayan.klm
```
 from jiayan import load_lm
from jiayan import CharHMMTokenizer

text = '是故内圣外王之道，暗而不明，郁而不发，天下之人各为其所欲焉以自为方。'

lm = load_lm('jiayan.klm')
tokenizer = CharHMMTokenizer(lm)
print(list(tokenizer.tokenize(text)))
```
  Results:
  ['是', '故', '内圣外王', '之', '道', '，', '暗', '而', '不', '明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各', '为', '其', '所', '欲', '焉', '以', '自', '为', '方', '。']
  Since there is no public tokenizing data for Classical Chinese, it's hard to do performance evaluation directly; however, we can compare the results with other popular modern Chinese NLP tools to check the performance:
  Compare the tokenizing result of LTP (3.4.0):
  ['是', '故内', '圣外王', '之', '道', '，', '暗而不明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各', '为', '其', '所', '欲', '焉以自为方', '。']
  Also, compare the tokenizing result of HanLP:
  ['是故', '内', '圣', '外', '王之道', '，', '暗', '而', '不明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各为其所欲焉', '以', '自为', '方', '。']
  It's apparent that Jiayan has much better tokenizing performance than general Chinese NLP tools.
2. Max probability path approaching tokenizing based on words
```
 from jiayan import WordNgramTokenizer

text = '是故内圣外王之道，暗而不明，郁而不发，天下之人各为其所欲焉以自为方。'
tokenizer = WordNgramTokenizer()
print(list(tokenizer.tokenize(text)))
```
  Results:
  ['是', '故', '内', '圣', '外', '王', '之', '道', '，', '暗', '而', '不', '明', '，', '郁', '而', '不', '发', '，', '天下', '之', '人', '各', '为', '其', '所', '欲', '焉', '以', '自', '为', '方', '。']

POS Tagging

 from jiayan import CRFPOSTagger

words = ['天下', '大乱', '，', '贤圣', '不', '明', '，', '道德', '不', '一', '，', '天下', '多', '得', '一', '察', '焉', '以', '自', '好', '。']

postagger = CRFPOSTagger()
postagger.load('pos_model')
print(postagger.postag(words))

Results:
['n', 'a', 'wp', 'n', 'd', 'a', 'wp', 'n', 'd', 'm', 'wp', 'n', 'a', 'u', 'm', 'v', 'r', 'p', 'r', 'a', 'wp']

Sentence Segmentation

 from jiayan import load_lm
from jiayan import CRFSentencizer

text = '天下大乱贤圣不明道德不一天下多得一察焉以自好譬如耳目皆有所明不能相通犹百家众技也皆有所长时有所用虽然不该不遍一之士也判天地之美析万物之理察古人之全寡能备于天地之美称神之容是故内圣外王之道暗而不明郁而不发天下之人各为其所欲焉以自为方悲夫百家往而不反必不合矣后世之学者不幸不见天地之纯古之大体道术将为天下裂'

lm = load_lm('jiayan.klm')
sentencizer = CRFSentencizer(lm)
sentencizer.load('cut_model')
print(sentencizer.sentencize(text))

Results:
['天下大乱', '贤圣不明', '道德不一', '天下多得一察焉以自好', '譬如耳目', '皆有所明', '不能相通', '犹百家众技也', '皆有所长', '时有所用', '虽然', '不该不遍', '一之士也', '判天地之美', '析万物之理', '察古人之全', '寡能备于天地之美', '称神之容', '是故内圣外王之道', '暗而不明', '郁而不发', '天下之人各为其所欲焉以自为方', '悲夫', '百家往而不反', '必不合矣', '后世之学者', '不幸不见天地之纯', '古之大体', '道术将为天下裂']

Punctuation

 from jiayan import load_lm
from jiayan import CRFPunctuator

text = '天下大乱贤圣不明道德不一天下多得一察焉以自好譬如耳目皆有所明不能相通犹百家众技也皆有所长时有所用虽然不该不遍一之士也判天地之美析万物之理察古人之全寡能备于天地之美称神之容是故内圣外王之道暗而不明郁而不发天下之人各为其所欲焉以自为方悲夫百家往而不反必不合矣后世之学者不幸不见天地之纯古之大体道术将为天下裂'

lm = load_lm('jiayan.klm')
punctuator = CRFPunctuator(lm, 'cut_model')
punctuator.load('punc_model')
print(punctuator.punctuate(text))

Results:
天下大乱，贤圣不明，道德不一，天下多得一察焉以自好，譬如耳目，皆有所明，不能相通，犹百家众技也，皆有所长，时有所用，虽然，不该不遍，一之士也，判天地之美，析万物之理，察古人之全，寡能备于天地之美，称神之容，是故内圣外王之道，暗而不明，郁而不发，天下之人各为其所欲焉以自为方，悲夫！百家往而不反，必不合矣，后世之学者，不幸不见天地之纯，古之大体，道术将为天下裂。

Versions

v0.0.21
- Divide the installation into two steps to ensure to get the latest version of kenlm.
v0.0.2
- POS tagging feature is open.
v0.0.1
- Add features of lexicon construction, tokenizing, sentence segmentation and automatic punctuation.

Expand

Additional Information

Version 1.0.0
Type Other source code
Update Time 2025-04-16
size 216.93KB
From Github

Related Applications

Google Dorks

2025-03-10
shepherd

2025-06-04
mongo express

2025-06-04
hidusbf

2025-02-14
Free Algorithms Books

2025-05-29
markdownpedia

2025-04-22

Recommended for You

chat.petals.dev

Other source code

1.0.0
GPT Prompt Templates

Other source code

1.0.0
GPTyped

Other source code

GPTyped 1.0.5
Google Dorks

Other source code

1.0
shepherd

Other source code

v6.1.6-react-shepherd: Prepare Release (#3063)
mongo express

Other source code

v1.1.0-rc-3
Google Dorks

Other source code

1.0
shepherd

Other source code

v6.1.6-react-shepherd: Prepare Release (#3063)
mongo express

Other source code

v1.1.0-rc-3

Related Information All