Data Analysis and Knowledge Discovery  2018, Vol. 2 Issue (12): 68-76    DOI: 10.11925/infotech.2096-3467.2018.0391
Classifying Chinese Texts with CapsNet
Guoming Feng,Xiaodong Zhang(),Suhui Liu
School of Economics and Management, University of Science and Technology Beijing, Beijing 100083, China
[Objective] This study tries to address the issues facing long text representation and use CapsNet to improve the accuracy of Chinese text classification. [Methods] First, we proposed a LDA matrix and word vector to represent the long texts. Then, we constructed a Chinese classification model based on CapsNet. Third, we examined the proposed model with Sogou news corpus and the text classification corpus of Fudan University. Finally, we compared our results with those of the classic models (e.g., TextCNN, DNN and so on). [Results] The performance of CapsNet model was better than other models. The classification accuracy in five categories of short and long texts reached 89.6% and 96.9% respectively. The convergence speed of the proposed model was almost two times faster than that of the CNN model. [Limitations] The computational complexity of the model is high, which limits the size of testing corpus. [Conclusions] The proposed Chinese text representation method and the modified CapsNet model have better accuracy, convergence speed and robustness than the existing ones.

Key wordsText Categorization      CapsNet      Deep Learning      Text Representation      TextCNN     
Received: 08 April 2018      Published: 16 January 2019

Cite this article:

Guoming Feng,Xiaodong Zhang,Suhui Liu. Classifying Chinese Texts with CapsNet. Data Analysis and Knowledge Discovery, 2018, 2(12): 68-76.

