Data Analysis and Knowledge Discovery  2017, Vol. 1 Issue (2): 28-34    DOI: 10.11925/infotech.2096-3467.2017.02.04
Extracting Keywords with Modified TextRank Model
Tian Xia()
Key Laboratory of Data Engineering and Knowledge Engineering of Ministry of Education, Renmin University of China, Beijing 100872, China
School of Information Resource Management, Renmin University of China, Beijing 100872, China
[Objective] This study aims to improve the single document keyword extraction algorithm by adding the world knowledge vector from the Wikipedia to the TextRank model. [Methods] First, we created a new word embedding model based on the Word2Vec model with Wikipedia’s Chinese data. Second, we clustered the nodes of TextRank wordgraph to adjust the voting importance of each cluster. Third, we calculated the random walk probability with additional factors of coverage and location. Finally, we got the node score with iterative computation of the transition matrix, and then selected the Top N words as the needed keywords. [Results] The performance of the new TextRank model was much better than other methods when the Top N value was less than or equal to 7. If we only retrieved three keywords, the F measure reached its maximum value, which was 3.374% higher than the best existing results. When the Top N value was larger than 7, the results were similar to the traditional TextRank method. [Limitations] The computation cost was increased due to the cluster analysis. [Conclusions] The new weighted TextRank model could extract keywords effectively.

Key wordsKeyword Extraction      Word Embedding      TextRank      Word2vec     
Received: 28 October 2016      Published: 27 March 2017

Cite this article:

Tian Xia. Extracting Keywords with Modified TextRank Model. Data Analysis and Knowledge Discovery, 2017, 1(2): 28-34.

