Abstract:This paper firstly analyzes a graphical framework for name disambiguation called GHOST, and then provides a modified name disambiguation algorithm combining with the text mining of literature information. The new algorithm is more suitable for literature database, making up for the limitations existed in GHOST. Based on selecting title and publication name as computing feature from the literature information, the experiment shows that the algorithm achieves high precision and recall value, and F1 reaches 84%, which is good enough for name disambiguation.
郭舒. 文献数据库中作者名消歧算法研究[J]. 现代图书情报技术, 2013, 29(7/8): 69-74.
Guo Shu. Research on Author Name Disambiguation Algorithm in the Literature Database. New Technology of Library and Information Service, 2013, 29(7/8): 69-74.
[1] Han H, Giles L, Zha H, et al. Two Supervised Learning Approaches for Name Disambiguation in Author Citations[C]. In: Proceedings of the 4th ACM/IEEE Joint Conference on Digital Libraries (JCDL '04). New York: ACM, 2004:296-305.[2] Treeratpituk P, Giles C L. Disambiguating Authors in Academic Publications Using Random Forests[C]. In:Proceedings of the 9th ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL'09). New York: ACM,2009:39-48.[3] Han H, Zha H, Giles C L. Name Disambiguation in Author Citations Using a K-way Spectral Clustering Method[C]. In: Proceedings of the 5th ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL'05). New York: ACM, 2005:334-343.[4] Fan X M, Wang J Y, Pu X, et al. On Graph-based Name Disambiguation[J]. Journal of Data and Information Quality, 2011, 2(2):23-56.[5] Pereira D A, Ribeiro-Neto B, Ziviani N, et al. Using Web Information for Author Name Disambiguation[C]. In: Proceedings of the 9th ACM/IEEE-CS Joint International Conference on Digital Libraries (JCDL'09). New York: ACM, 2009:49-58.[6] Song Y, Huang J, Councill I G, et al. Efficient Topic-based Unsupervised Name Disambiguation[C]. In: Proceedings of the 7th ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL'07). New York: ACM, 2007:342-351.[7] 蒲旭, 王建勇, 范小明. GHOST:作者名字排歧系统[J]. 计算机研究与发展, 2010,47(S1):512-515.(Pu Xu, Wang Jianyong, Fan Xiaoming. GHOST: An Author Name Disambiguation System[J]. Journal of Computer Research and Development, 2010,47(S1):512-515.)[8] DBLP[EB/OL].[2013-04-13]. http://www.informatik.uni-trier.de/~ley/db/index.html.[9] Lucene[EB/OL].[2013-04-04]. http://lucene.apache.org/.[10] Manning C D, Raghavan P, Schütze H. Introduction to Information Retrieval[M]. New York: Cambridge University Press, 2008.[11] Cota R G, Ferreira A A, Nascimento C, et al. An Unsupervised Heuristic-based Hierarchical Method for Name Disambiguation in Bibliographic Citations[J].Journal of the American Society for Information Science and Technology, 2010, 61(9):1853-1870.[12] Robertson S. Understanding Inverse Document Frequency: On Theoretical Argument for IDF[J]. Journal of Documentation, 2004, 60(5):503-520.[13] 肖晶, 梁冰, 张晓丹, 等. 一种面向篇级数据的作者名消歧规则和算法[J]. 现代图书情报技术, 2012(5):55-59.(Xiao Jing, Liang Bing, Zhang Xiaodan, et al. Author Disambiguation Rules and Algorithm for Article Level Data[J]. New Technology of Library and Information Service, 2012(5):55-59.)