
BioSnowball: Automated Population of Wikis
课程网址: http://videolectures.net/kdd2010_nie_bsap/  
主讲教师: Zaiqing Nie
开课单位: 微软公司
开课时间: 2010-10-01
课程语种: 英语
课程简介: Internet users regularly have the need to find biographies and facts of people of interest. Wikipedia has become the first stop for celebrity biographies and facts. However, Wikipedia can only provide information for celebrities because of its neutral point of view (NPOV) editorial policy. In this paper we propose an integrated bootstrapping framework named BioSnowball to automatically summarize the Web to generate Wikipedia-style pages for any person with a modest web presence. In BioSnowball, biography ranking and fact extraction are performed together in a single integrated training and inference process using Markov Logic Networks (MLNs) as its underlying statistical model. The bootstrapping framework starts with only a small number of seeds and iteratively finds new facts and biographies. As biography paragraphs on the Web are composed of the most important facts, our joint summarization model can improve the accuracy of both fact extraction and biography ranking compared to decoupled methods in the literature. Empirical results on both a small labeled data set and a real Web-scale data set show the effectiveness of BioSnowball. We also empirically show that BioSnowball outperforms the decoupled methods.
关 键 词: 互联网用户; 维基百科; 马尔可夫逻辑网络; 引导框架
课程来源: 视频讲座网
最后编审: 2020-04-13:chenxin
阅读次数: 47