
Large Scale High-Precision Topic Modeling on Twitter
课程网址: http://videolectures.net/kdd2014_yang_twitter/  
主讲教师: Shuang-Hong Yang
开课单位: 推特股份有限公司
开课时间: 2014-10-07
课程语种: 英语
课程简介: We are interested in organizing a continuous stream of sparse and noisy texts, known as "tweets", in real time into an ontology of hundreds of topics with measurable and stringently high precision. This inference is performed over a full-scale stream of Twitter data, whose statistical distribution evolves rapidly over time. The implementation in an industrial setting with the potential of affecting and being visible to real users made it necessary to overcome a host of practical challenges. We present a spectrum of topic modeling techniques that contribute to a deployed system. These include non-topical tweet detection, automatic labeled data acquisition, evaluation with human computation, diagnostic and corrective learning and, most importantly, high-precision topic inference. The latter represents a novel two-stage training algorithm for tweet text classification and a close-loop inference mechanism for combining texts with additional sources of information. The resulting system achieves 93% precision at substantial overall coverage.
关 键 词: 统计分布; 主题建模; 推文检测
课程来源: 视频讲座网
数据采集: 2023-03-26:chenxin01
最后编审: 2023-05-22:chenxin01
阅读次数: 26