
Collaboratively Improving Topic Discovery and Word Embeddings by Coordinating Global and Local Contexts
课程网址: http://videolectures.net/kdd2017_xun_contexts/  
主讲教师: Guangxu Xun
开课单位: 布法罗大学
开课时间: 2017-10-09
课程语种: 英语
课程简介: A text corpus typically contains two types of context information -- global context and local context. Global context carries topical information which can be utilized by topic models to discover topic structures from the text corpus, while local context can train word embeddings to capture semantic regularities reflected in the text corpus. This encourages us to exploit the useful information in both the global and the local context information. In this paper, we propose a unified language model based on matrix factorization techniques which 1) takes the complementary global and local context information into consideration simultaneously, and 2) models topics and learns word embeddings collaboratively. We empirically show that by incorporating both global and local context, this collaborative model can not only significantly improve the performance of topic discovery over the baseline topic models, but also learn better word embeddings than the baseline word embedding models. We also provide qualitative analysis that explains how the cooperation of global and local context information can result in better topic structures and word embeddings.
关 键 词: 词嵌入; 文本语料库; 数据科学
课程来源: 视频讲座网
数据采集: 2023-12-27:wujk
最后编审: 2023-12-27:wujk
阅读次数: 9