0


PSkip:从网络搜索点击的数据相关性来排序质量

PSkip: Estimating relevance ranking quality from web search clickthrough data
课程网址: http://videolectures.net/kdd09_wang_pserrqwscd/  
主讲教师: Kuansan Wang
开课单位: 微软公司
开课时间: 2009-08-14
课程语种: 英语
中文简介:
52002:SYSTEM ERROR
课程简介: In this article, we report our efforts in mining the information encoded as clickthrough data in the server logs to evaluate and monitor the relevance ranking quality of a commercial web search engine. We describe a metric called pSkip that aims to quantify the ranking quality by estimating the probability of users encountering non relevant results that cost them the efforts to read and skip. A search engine with a lower pSkip is regarded as having a better ranking quality. A key design goal of pSkip is to integrate the findings from two sets of user studies that utilize eye-tracking devices to track users browsing patterns on the search result pages, and that use specially instrumented browsers to actively solicit users explicit judgments on their search activities. We present the derivation of the maximum likelihood estimation of pSkip and demonstrate its efficacy in describing the user study data. The mathematical properties of pSkip are further analyzed and compared with several objective metrics as well as the cumulated gain method that uses subjective judgments. Experimental data show that pSkip can measure aspects of the search quality that these existing metrics are not designed or fail to address, such as identifying the real search intents expressed in the ambiguous queries. Although effective and superior in many ways, we also report a series of experiments that show pSkip may be influenced by system issues that are not directly related to relevance ranking, suggesting that measurements complementary to pSkip are still needed in order to form a holistic and accurate characterization of the ranking quality.
关 键 词: 计算机科学; Web挖掘; 搜索引擎
课程来源: 视频讲座网
最后编审: 2020-04-20:chenxin
阅读次数: 44