
Explanation of SVM's behaviour in text classification
课程网址: http://videolectures.net/solomon_colas_esvm/  
主讲教师: Fabrice Colas
开课单位: 莱顿大学
开课时间: 2007-08-24
课程语种: 英语
课程简介: We are concerned with the problem of learning classification rules in text categorization where many authors presented Support Vector Machines (SVM) as leading classification method. Number of studies, however, repeatedly pointed out that in some situations SVM is outperformed by simpler methods such as naive Bayes or nearest-neighbor rule. In this paper, we aim at developing better understanding of SVM behaviour in typical text categorization problems represented by sparse bag of words feature spaces. We study in details the performance and the number of support vectors when varying the training set size, the number of features and, unlike existing studies, also SVM free parameter C, which is the Lagrange multipliers upper bound in SVM dual. We show that SVM solutions with small C are high performers. However, most training documents are then bounded support vectors sharing a same weight C . Thus, SVM reduce to a nearest mean classifier; this raises an interesting question on SVM merits in sparse bag of words feature spaces. Additionally, SVM suffer from performance deterioration for particular training set size/number of features combinations.
关 键 词: 分类规则; 特征空间袋; 乘数上限
课程来源: 视频讲座网
最后编审: 2019-09-21:cwx
阅读次数: 34