
WIKImage: Correlated image and text datasets
课程网址: http://videolectures.net/sikdd2011_pracner_wikimage/  
主讲教师: Doni Pracner
开课单位: 诺维萨德大学
开课时间: 2011-11-04
课程语种: 日语
课程简介: This paper presents work towards the creation of free and redistributable datasets of correlated images and text. Collections of free images and related text were extracted from Wikipedia with our new tool WIKImage. An additional tool – WIKImage browser – was introduced to visualize the resulting dataset, and was expanded into a manual labeling tool. The paper presents a starting dataset of 1007 images labeled with any combination of 14 tags. The images were processed into a number of scale invariant (SIFT) and color histogram features, and the captions were transformed into a bag-of-words (BOW) representation. Experiments were then performed with the aim of classifying data with respect to each of the labels on dataset variants with just the image information, just the textual data, and both, in order to estimate the difficulty of the dataset in the context of different feature spaces. Results indicate improvements in precision, recall and the F-measure when using the combined representation with support vector machines as well as the k-nearest neighbor classifier with the cosine similarity measure.
关 键 词: 计算机科学; 文本挖掘; 机器学习; 图像分析
课程来源: 视频讲座网
最后编审: 2020-06-06:zyk
阅读次数: 71