会议专题

Incomplete Data Classification Based on Multiple Views

  Missing values have negative impacts on big data analysis.However,in absence of extra knowledge,exact imputation can hardly be conducted for many data sets.Therefore,we have to tolerate missing values and perform data mining on incomplete data sets directly.To achieve high quality data mining on incomplete data,we propose a classification approach based on multiple views.We use various complete views of the data set to generate the base classifiers and combine the results of base classifiers.Since the amount of base classifiers will affect the effectiveness and efficiency of the classification,we aim to find proper view sets.We prove that the view set selection problem is an NP-hard problem and develop an approximation algorithm with approximate ratio ln|S| + 1 where S is the feature set of original data set.Extensive experimental results demonstrate the efficiency and effectiveness of the proposed approaches.

Ming Sun Hongzhi Wang Fanshan Meng Jianzhong Li Hong Gao

Departmemt of Computer Science and Technology,Harbin Institute of Technology,Harbin,China

国际会议

International Asia-Pacific Web Conference(第18届国际亚太互联网大会)

苏州

英文

239-250

2016-09-23(万方平台首次上网日期,不代表论文的发表时间)