Publications

Journal

Incorporating receiver operating characteristics into naive Bayes for unbalanced data classification

페이지 정보

profile_image
작성자 관리자
조회 1,665회 작성일 20-12-19 21:29

본문

Journal Computing, 99(3), 203-208.
Name Kim, T., Chung, B. D. and Lee, J.-S.
Year 2017

Naive Bayesian classification has been widely used in data mining area because of its simplicity and robustness to missing values and irrelevant attributes. However, naive Bayes classifiers sometimes show poor performance due to their unrealistic assumption that all attributes are equally important and conditionally independent of each other. In this research, we dispense with the former assumption by proposing a new attribute weighting method. The proposed method considers each attribute as a single classifier and measures its discriminating ability using the area under an ROC curve (AUC). Each AUC value is then used to weight the corresponding attribute. In addition, we try to reduce the complexity of classification models by selecting high AUC attributes. Using 20 real datasets from the machine learning repository at UC Irvine (UCI), we conduct a numerical experiment to show that the proposed method is an improvement over standard naive Bayes classification and existing weighting methods.