ARTICLE
TITLE

SIMILARITY BASED ENTROPY ON FEATURE SELECTION FOR HIGH DIMENSIONAL DATA CLASSIFICATION

SUMMARY

Abstract Curse of dimensionality is a major problem in most classification tasks. Feature transformation and feature selection as a feature reduction method can be applied to overcome this problem. Despite of its good performance, feature transformation is not easily interpretable because the physical meaning of the original features cannot be retrieved. On the other side, feature selection with its simple computational process is able to reduce unwanted features and visualize the data to facilitate data understanding. We propose a new feature selection method using similarity based entropy to overcome the high dimensional data problem. Using 6 datasets with high dimensional feature, we have computed the similarity between feature vector and class vector. Then we find the maximum similarity that can be used for calculating the entropy values of each feature. The selected features are features that having higher entropy than mean entropy of overall features. The fuzzy k-NN classifier was implemented to evaluate the selected features. The experiment result shows that proposed method is able to deal with high dimensional data problem with average accuracy of 80.5%.

 Articles related

Mardi Siswo Utomo,Edi Winarko    

Abstract— Document similarity can be used as a reference for other information searches similar. So as to reduce the time-re-appointment for information following a similar document. Document similarity search capability is usually implemented on the fea... see more


N. E. Kondruk    

Context. The study is devoted to the development of a flexible mathematical apparatus, which should have a sufficiently wide range ofmeans for grouping objects into different types of similarity measures. This makes it possible, within the framework of t... see more


Herlina Jayadianti, Budi Santosa, Judanti Cahyaning, Shoffan Saifullah, Rafal Drezewski    

Writing errors on e-essay exams reduce scores. Thus, detecting and correcting errors automatically in writing answers is necessary. The implementation of Levenshtein Distance and N-Gram can detect writing errors. However, this process needed a long time ... see more


Rakhmat Arianto,Alwy Abdullah,Usman Nurhasan,Rokhimatul Wakhidah    

Meeting minutes are important because they can track decisions and agreements made during the meeting. Meeting minutes can also be used as a benchmark for whether the meeting objectives have been achieved or not. Minutes are taken during the meeting unti... see more

Revista: SISFORMA

Aszani Aszani,Hayyu Ilham Wicaksono,Uffi Nadzima,Lukman Heryawan    

 The growth of medical records continues to increase and needs to be used to improve doctors' performance in diagnosing a disease. A retrieval method returns proposed information to provide diagnostic recommendations based on symptoms from medical r... see more