Redundant feature elimination for multi-class problemsAnnalisa Appice, Michelangelo Ceci, Simon Rawles, Peter Flach, Redundant feature elimination for multi-class problems. Proceedings of the 21st International Conference on Machine Learning (ICML 2004). Russ Greiner , Dale Schuurmans, (eds.). ISBN 1-58113-838-5, pp. 33–40. July 2004. PDF, 1056 Kbytes.
We consider the problem of eliminating redundant Boolean features for a given data set, where a feature is redundant if it separates the classes less well than another feature or set of features. Lavrac et al. proposed the algorithm REDUCE that works by pairwise comparison of features, i.e., it eliminates a feature if it is redundant with respect to another feature. Their algorithm operates in an ILP setting and is restricted to two-class problems. In this paper we improve their method and extend it to multiple classes. Central to our approach is the notion of a neighbourhood of examples: a set of examples of the same class where the number of different features between examples is relatively small. Redundant features are eliminated by applying a revised version of the REDUCE method to each pair of neighbourhoods of different class. We analyse the performance of our method on a range of data sets.