Volume 10, Number 5

A Hybrid Learning Algorithm in Automated Text Categorization of Legacy Data

  Authors

Dali Wang1, Ying Bai2 and David Hamblin1, 1Christopher Newport University, USA and 2Johnson C. Smith University, USA

  Abstract

The goal of this research is to develop an algorithm to automatically classify measurement types from NASA’s airborne measurement data archive. The product has to meet specific metrics in term of accuracy, robustness and usability, as the initial decision-tree based development has shown limited applicability due to its resource intensive characteristics. We have developed an innovative solution that is much more efficient while offering comparable performance. Similar to many industrial applications, the data available are noisy and correlated; and there is a wide range of features that are associated with the type of measurement to be identified. The proposed algorithm uses a decision tree to select features and determine their weights.A weighted Naive Bayes is used due to the presence of highly correlated inputs. The development has been successfully deployed in an industrial scale, and the results show that the development is well-balanced in term of performance and resource requirements.

  Keywords

classification, machine learning, atmospheric measurement