Global ETD Search

1	Exploring Methods for Comparing Similarity of Dimensionally Inconsistent Multivariate Numerical Data Micic, Natasha, Neagu, Daniel, Torgunov, Denis, Campean, Felician 28 June 2018 (has links) no / When developing multivariate data classiﬁcation and clustering methodologies for data mining, it is clear that most literature contributions only really consider data that contain consistently the same attributes. There are however many cases in current big data analytics applications where for same topic and even same source data sets there are diﬀering attributes being measured, for a multitude of reasons (whether the speciﬁc design of an experiment or poor data quality and consistency). We deﬁne this class of data a dimensionally inconsistent multivariate data, a topic that can be considered a subclass of the Big Data Variety research. This paper explores some classiﬁcation methodologies commonly used in multivariate classiﬁcation and clustering tasks and considers how these traditional methodologies could be adapted to compare dimensionally inconsistent data sets. The study focuses on adapting two similarity measures: Robinson-Foulds tree distance metrics and Variation of Information; for comparing clustering of hierarchical cluster algorithms (such clusters are derived from the raw multivariate data). The results from experiments on engineering data highlight that adapting pairwise measures to exclude non-common attributes from the traditional distance metrics may not be the best method of classiﬁcation. We suggest that more specialised metrics of similarity are required to address challenges presented by dimensionally inconsistent multivariate data, with speciﬁc applications for big engineering data analytics. / Jaguar Land-Rover Big data Clustering Heterogeneous data sets Classiﬁcation methodologies Inconsistent multivariate data

Search results

Exploring Methods for Comparing Similarity of Dimensionally Inconsistent Multivariate Numerical Data