Global ETD Search

Return to search

Mixtures of Skew-t Factor Analyzers

Model-based clustering allows for the identification of subgroups in a data set through the use of finite mixture models. When applied to high-dimensional microarray data, we can discover groups of genes characterized by their gene expression profiles. In this thesis, a mixture of skew-t factor analyzers is introduced for the clustering of high-dimensional data. Notably, we make use of a version of the skew-t distribution which has not previously appeared in mixture-modelling literature. Allowing a constraint on the factor loading matrix leads to two mixtures of skew-t factor analyzers models. These models are implemented using the alternating expectation-conditional maximization algorithm for parameter estimation with an Aitken's acceleration stopping criterion used to determine convergence. The Bayesian information criterion is used for model selection and the performance of each model is assessed using the adjusted Rand index. The models are applied to both real and simulated data, obtaining clustering results which are equivalent or superior to those of established clustering methods.

http://hdl.handle.net/10214/5274

Identifer	oai:union.ndltd.org:LACETR/oai:collectionscanada.gc.ca:OGU.10214/5274
Date	11 1900
Creators	Murray, Paula
Contributors	McNicholas, Paul
Source Sets	Library and Archives Canada ETDs Repository / Centre d'archives des thèses électroniques de Bibliothèque et Archives Canada
Language	English
Detected Language	English
Type	Thesis

Page generated in 0.0021 seconds

Mixtures of Skew-t Factor Analyzers

Description

Links & Downloads

Tags

Additional Fields