Global ETD Search

Return to search

Cell Classification for Layout Recognition in Spreadsheets

Spreadsheets compose a notably large and valuable dataset of documents within the enterprise settings and on the Web. Although spreadsheets are intuitive to use and equipped with powerful functionalities, extracting and reusing data from them remains a cumbersome and mostly manual task. Their greatest strength, the large degree of freedom they provide to the user, is at the same time also their greatest weakness, since data can be arbitrarily structured. Therefore, in this paper we propose a supervised learning approach for layout recognition in spreadsheets. We work on the cell level, aiming at predicting their correct layout role, out of five predefined alternatives. For this task we have considered a large number of features not covered before by related work. Moreover, we gather a considerably large dataset of annotated cells, from spreadsheets exhibiting variability in format and content. Our experiments, with five different classification algorithms, show that we can predict cell layout roles with high accuracy. Subsequently, in this paper we focus on revising the classification results, with the aim of repairing misclassifications. We propose a sophisticated approach, composed of three steps, which effectively corrects a reasonable number of inaccurate predictions.

info:eu-repo/classification/ddc/004

ddc:004

Identifer	oai:union.ndltd.org:DRESDEN/oai:qucosa:de:qucosa:75562
Date	28 July 2021
Creators	Koci, Elvis, Thiele, Maik, Romero, Oscar, Lehner, Wolfgang
Publisher	Springer
Source Sets	Hochschulschriftenserver (HSSS) der SLUB Dresden
Language	English
Detected Language	English
Type	info:eu-repo/semantics/acceptedVersion, doc-type:conferenceObject, info:eu-repo/semantics/conferenceObject, doc-type:Text
Rights	info:eu-repo/semantics/openAccess
Relation	978-3-319-99701-8, 10.1007/978-3-319-99701-8_4, info:eu-repo/grantAgreement/European Commission/Erasmus Mundus Joint Doctorate/IT4BI-DC//Information Technologies for Business Intelligence - Doctoral College

Page generated in 0.013 seconds

Cell Classification for Layout Recognition in Spreadsheets

Description

Links & Downloads

Tags

Additional Fields