Skip to content

Opening book details…

Can I read FlashExtract: a framework for data extraction by examples on EtoBox?

FlashExtract: a framework for data extraction by examples by Vu Le; Sumit Gulwani is a Computer Science article available to read on EtoBox.

What is FlashExtract: a framework for data extraction by examples about?

Various document types that combine model and view (e.g., text files, webpages, spreadsheets) make it easy to organize (possibly hierarchical) data, but make it difficult to extract raw data for any further manipulation or querying. We present a general framework FlashExtract to extract relevant data from semi-structured documents using examples. It includes: (a) an interaction model that allows end-users to give examples to extract various fields and to relate them in a hierarchical organization using structure and sequence constructs. (b) an inductive synthesis algorithm to synthesize the intended program from few examples in __any__ underlying domain-specific language for data extraction that has been built using our specified algebra of few core operators (map, filter, merge, and pair). We describe instantiation of our framework to three different domains: text files, webpages, and spreadsheets. On our benchmark comprising 75 documents, FlashExtract is able to extract intended data using an average of 2.36 examples in 0.84 seconds per field.

Who reads FlashExtract: a framework for data extraction by examples?

It is typically read by researchers, students, and practitioners in Computer Science.

Author
Vu Le; Sumit Gulwani
Publisher
ACM
Published
2014
Language
EN
Field
Computer Science (Physical Sciences)

More by Vu Le; Sumit Gulwani

Browse all works by Vu Le; Sumit Gulwani