Exploiting the Information Web
Business intelligence, information extraction, information retrieval, intelligent systems
Digital Object Identifier (DOI)
The World Wide Web is an increasingly important data source for business decision making; however, extracting information from the Web remains one of the challenging issues related to Web business intelligence applications. To use heterogeneous Web data for decision making, documents containing relevant data must be located, and the data of interest within the documents must be identified and extracted. Currently, most automatic information extraction systems can only cope with a limited set of document formats or do not adapt well to changes in document structure, as a result, many real-world data sources with complex document structures cannot be consistently interpreted using a single information extraction system. This paper presents an adaptive information extraction system prototype that combines multiple information extraction approaches to allow more accurate and resilient data extraction for a wide variety of Web sources. The Amorphic Web information extraction system prototype can locate data of interest based on domain knowledge or page structure, can automatically generate a wrapper for a data source, and can detect when the structure of a Web-based resource has changed and act on this to search the updated resource to locate the desired data. The prototype Amorphic information extraction system demonstrated improved information extraction accuracy for the four different extraction scenarios examined when compared with traditional data extraction approaches.
Was this content written or created while at USF?
Citation / Publisher Attribution
IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), v. 37, issue 1, p. 109-125
Scholar Commons Citation
Gregg, Dawn G. and Walczak, Steven, "Exploiting the Information Web" (2007). School of Information Faculty Publications. 182.