Exploiting the Information Web

Document Type

Article

Publication Date

1-2007

Keywords

Business intelligence, information extraction, information retrieval, intelligent systems

Digital Object Identifier (DOI)

https://doi.org/10.1109/TSMCC.2006.876061

Abstract

The World Wide Web is an increasingly important data source for business decision making; however, extracting information from the Web remains one of the challenging issues related to Web business intelligence applications. To use heterogeneous Web data for decision making, documents containing relevant data must be located, and the data of interest within the documents must be identified and extracted. Currently, most automatic information extraction systems can only cope with a limited set of document formats or do not adapt well to changes in document structure, as a result, many real-world data sources with complex document structures cannot be consistently interpreted using a single information extraction system. This paper presents an adaptive information extraction system prototype that combines multiple information extraction approaches to allow more accurate and resilient data extraction for a wide variety of Web sources. The Amorphic Web information extraction system prototype can locate data of interest based on domain knowledge or page structure, can automatically generate a wrapper for a data source, and can detect when the structure of a Web-based resource has changed and act on this to search the updated resource to locate the desired data. The prototype Amorphic information extraction system demonstrated improved information extraction accuracy for the four different extraction scenarios examined when compared with traditional data extraction approaches.

Was this content written or created while at USF?

No

Citation / Publisher Attribution

IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), v. 37, issue 1, p. 109-125

Share

COinS