PROJECT TITLE :

Combining Tag and Value Similarity for Data Extraction and Alignment

ABSTRACT:

Web databases generate query result pages based on a user's query. Automatically extracting the data from these query result pages is very important for many applications, such as data integration, which need to cooperate with multiple web databases. We present a novel data extraction and alignment method called CTVS that combines both tag and value similarity. CTVS automatically extracts data from query result pages by first identifying and segmenting the query result records (QRRs) in the query result pages and then aligning the segmented QRRs into a table, in which the data values from the same attribute are put into the same column. Specifically, we propose new techniques to handle the case when the QRRs are not contiguous, which may be due to the presence of auxiliary information, such as a comment, recommendation or advertisement, and for handling any nested structure that may exist in the QRRs. We also design a new record alignment algorithm that aligns the attributes in a record, first pairwise and then holistically, by combining the tag and data value similarity information. Experimental results show that CTVS achieves high precision and outperforms existing state-of-the-art data extraction methods.


Did you like this research project?

To get this research project Guidelines, Training and Code... Click Here


PROJECT TITLE :Combining Solar Energy Harvesting with Wireless Charging for Hybrid Wireless Sensor Networks - 2018ABSTRACT:The appliance of wireless charging technology in ancient battery-powered wireless sensor networks (WSNs)
PROJECT TITLE :An Enhanced MPPT Method Combining Fractional-Order and Fuzzy Logic Control - 2017ABSTRACT:A fractional-order fuzzy logic control (FOFLC) method for maximum power point tracking (MPPT) in a photovoltaic (PV) system
PROJECT TITLE : Combining inertial measurements with blind Image deblurring using distance transform - 2016 ABSTRACT: Camera motion throughout exposure results in a blurry image. We tend to propose an image deblurring method
PROJECT TITLE :Detection of Manhole Covers in High-Resolution Aerial Images of Urban Areas by Combining Two MethodsABSTRACT:Mispositioning of buried utilities is an increasingly important problem both in industrialized and developing
PROJECT TITLE :On Combining Multiple-Instance Learning and Active Learning for Computer-Aided Detection of TuberculosisABSTRACT:The major advantage of multiple-instance learning (MIL) applied to a computer-aided detection (CAD)

Ready to Complete Your Academic MTech Project Work In Affordable Price ?

Project Enquiry