Extraction of Structured Rules from Web Pages and Maintenance of Mutual Consistency: XRML Approach

نویسندگان

  • Juyoung Kang
  • Jae Kyu Lee
چکیده

Web pages provide valuable knowledge for human comprehension in text, tables, and mathematical notations. However, the extraction and maintenance of structured rules from the Web pages are not easy tasks. To tackle these problems, we adopt the eXtensible Rule Markup Language framework. The RIML (Rule Identification Markup Language) and RSML (Rule Structure Markup Language) are two compliant representations in XRML for this purpose. RIML identifies the implicit rules in the Web pages possibly using multiple pages to make a rule or rule group. RSML specifies the complete rule structure to be processed by software agents or expert systems. In this study, we cover the natural text, tables, and implicit numeric functions in the texts. In order to fulfill the research goal, we define the necessary tags for the rule extraction and maintenance in XRML. Typical ones include tags for rule grouping, tabular rules, numeric operators, and functions. The rule acquisition process consists of rule base design, rule identification with RIML, and rule structuring with RSML. The maintenance process for the revisions that may occur either in Web pages and structured rules is also described. The approach is demonstrated with the shipping cost comparison on the electronic book stores.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

The Performance of Rule Identification from Web Pages

In the world of Web pages, there are oceans of documents in natural language texts and tables. To extract rules from Web pages and maintain consistency between them, we have developed the framework of XRML (eXtensible Rule Markup Language). XRML allows the identification of rules on Web pages and generates the identified rules automatically. For this purpose, we have designed the Rule Identific...

متن کامل

Enhanced Knowledge Management with eXtensible Rule Markup Language

XML has become the standard platform for structured data exchange on the Web. Next concern of Semantic Web is the exchange of rules in markup language form. The rules should be represented in such a way as to allow software agents to process and browse them for human comprehension. For this purpose, we propose a language eXtensible Rule Markup Language (XRML). XRML is composed of rule identific...

متن کامل

Presenting a method for extracting structured domain-dependent information from Farsi Web pages

Extracting structured information about entities from web texts is an important task in web mining, natural language processing, and information extraction. Information extraction is useful in many applications including search engines, question-answering systems, recommender systems, machine translation, etc. An information extraction system aims to identify the entities from the text and extr...

متن کامل

Contextual Data Extraction and Instance-Based Integration

We propose a formal framework for an unsupervised approach tacking at the same time two problems: the data extraction problem, for generating the extraction rules needed to gain data from web pages, and the data integration problem, to integrate the data coming from several sources. We motivate the approach by discussing its advantages with regard to the traditional “waterfall approach”, in whi...

متن کامل

Rule-based personalized comparison shopping including delivery cost

Comparison shopping allows customers to reduce time and effort when searching for product information and prices. However, traditional comparison sites mainly compare product prices without using precise information on delivery cost. To overcome this limitation, we adopted a rule-based comparison shopping framework using the eXtensible Rule Markup Language (XRML) architecture, which computes th...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2003