Mining Top-K Path Traversal Patterns over Streaming Web Click-Sequences

نویسندگان

  • Hua-Fu Li
  • Suh-Yin Lee
چکیده

Online, one-pass mining Web click streams poses some interesting computational issues, such as unbounded length of streaming data, possibly very fast arrival rate, and just one scan over previously arrived Web click-sequences. In this paper, we propose a new, single-pass algorithm, called DSM-TKP (Data Stream Mining for Top-K Path traversal patterns), for mining a set of top-k path traversal patterns, where k is the desired number of path traversal patterns to be mined. An effective summary data structure, called TKP-forest (a forest of Top-K Path traversal patterns), is used to maintain the essential information about the top-k path traversal patterns generated so far. Experimental studies show that the proposed DSM-TKP algorithm uses stable memory usage and makes only one pass over the streaming Web click-sequences.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

DSM-PLW: Single-pass mining of path traversal patterns over streaming Web click-sequences

Mining Web click streams is an important data mining problem with broad applications. However, it is also a difficult problem since the streaming data possess some interesting characteristics, such as unknown or unbounded length, possibly a very fast arrival rate, inability to backtrack over previously arrived click-sequences, and a lack of system control over the order in which the data arrive...

متن کامل

A Top-Down Algorithm for Mining Maximal Traversal Paths in Web Log Sessions

Mining of frequent traversal paths in web logs is an application of sequence mining and useful with many applications that include web recommendation, caching, pre-fetching etc. Most of the existing algorithms follow a bottom-up approach to mine sequence patterns in a database. In this paper, a fast top-down algorithm is presented to discover maximal traversal paths which are contiguous sequenc...

متن کامل

Clickstreams, The Basis to Establish User Navigation Patterns on Web Sites

Collecting and mining clickstream data from e-commerce sites has become increasingly important for marketing, advertising, and traffic analysis activities. Organizations are promoting many initiatives concerning user’s navigation pattern discovering, in order to implement better sites, more functional and close to customers’ needs. Basically, the main idea is to provide more quality of attendan...

متن کامل

Web Users Session Analysis Using DBSCAN and Two Phase Utility Mining Algorithms

One of the important issues in data mining is the interestingness problem. Typically, in a data mining process, the number of patterns discovered can easily exceed the capabilities of a human user to identify interesting results. To address this problem, utility measures have been used to reduce the patterns prior to presenting them to the user. A frequent itemset only reflects the statistical ...

متن کامل

Data Mining for Path Traversal Patterns in a

In this paper, we explore a new data mining capability which involves mining path traversal patterns in a distributed information providing environment like worldwide web. First, we convert the original sequence of log data into a set of maximal forward references and lter out the eeect of some backward references which are mainly made for ease of traveling. Second, we derive algorithms to dete...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:
  • J. Inf. Sci. Eng.

دوره 25  شماره 

صفحات  -

تاریخ انتشار 2009