سال انتشار: ۱۳۸۲

محل انتشار: نهمین کنفرانس سالانه انجمن کامپیوتر ایران

تعداد صفحات: ۸

نویسنده(ها):

Ahmad Abdollahzadeh Barfourosh – Computer Eng. & IT Faculty , Amirkabir University of Technology Tehran, Iran
Hamid Reza Motahari Nezhad – Computer Eng. & IT Faculty , Amirkabir University of Technology Tehran, Iran

چکیده:

The main contribution of this paper is introducing an approach for expanding the crawling methods of Cora spider, as a RL-based spider. We have introduced novel methods for calculating the Q-Value in reinforcement learning module of the spider. The proposed crawlers can find the target pages faster and earn more rewards over the crawl than Cora’s crawlers. We have used support Vector Machines (SVMs) classifier for the first time as a text learner in Web crawlers and compared the results with crawlers which use Naïve Bayes (NB) classifier for this purpose. The results show that crawlers using SVMs outperform crawlers which use NB in the first half of crawling a web site and find the target pages more quickly. The test bed for the evaluation of our approaches was Web sites of four computer science departments of four universities, which have been made available offline.