Skip to search form
Skip to main content
Skip to account menu
Semantic Scholar
Semantic Scholar's Logo
Search 238,208,002 papers from all fields of science
Search
Sign In
Create Free Account
Web crawler
Known as:
Webcrawler
, Crawl site
, RBSE
Expand
A Web crawler is an Internet bot which systematically browses the World Wide Web, typically for the purpose of Web indexing (web spidering). Web…
Expand
Wikipedia
(opens in a new tab)
Create Alert
Alert
Related topics
Related topics
50 relations
Ajax (programming)
Apache Hadoop
Apache Nutch
Apache Storm
Expand
Papers overview
Semantic Scholar uses AI to extract papers important to this topic.
Review
2013
Review
2013
Intelligent Web Crawling
Denis Shestakov
The IEEE intelligent informatics bulletin
2013
Corpus ID: 11326025
—Web crawling, a process of collecting web pages in an automated manner, is the primary and ubiquitous operation used by a large…
Expand
2011
2011
An improved topic relevance algorithm for focused crawling
Hongwei Hao
,
Cui-Xia Mu
,
Xu-Cheng Yin
,
Shen Li
,
Zhi-Bin Wang
IEEE International Conference on Systems, Man and…
2011
Corpus ID: 36558401
Topic relevance of pages and hyperlinks is the key issue in focused crawling. In this paper, an improved topic relevance…
Expand
2009
2009
Quantifying performance and quality gains in distributed web search engines
B. B. Cambazoglu
,
Vassilis Plachouras
,
R. Baeza-Yates
Annual International ACM SIGIR Conference on…
2009
Corpus ID: 17559264
Distributed search engines based on geographical partitioning of a central Web index emerge as a feasible solution to the immense…
Expand
2007
2007
Analyzing peer behavior in KAD
Moritz Steiner
2007
Corpus ID: 21226260
2007
2007
HiNRG Technical Report: 01-10-2007 Measuring the Storm Worm Network
S. Sarat
,
A. Terzis
2007
Corpus ID: 17889337
The Storm worm is a botnet which appeared in the early months of 2007. Its prolific growth, the use of decentralized command and…
Expand
2007
2007
Discovering Web Communities in the Blogspace
Ying Zhou
,
Joseph G. Davis
Hawaii International Conference on System…
2007
Corpus ID: 206703211
With the emergence of a range of second generation Internet based services such as Weblogs and their hosting services, many…
Expand
2002
2002
Asynchronous maturation of the sexes may limit close inbreeding in a subsocial spider
Todd C. Bukowski
,
L. Avilés
2002
Corpus ID: 54816128
We studied the temporal patterns of maturation and sexual receptivity of a subsocial spider, Anelosimus cf. jucundus, in southern…
Expand
1999
1999
Eecient Web Spidering with Reinforcement Learning
Jason Renniey
,
Andrew McCallumzy
1999
Corpus ID: 17330918
Consider the task of exploring the Web in order to nd pages of a particular kind or on a particular topic. This task arises in…
Expand
1995
1995
A revision of the tracheline spiders (Araneae, Corinnidae) of southern South America. American Museum novitates ; no. 3128
N. Platnick
,
C. Ewing
1995
Corpus ID: 260498832
1964
1964
An Investigation of Certain Components of the Venom of the Female Sydney Funnel Web Spider, Atrax Robustus Cambr.
C. M. Gilbo
,
N. Coles
1964
Corpus ID: 55249548
Venom of the female fUnnel web spider, A. robustus, was heated and diaIysed. From the diffusate was isolated a toxic component…
Expand