Skip to search form
Skip to main content
Skip to account menu
Semantic Scholar
Semantic Scholar's Logo
Search 236,570,914 papers from all fields of science
Search
Sign In
Create Free Account
Robots exclusion standard
Known as:
Robot Exclusion Protocol
, Robots exclusion protocol
, Robots exclusion file
Expand
The robots exclusion standard, also known as the robots exclusion protocol or simply robots.txt, is a standard used by websites to communicate with…
Expand
Wikipedia
(opens in a new tab)
Create Alert
Alert
Related topics
Related topics
24 relations
.htaccess
Apache Nutch
Automated Content Access Protocol
Distributed web crawling
Expand
Broader (1)
World Wide Web
Papers overview
Semantic Scholar uses AI to extract papers important to this topic.
2018
2018
What the HAK? Estimating Ranking Deviations in Incomplete Graphs
Helge H olzmann
,
Avishek A nand
,
Megha K hosla
2018
Corpus ID: 51883334
Most real-world graphs collected from the Web like Web graphs and social network graphs are incomplete . This leads to inaccurate…
Expand
2018
2018
Robots.txt y su influencia en las estrategias SEO
E. Ribas
2018
Corpus ID: 208126189
2017
2017
Analysis of Robot Detection approaches for ethical and unethical robots on Web server log
Mitali Srivastava
,
A. Srivastava
,
Rakhi Garg
,
P. Mishra
2017
Corpus ID: 69749006
:Due to proliferation of Web robots, it is becoming important to detect robots on commercial and educational websites. Web robots…
Expand
2013
2013
A Website Owner's Practical Guide to the Wayback Machine
H. Andersen
Journal on Telecommunications and High Technology…
2013
Corpus ID: 40365780
INTRODUCTION ..................................................................................... 251 I. THE WAYBACK MACHINE AND…
Expand
2013
2013
Identification and characterization of crawlers through analysis of web logs
N. Algiriyage
,
Sanath Jayasena
,
G. Dias
,
Amila B. Perera
,
K.H.N.K. Dayananda
IEEE 8th International Conference on Industrial…
2013
Corpus ID: 14043287
Web crawlers are software programs that automatically traverse the hyperlink structure of the world-wide web in order to locate…
Expand
2012
2012
Hotel Information Exposure in Cyberspace: The Case of Hong Kong
Rosanna Leung
,
R. Law
Information and Communication Technologies in…
2012
Corpus ID: 59621899
Search engines are an everyday tool for Internet surfing. They are also a critical factor that affects e-business performance…
Expand
2010
2010
Plain text list of URIs for www.lincoln.ac.uk (as of 14/02/2010)
A. Bilbie
2010
Corpus ID: 59100058
This dataset was generated using this tool http://www.auditmypc.com/xml-sitemap.asp. Only publicly accessible pages are included…
Expand
2008
2008
Generador de robots.txt - Serco Consultores
S. Consultores
2008
Corpus ID: 56538687
2008
2008
File robots.txt
DiGi
2008
Corpus ID: 235993964
2006
2006
ANALYSIS OF THE USAGE STATISTICS OF ROBOTS EXCLUSION STANDARD
S. Ajay
,
Jaliya Ekanayake
2006
Corpus ID: 13936388
Robots Exclusion standard [4] is a de-facto standard that is used to inform the crawlers, spiders or web robots about the…
Expand