Description
Now Publishers Inc Web Crawling by Christopher Olston, Marc Najork
This is a survey of the science and practice of web crawling. While at first glance web crawling may appear to be merely an application of breadth-first-search, the truth is that there are many challenges ranging from systems concerns such as managing very large data structures, to theoretical questions such as how often to revisit evolving content sources._x000D__x000D_This survey outlines the fundamental challenges and describes the state-of-the-art models and solutions. It also highlights avenues for future work._x000D_ Table of contents :- _x000D_
1: Introduction 2: Crawler Architecture 3: Crawl Ordering Problem 4: Batch Crawl Ordering 5: Incremental Crawl Ordering 6: Avoiding Problematic and Undesirable Content 7: Deep Web Crawling 8: Future Directions. References._x000D_