Massive Technical Interviews Tips: How to avoid getting into infinite loops when designing a web crawler

Wednesday, December 24, 2014

How to avoid getting into infinite loops when designing a web crawler | Runhe Tian Coding Practice

How to avoid getting into infinite loops when designing a web crawler | Runhe Tian Coding Practice
First, how does the crawler get into a loop? The answer is very simple: when we re-parse an already parsed page. This would mean that we revisit all the links found in that page, and this would continue in a circular fashion.
Be careful about what the interviewer considers the "same" page. Is it URL or content? One could easily get redirected to a previously crawled page.
So how do we stop visiting an already visited page? The web is a graph-based structure, and we commonly use DFS (depth first search) and BFS (breadth first search) for traversing graphs. We can mark already visited pages the same way that we would in a BFS/DFS.
We can easily prove that this algorithm will terminate in any case. We know that each step of the algorithm will parse only new pages, not already visited pages. So, if we assume that we have N number of unvisited pages, then at every step we are reducing N (N-1) by 1. That proves that our algorithm will continue until they are only N steps.
http://www.hawstein.com/posts/12.5.html

http://www.quora.com/How-do-web-crawlers-avoid-getting-into-infinite-loops
By relying on the Canonical URL (Specify your canonical), limiting the crawling depth and in some cases extracting a unique identifier from the scanned page (for example, a product number when crawling an e-commerce site).

http://www.quora.com/What-are-some-Web-crawler-tips-to-avoid-crawler-traps

Honeypot Traps
Don't follow the same crawling pattern
Be polite to your target website

https://tianrunhe.wordpress.com/2012/04/08/how-to-avoid-getting-into-infinite-loops-when-designing-a-web-crawler/
Be careful about what the interviewer considers the “same” page. Is it URL or content? One could easily get redirected to a previously crawled page.
http://www.quora.com/How-do-web-crawlers-avoid-getting-into-infinite-loops
By relying on the Canonical URL (Specify your canonical), limiting the crawling depth and in some cases extracting a unique identifier from the scanned page (for example, a product number when crawling an e-commerce site).

http://stackoverflow.com/questions/5834808/designing-a-web-crawler
Read full article from How to avoid getting into infinite loops when designing a web crawler | Runhe Tian Coding Practice

Wednesday, December 24, 2014

How to avoid getting into infinite loops when designing a web crawler | Runhe Tian Coding Practice

Labels

Popular Posts