Can I read Priority Crawler for URL Selection on EtoBox?
Priority Crawler for URL Selection by Khushbu Pareek is a document available to read on EtoBox.
What is Priority Crawler for URL Selection about?
The document describes a priority crawler that was implemented to analyze a large dataset of URLs over 120 days. The dataset contained 16-20000 URLs per day with 3.2 million attributes for each URL. The crawler aimed to selectively choose URLs to keep in a limited queue based on data quality and importance. It did this by calculating entropy and weighted sums of attributes to determine the top 10 URLs for each day, handling the huge anonymized data through preprocessing and feature selection techniques.
- Author
- Khushbu Pareek
- Language
- EN