About this document
Near Duplicate Document Detection For Large Information Flows by ayushigupta18apr is a document available to read on EtoBox.
The paper presents a new algorithm for detecting near duplicate documents using q-grams to enhance information quality in large data flows. It introduces an efficient indexing method and a similarity measure for comparing document q-gram occurrences, tested in a multifeed news content management system. Experimental results indicate the algorithm
- Author
- ayushigupta18apr
- Language
- EN