text-dedup is a Python library that enables efficient deduplication of large text corpora by using MinHash and other probabilistic techniques to detect near-duplicate content. This is especially useful for NLP tasks where duplicated training data can skew model performance. text-dedup scales to billions of documents and offers tools for chunking, hashing, and comparing text efficiently with low memory usage. It supports Jaccard similarity thresholding, parallel execution, and flexible deduplication strategies, making it ideal for cleaning web-scraped data, language model training datasets, or document archives.
Features
- Fast and scalable near-duplicate detection
- Uses MinHash and Jaccard similarity for fuzzy matching
- Designed for web-scale datasets with billions of documents
- Supports customizable deduplication thresholds
- Multi-threaded and memory-efficient processing
- Hashing-based representation of text chunks
- Optional GPU acceleration for faster computation
- Suitable for cleaning NLP and LLM training data
Categories
Stream ProcessingLicense
Apache License V2.0Follow text-dedup
Other Useful Business Software
Enterprise-grade ITSM, for every business
Freshservice is an intuitive, AI-powered platform that helps IT, operations, and business teams deliver exceptional service without the usual complexity. Automate repetitive tasks, resolve issues faster, and provide seamless support across the organization. From managing incidents and assets to driving smarter decisions, Freshservice makes it easy to stay efficient and scale with confidence.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of text-dedup!