Events

Qualifying Exam

Efficient and Scalable Near-Duplicate Text Alignment Algorithms

 

Download as iCal file

Friday, July 24, 2026, 11:00am - 12:00pm

 

Speaker: Yuheng Zhang

Bio

Location : CoRE 301

Committee

Associate Professor Dong Deng

Associate Professor Yongfeng Zhang

Professor Lirong Xia

Assistant Professor Qiong Zhang

Event Type: Qualifying Exam

Abstract: Near-duplicate text alignment locates and aligns the specific subsequences of one text that are near-duplicates of subsequences of another—a core operation for large-language-model data-contamination detection, training-data deduplication, and retrieval-augmented generation. The problem is computationally hard because a text has quadratically many subsequences and the underlying Jaccard / weighted Jaccard similarities are costly to evaluate. This exam studies near-duplicate text alignment from three complementary algorithmic directions: (i) one-permutation hashing, which computes passage signatures in a single pass under Jaccard similarity; (ii) consistent weighted sampling, which extends alignment to weighted Jaccard similarity with provably optimal grouping; and (iii) LSHAlign, a locality-sensitive-hashing scheme that further supports the all-pair alignment setting with linear time and space complexity. Extensive experiments on real text collections demonstrate that the proposed methods substantially reduce index-construction time, index size, and query latency compared with prior approaches, while preserving alignment accuracy. This work offers practical algorithmic foundations for scaling near-duplicate text analysis to the large corpora used in modern language-model pipelines.

Organization

Contact  Associate Professor Dong Deng

Zoom Link: https://rutgers.zoom.us/j/7958351323?pwd=WlhjaWhQeW9VaWZGZmcyRFc0dlNCQT09