Events
Qualifying ExamEfficient and Scalable Near-Duplicate Text Alignment Algorithms |
|
||
Friday, July 24, 2026, 11:00am - 12:00pm |
|||
Speaker: Yuheng Zhang
Bio
Location : CoRE 301
Committee:
Associate Professor Dong Deng
Associate Professor Yongfeng Zhang
Professor Lirong Xia
Assistant Professor Qiong Zhang
Event Type: Qualifying Exam
Abstract: Near-duplicate text alignment locates and aligns the specific subsequences of one text that are near-duplicates of subsequences of another—a core operation for large-language-model data-contamination detection, training-data deduplication, and retrieval-augmented generation. The problem is computationally hard because a text has quadratically many subsequences and the underlying Jaccard / weighted Jaccard similarities are costly to evaluate. This exam studies near-duplicate text alignment from three complementary algorithmic directions: (i) one-permutation hashing, which computes passage signatures in a single pass under Jaccard similarity; (ii) consistent weighted sampling, which extends alignment to weighted Jaccard similarity with provably optimal grouping; and (iii) LSHAlign, a locality-sensitive-hashing scheme that further supports the all-pair alignment setting with linear time and space complexity. Extensive experiments on real text collections demonstrate that the proposed methods substantially reduce index-construction time, index size, and query latency compared with prior approaches, while preserving alignment accuracy. This work offers practical algorithmic foundations for scaling near-duplicate text analysis to the large corpora used in modern language-model pipelines.
Organization:
Contact Associate Professor Dong Deng
Zoom Link: https://rutgers.zoom.us/j/7958351323?pwd=WlhjaWhQeW9VaWZGZmcyRFc0dlNCQT09