상세 보기
Cluster-Aggregated Transformer: Enhancing lightweight parameter models
- Guo, Zikun;
- Adedigba, Adeyinka P.;
- Mallipeddi, Rammohan
WEB OF SCIENCE
2SCOPUS
3초록
Despite the advancements made by Natural Language Processing (NLP) models in handling complex tasks, capturing long-range dependencies remains a significant challenge, particularly in applications that require long input tokens, such as document and book summarization, machine translation, and sentiment analysis of lengthy user comments. This is because NLP models have fixed but limited input token sizes (typically 512 or 1024 tokens). Hence, they struggle with the complexity and dependencies of extended text sequence sizes. Although many research efforts strive to address this problem and enhance the NLP model's performance, two significant issues remain unresolved: the high memory and computational costs, as well as the performance degradation when the model is compressed for deployment in downstream tasks. This paper introduces the cluster aggregation transformer that uses an attention clustering mechanism to replace the classic transformer with the O(n2 & sdot; d) time complexity. This study not only reduces the high computational cost of processing long text input but also maintains high performance when the model is compressed for downstream tasks, making the model deployment more convenient and faster. The model's efficiency is demonstrated on two public language summarization datasets, namely the Government Report (GovReport) with a token size of 9616 and Book Summarization (BookSum) with a token size of 143,301, in which our method significantly improved the Bidirectional and Auto-Regressive Transformers (BART) model performance. Specifically, our approach achieved 24.5%, 53.6%, and 33.3% improvements on the Recall-Oriented Understudy for Gisting Evaluation (ROUGE) 1, 2 & L scores, respectively, and this performance is maintained up to a 0.7 compression ratio of the BART model, making it suitable for downstream tasks and model deployment. In addition, our approach significantly improves Unlimiformer (the current state-of-the-art), achieving 17.1%, 12.2%, and 5.2% improvements on the ROUGE 1, 2 & L scores, respectively.
키워드
- 제목
- Cluster-Aggregated Transformer: Enhancing lightweight parameter models
- 저자
- Guo, Zikun; Adedigba, Adeyinka P.; Mallipeddi, Rammohan
- 발행일
- 2025-11-01
- 유형
- Article
- 권
- 159
- 언어
- ENG
- 출판사
- PERGAMON-ELSEVIER SCIENCE LTD
- 발행국가
- 영국
- ISSN
- E 1873-6769
P 0952-1976