Addressing Context Fragmentation in Mamba-Based RAG Using Adaptive Semantic Chunking for Retrieval Accuracy and Linear Efficiencytive Semantic Chunking
DOI:
https://doi.org/10.30812/bite.v8i1.6326Keywords:
Adaptive Semantic Chunking , Contextual Relevance, Cosine Similarity, Mamba, RAGAbstract
Background: The limitations of Large Language Models (LLMs) regarding context relevance and computational efficiency in handling long text sequences present a major challenge in implementing Retrieval-Augmented Generation (RAG).
Objective: This study aims to optimize contextual relevance through an Adaptive Semantic Chunking mechanism integrated into the Mamba linear sequential model architecture.
Methods: Unlike static (fixed-size) cutting methods, the proposed algorithm dynamically determines text chunk boundaries based on semantic coherence thresholds () using Cosine Similarity. Experiments were conducted using the WikiQA dataset to evaluate retrieval accuracy and inference efficiency.
Result: The results demonstrate that a value of = 0.7 represents the optimal point, producing intact semantic units with a relevance score of 0.7037. In terms of performance, this integration enables the Mamba model to achieve highly efficient inference times of 0.0412 seconds with linear time complexity.
Conclusion: This adaptive approach successfully eliminates information fragmentation and minimizes noise within the Mamba model's hidden states. This study concludes that adaptive semantic grouping significantly enhances information density and answer accuracy in Mamba-based RAG systems, offering crucial implications for the development of real-time question-answering systems
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Irwan Darmawan, Citra Siwi Hanayanti, Nilam Ramadhani, Yuliana Trisanti

This work is licensed under a Creative Commons Attribution 4.0 International License.
Irwan Darmawan








