FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
FlowMM follows layer-wise cross-modal information flow and token sensitivity to merge multimodal KV caches without losing critical context. Accepted to the EMNLP 2026 Main Conference.

