Token Compression in the AI Agent Lifecycle: A Comprehensive Survey from Perceptual Inputs to Semantic Contexts
Manuscript submitted to IEEE Communications Surveys & Tutorials PDF 下载
Authors
Fengxi Zhang1, Zhengxue Cheng1,*, Guo Lu1, Li Song1,*, Zhu Li2, Zhiyong Chen1, Meixia Tao1, and Wenjun Zhang1
Affiliations
1 School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, Shanghai 200240, China
2 School of Science and Engineering, University of Missouri–Kansas City, Kansas City, MO 64110, USA
* Corresponding authors: Zhengxue Cheng and Li Song
Abstract
Tokens serve as the fundamental interactive units through which foundation models represent inputs, maintain contexts, and support reasoning, thereby enabling communication and collaboration between humans and AI systems. In AI agents, tokens arise not only from current perceptual inputs, but also from workflow contexts accumulated across multi-step execution, including retrieval results, reasoning traces, action-observation histories, and memory records. Dense multimodal inputs and iterative workflow accumulation lead to token explosion, increasing inference cost and making active-context management critical for efficient and reliable agent execution. This survey reviews token compression from the perspective of the AI agent lifecycle. We propose an active context optimization problem formulation for token compression, which aims to reduce the token cost of active contexts while preserving task utility under limited context window and computation budgets. We then organize existing methods according to where tokens arise in the lifecycle. Perception compression targets current input-side tokens derived from text, image, video, and audio, and is reviewed through transformation, token selection, token aggregation, and token resampling. Semantic compression targets workflow contexts introduced or accumulated during agent execution, covering retrieval, thought, action-observation, and memory token compression. Finally, we discuss open challenges and future directions from a unified AI-agent perspective, highlighting token compression as an essential technique for managing active contexts in long-horizon agent execution. A repository of token compression methods is available at: https://github.com/XCJinggai/Awesome_Token_Compression_in_AI_Agent_Lifecycle .
Keywords
Token compression; AI agents; large language models; multimodal large language models.
Figures Preview




