<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Meixia Tao | SJTU Media Lab</title>
    <link>https://medialab.sjtu.edu.cn/author/meixia-tao/</link>
      <atom:link href="https://medialab.sjtu.edu.cn/author/meixia-tao/index.xml" rel="self" type="application/rss+xml" />
    <description>Meixia Tao</description>
    <generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sat, 25 Jul 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://medialab.sjtu.edu.cn/media/icon_hu8d5d95fcee90aa32aafdcbe9e477e4aa_72543_512x512_fill_lanczos_center_3.png</url>
      <title>Meixia Tao</title>
      <link>https://medialab.sjtu.edu.cn/author/meixia-tao/</link>
    </image>
    
    <item>
      <title>Token Compression in the AI Agent Lifecycle: A Comprehensive Survey from Perceptual Inputs to Semantic Contexts</title>
      <link>https://medialab.sjtu.edu.cn/post/26-07-25-token-compression-in-the-ai-agent-lifecycle/</link>
      <pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://medialab.sjtu.edu.cn/post/26-07-25-token-compression-in-the-ai-agent-lifecycle/</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Manuscript submitted to IEEE Communications Surveys &amp;amp; Tutorials   &lt;a href=&#34;A_Survey_of_Token_Compression_in_the_AI_Agent_Lifecycle_Preprint_v0725.pdf&#34; class=&#34;btn btn-primary btn-sm&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;&lt;i class=&#34;fas fa-file-pdf pr-1&#34;&gt;&lt;/i&gt;PDF 下载&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&#34;authors&#34;&gt;Authors&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Fengxi Zhang&lt;/strong&gt;&lt;sup&gt;1&lt;/sup&gt;, &lt;strong&gt;Zhengxue Cheng&lt;/strong&gt;&lt;sup&gt;1,*&lt;/sup&gt;, &lt;strong&gt;Guo Lu&lt;/strong&gt;&lt;sup&gt;1&lt;/sup&gt;, &lt;strong&gt;Li Song&lt;/strong&gt;&lt;sup&gt;1,*&lt;/sup&gt;, &lt;strong&gt;Zhu Li&lt;/strong&gt;&lt;sup&gt;2&lt;/sup&gt;, &lt;strong&gt;Zhiyong Chen&lt;/strong&gt;&lt;sup&gt;1&lt;/sup&gt;, &lt;strong&gt;Meixia Tao&lt;/strong&gt;&lt;sup&gt;1&lt;/sup&gt;, and &lt;strong&gt;Wenjun Zhang&lt;/strong&gt;&lt;sup&gt;1&lt;/sup&gt;&lt;/p&gt;
&lt;h2 id=&#34;affiliations&#34;&gt;Affiliations&lt;/h2&gt;
&lt;p&gt;&lt;sup&gt;1&lt;/sup&gt; School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, Shanghai 200240, China&lt;/p&gt;
&lt;p&gt;&lt;sup&gt;2&lt;/sup&gt; School of Science and Engineering, University of Missouri–Kansas City, Kansas City, MO 64110, USA&lt;/p&gt;
&lt;p&gt;&lt;sup&gt;*&lt;/sup&gt; Corresponding authors: Zhengxue Cheng and Li Song&lt;/p&gt;
&lt;h2 id=&#34;abstract&#34;&gt;Abstract&lt;/h2&gt;
&lt;p&gt;Tokens serve as the fundamental interactive units through which foundation models represent inputs, maintain contexts, and support reasoning, thereby enabling communication and collaboration between humans and AI systems. In AI agents, tokens arise not only from current perceptual inputs, but also from workflow contexts accumulated across multi-step execution, including retrieval results, reasoning traces, action-observation histories, and memory records. Dense multimodal inputs and iterative workflow accumulation lead to token explosion, increasing inference cost and making active-context management critical for efficient and reliable agent execution. This survey reviews token compression from the perspective of the AI agent lifecycle. We propose an active context optimization problem formulation for token compression, which aims to reduce the token cost of active contexts while preserving task utility under limited context window and computation budgets. We then organize existing methods according to where tokens arise in the lifecycle. Perception compression targets current input-side tokens derived from text, image, video, and audio, and is reviewed through transformation, token selection, token aggregation, and token resampling. Semantic compression targets workflow contexts introduced or accumulated during agent execution, covering retrieval, thought, action-observation, and memory token compression. Finally, we discuss open challenges and future directions from a unified AI-agent perspective, highlighting token compression as an essential technique for managing active contexts in long-horizon agent execution. A repository of token compression methods is available at: &lt;a href=&#34;https://github.com/XCJinggai/Awesome_Token_Compression_in_AI_Agent_Lifecycle&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://github.com/XCJinggai/Awesome_Token_Compression_in_AI_Agent_Lifecycle&lt;/a&gt; .&lt;/p&gt;
&lt;h2 id=&#34;keywords&#34;&gt;Keywords&lt;/h2&gt;
&lt;p&gt;Token compression; AI agents; large language models; multimodal large language models.&lt;/p&gt;
&lt;h2 id=&#34;figures-preview&#34;&gt;Figures Preview&lt;/h2&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;&#34; srcset=&#34;
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/01_hu3f91d60bfcd598f9f4d063281939d3f3_201346_a273920f967de4530d1c282513b3ac49.webp 400w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/01_hu3f91d60bfcd598f9f4d063281939d3f3_201346_5b908830276f13b28311b831fe4e36df.webp 760w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/01_hu3f91d60bfcd598f9f4d063281939d3f3_201346_1200x1200_fit_q75_h2_lanczos_3.webp 1200w&#34;
               src=&#34;https://medialab.sjtu.edu.cn/post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/01_hu3f91d60bfcd598f9f4d063281939d3f3_201346_a273920f967de4530d1c282513b3ac49.webp&#34;
               width=&#34;760&#34;
               height=&#34;649&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;center style=&#34;color:#888888&#34;&gt;Figure 1. Token Explosion in the AI Agent Lifecycle. As agents execute long-horizon tasks, dense multimodal perceptual inputs and multi-turn interactions continuously expand the active context, resulting in token explosion. This rapid token growth substantially increases token cost and inference latency. Additionally, excessive context length can degrade task performance and increase the risk of task failure. &lt;/center&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;&#34; srcset=&#34;
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/02_hu1817b54c33d4045e1b1df66baaac4d35_313113_59aaa6cd50e06add90837ede836dc1a8.webp 400w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/02_hu1817b54c33d4045e1b1df66baaac4d35_313113_4b9e1ac3da710707bd14bcbbc48dc3c5.webp 760w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/02_hu1817b54c33d4045e1b1df66baaac4d35_313113_1200x1200_fit_q75_h2_lanczos_3.webp 1200w&#34;
               src=&#34;https://medialab.sjtu.edu.cn/post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/02_hu1817b54c33d4045e1b1df66baaac4d35_313113_59aaa6cd50e06add90837ede836dc1a8.webp&#34;
               width=&#34;760&#34;
               height=&#34;571&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;center style=&#34;color:#888888&#34;&gt;Figure 2. Token Explosion in the AI Agent Lifecycle. As agents execute long-horizon tasks, dense multimodal perceptual inputs and multi-turn interactions continuously expand the active context, resulting in token explosion. This rapid token growth substantially increases token cost and inference latency. Additionally, excessive context length can degrade task performance and increase the risk of task failure.&lt;/center&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;&#34; srcset=&#34;
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/03_hu99f4896932c0d5741555ae9c60c8c3c1_343339_1ba3bd7e9be609cf97537f453bad2ae9.webp 400w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/03_hu99f4896932c0d5741555ae9c60c8c3c1_343339_17ea5f3c9a62d3969efebea1728b3f14.webp 760w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/03_hu99f4896932c0d5741555ae9c60c8c3c1_343339_1200x1200_fit_q75_h2_lanczos_3.webp 1200w&#34;
               src=&#34;https://medialab.sjtu.edu.cn/post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/03_hu99f4896932c0d5741555ae9c60c8c3c1_343339_1ba3bd7e9be609cf97537f453bad2ae9.webp&#34;
               width=&#34;760&#34;
               height=&#34;695&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;center style=&#34;color:#888888&#34;&gt;Figure 4. Agent-centric taxonomy of token compression. Token compression is organized into perception compression for current perceptual inputs by users and semantic compression for contexts accumulated across the AI agent lifecycle. Perception compression is further categorized by compression mechanisms, while semantic compression follows the agent workflow.&lt;/center&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;&#34; srcset=&#34;
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/04_hub5586bb58bf5faa5b2ca66c735ff691c_139641_eea8c2bdd221b3394f6470cc21bd1b13.webp 400w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/04_hub5586bb58bf5faa5b2ca66c735ff691c_139641_94a4810ffd8f7c57de2c6faa48abccc5.webp 760w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/04_hub5586bb58bf5faa5b2ca66c735ff691c_139641_1200x1200_fit_q75_h2_lanczos_3.webp 1200w&#34;
               src=&#34;https://medialab.sjtu.edu.cn/post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/04_hub5586bb58bf5faa5b2ca66c735ff691c_139641_eea8c2bdd221b3394f6470cc21bd1b13.webp&#34;
               width=&#34;760&#34;
               height=&#34;226&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;center style=&#34;color:#888888&#34;&gt;Figure 5. Representative application scenarios of token compression. (a) Human-MLLM interaction compresses multimodal perceptual tokens transmitted from users for efficient real-time responses. (b) Human-agent collaboration compresses active contexts exchanged across execution steps for reliable long-horizon task completion. (c) Multi-agent collaboration compresses exchanged semantic messages for scalable inter-agent coordination.&lt;/center&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;&#34; srcset=&#34;
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/05_hu14f149ad0fb2d7e620317014af56c7bc_138132_48eb4fc161dea4a3348f9d6611835f3b.webp 400w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/05_hu14f149ad0fb2d7e620317014af56c7bc_138132_0716fb0d557859158bb40bbba9c6dec6.webp 760w,
               /post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/05_hu14f149ad0fb2d7e620317014af56c7bc_138132_1200x1200_fit_q75_h2_lanczos_3.webp 1200w&#34;
               src=&#34;https://medialab.sjtu.edu.cn/post/26-07-25-token-compression-in-the-ai-agent-lifecycle/figs/05_hu14f149ad0fb2d7e620317014af56c7bc_138132_48eb4fc161dea4a3348f9d6611835f3b.webp&#34;
               width=&#34;760&#34;
               height=&#34;267&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;center style=&#34;color:#888888&#34;&gt;Figure 6. Future directions for token compression in agent workflows, organized from method design through evaluation to deployment.&lt;/center&gt;
</description>
    </item>
    
  </channel>
</rss>
