LayerExit: Adaptive Intermediate-Layer Exiting for Context Compression
LayerExit adaptively chooses between an efficient intermediate encoder layer and the final layer, improving question-answering accuracy while reducing compression cost.
I am now a second-year PhD candidate at AML Lab, City University of Hong Kong, supervised by Prof. Xiangyu Zhao. I received my bachelor's degree from the School of Artificial Intelligence at Nanjing University.
My research focuses on LLM context compression, with the long-term goal of building AI systems that proactively learn and evolve through deployment (see Ilya Sutskever's discussion of continual learning).
I am currently an intern with ByteDance's Content Consumption team.
For the full publication list, please refer to my Google Scholar.
LayerExit adaptively chooses between an efficient intermediate encoder layer and the final layer, improving question-answering accuracy while reducing compression cost.
DiVA-Former uses visual tokens to distill long serialized tables into compact representations that preserve both structure and fine-grained text.
A single learned Behavior-Equivalent Token can replace a long system prompt while preserving its downstream behavior and sharply reducing prompt overhead.
Threshold Filtering Packing groups related yet diverse samples into training packs to reduce cross-sample interference during supervised fine-tuning.