Skip to main navigation Skip to search Skip to main content

Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework

  • Maolin Wang
  • , Jun Chu
  • , Sicong Xie
  • , Xiaoling Zang
  • , Yao Zhao
  • , Wenliang Zhong
  • , Xiangyu Zhao*
  • *Corresponding author for this work
  • City University of Hong Kong
  • Ant Group

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In the era of mobile computing, deploying efficient Natural Language Processing (NLP) models in resource-restricted edge settings presents significant challenges, particularly in environments requiring strict privacy compliance, real-time responsiveness, and diverse multi-tasking capabilities. These challenges create a fundamental need for ultra-compact models that maintain strong performance across various NLP tasks while adhering to stringent memory constraints. To this end, we introduce Edge ultra-lIte BERT framework (EI-BERT) with a novel cross-distillation method. EI-BERT efficiently compresses models through a comprehensive pipeline including hard token pruning, cross-distillation, parameter quantization, and plugin-and-play deployment. Specifically, the cross-distillation method uniquely positions the teacher model to understand the student model's perspective, ensuring efficient knowledge transfer through parameter integration and the mutual interplay between models. Through extensive experiments, we achieve a remarkably compact BERT-based model of only 1.91 MB - the smallest to date for Natural Language Understanding (NLU) tasks. This ultra-compact model has been successfully deployed across multiple scenarios within the Alipay ecosystem, demonstrating significant improvements in real-world applications. For example, it has been integrated into Alipay's live Edge Recommendation system since January 2024, currently serving the app's recommendation traffic across 8.4 million daily active devices.

Original languageEnglish
Title of host publicationKDD 2025 - Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining
PublisherAssociation for Computing Machinery
Pages4975-4985
Number of pages11
ISBN (Electronic)9798400714542
DOIs
StatePublished - 3 Aug 2025
Externally publishedYes
Event31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025 - Toronto, Canada
Duration: 3 Aug 20257 Aug 2025

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Volume2
ISSN (Print)2154-817X

Conference

Conference31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025
Country/TerritoryCanada
CityToronto
Period3/08/257/08/25

Keywords

  • alipay
  • knowledge distillation
  • language model compression
  • model deployment
  • natural language processing
  • natural language understanding

Fingerprint

Dive into the research topics of 'Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework'. Together they form a unique fingerprint.

Cite this