MARATTO

article

A Survey of Model Compression Techniques for TinyML Applications

202514 citations

Abstract

The convergence of embedded systems and artificial intelligence has led to the rapid growth of TinyML, which enables ML inference on highly resource-constrained devices. However, deploying modern DL models in these environments poses signif-icant challenges due to limited memory, processing power, and energy availability. Model compression techniques have emerged as essential tools to bridge this gap, allowing for the deployment of efficient, low-latency, and accurate models at the edge. This survey provides a comprehensive and structured review of key model compression strategies-including low-rank factorization, neural architecture search, pruning, quantization, knowledge distillation, and other emerging methods-highlighting their principles, mathematical formulations, and application relevance to TinyML. We also discuss the trade-offs and challenges as-sociated with these methods, such as accuracy loss, hardware-software co-design, and deployment constraints. Our goal is to equip researchers and practitioners with a practical reference that informs the design of lightweight AI solutions suitable for the next generation of intelligent edge devices.

Research topics

  • Algorithms and Data Compression

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/iccsc66714.2025.11135279

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.