전체 글 썸네일형 리스트형 [ICLR 2026] When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models https://arxiv.org/abs/2504.02010, ICLR 2026 When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning ModelsCompression methods, including quantization, distillation, and pruning, improve the computational efficiency of large reasoning models (LRMs). However, existing studies either fail to sufficiently compare all three compression methods on LRMs or lac.. 더보기 [ICML 2025] Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers https://arxiv.org/abs/2505.22167 ICML 2025 Accepted Paper Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion TransformersDiffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage r.. 더보기 [ICLR 2025] DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models https://arxiv.org/abs/2501.04304 AbstractText-to-Image 생성을 위한 Diffusion모델은 높은 계산량과 메모리사용으로 인해 실제 적용에 제약을 줌.이를 하기 위해 Quantization 기법이 이용되는데, 기존 Diffusion모델의 quantization은 낮은 비트에서 이미지 품질과 텍스트-이미지 alignment를 유지하는데 한계를 지님.본 논문에서는 activation에서 outlier가 있으며, 이는 이미지 품질을 결정하는데 중요한 역할을 한다는 것을 분석함. 또한, 텍스트-이미지 alignment에 cross-attention이 중요한 역할을 한다는 것을 분석함.논문에서 제안하는 DGQ(Distribution-aware Group Quantizat.. 더보기 [ICLR 2025] ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation https://arxiv.org/abs/2406.02540 ICLR 2025 ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video GenerationDiffusion transformers have demonstrated remarkable performance in visual generation tasks, such as generating realistic images or videos based on textual instructions. However, larger model sizes and multi-frame processing for video generation lead to i.. 더보기 [CVPR 2025] Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers https://arxiv.org/abs/2406.17343 , CVPR 2025 (Accepted) Q-DiT: Accurate Post-Training Quantization for Diffusion TransformersRecent advancements in diffusion models, particularly the architectural transformation from UNet-based models to Diffusion Transformers (DiTs), significantly improve the quality and scalability of image and video generation. However, despite their impressiarxiv.org Abstrac.. 더보기 [ECCV2024] Post-training Quantization for Text-to-Image Diffusion Models with Progressive Calibration and Activation Relaxing https://arxiv.org/abs/2311.06322 , ECCV 2024 Post-training Quantization for Text-to-Image Diffusion Models with Progressive Calibration and Activation RelaxingHigh computational overhead is a troublesome problem for diffusion models. Recent studies have leveraged post-training quantization (PTQ) to compress diffusion models. However, most of them only focus on unconditional models, leaving the q.. 더보기 [ACM SAC 2025] Advanced Knowledge Transfer: Refined Feature Distillation for Zero-Shot Quantization in Edge Computing https://dl.acm.org/doi/abs/10.1145/3672608.3707816Abstract기존 Zero-Shot Quantization(ZSQ, Data-Free Quantizaton) 분야에서는 full-precision(FP) Model로부터 높은 quality의 데이터를 생성하는 데 초점을 두는 연구가 진행되고 있음. 하지만, low-bit(높은 압축률) 환경에서 Quantized Model을 학습할 때는 Quantized Model이 정보 수용량 관련 한계를 갖기 때문에 데이터를 생성하는 기법만으로는 적절한 학습이 이루어지지 않음. 이러한 한계를 개선하기 위해 본 논문에서는 Quantized Model을 효과적으로 학습하기 위한 AKT(Advanced Knowledge Transfe.. 더보기 [ECCV2024] GenQ: Quantization in Low Data Regimes with Generative Synthetic Data https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/02058.pdf ECCV 2024 Abstract저 비트 양자화(Low-bit Quantization)에서 발생하는 오류를 줄이기 위해 훈련데이터를 통한 모델 재학습이 필요함.하지만, 데이터 접근이 어려운 환경에서는 모델 재학습이 불가능함. 이를 개선하기 위해 GenQ라는 새로운 method를 제안하며, 이 방식을 통해 더 real 하고 고해상도의 합성데이터를 생성할 수 있음을 보여줌. 또한, 이 방식이 기존 방식들에 비해 대규모 데이터셋의 데이터를 구현하는데 보이는 한계를 극복함.GenQ는 2가지의 필터링 메커니즘을 통해 합성데이터를 실제 훈련데이터와 밀접하게 일치되도록 함.실험을 통해 GenQ는 데.. 더보기 이전 1 2 3 4 ··· 7 다음