Quantization
NVIDIA Jetson Memory Tricks Let Edge Devices Run 10B Parameter AI Models
NVIDIA reveals optimization techniques that reclaim up to 12GB of memory on Jetson devices, enabling multi-billion parameter LLMs to run on edge hardware.
Enhancing AI Model Efficiency with Quantization Aware Training and Distillation
Explore how Quantization Aware Training (QAT) and Quantization Aware Distillation (QAD) optimize AI models for low-precision environments, enhancing accuracy and inference performance.
Enhancing Large Language Models: NVIDIA's Post-Training Quantization Techniques
NVIDIA's post-training quantization (PTQ) advances performance and efficiency in AI models, leveraging formats like NVFP4 for optimized inference without retraining, according to NVIDIA.
Nexa AI Enhances DeepSeek R1 Distill Performance with NexaQuant on AMD Platforms
Nexa AI introduces NexaQuant technology for DeepSeek R1 Distills, optimizing performance on AMD platforms with improved inference capabilities and reduced memory footprint.
QTIP Revolutionizes LLM Quantization with Enhanced Speed and Quality
QTIP introduces a novel approach to LLM quantization, enhancing speed and quality by utilizing trellis coded quantization. Discover how it outperforms previous methods and its implications for memory-bound inference.