NVIDIA Launches Accelerated Llama 3.2 Models for Edge and Cloud Deployment

Jessie A Ellis Sep 25, 2024 22:15

NVIDIA introduces Llama 3.2 models with enhanced vision language capabilities, optimized for deployment from edge devices to cloud infrastructure.

NVIDIA Launches Accelerated Llama 3.2 Models for Edge and Cloud Deployment

NVIDIA has expanded its open-source Meta Llama collection with the release of Llama 3.2, which includes vision language models (VLMs), small language models (SLMs), and an updated Llama Guard model. These new models, when paired with NVIDIA's accelerated computing platform, offer developers, researchers, and enterprises enhanced capabilities for generative AI applications, according to the NVIDIA Technical Blog.

Enhanced Capabilities of Llama 3.2

The Llama 3.2 models, trained on NVIDIA H100 Tensor Core GPUs, are available in various sizes. The SLMs come in 1B and 3B sizes, ideal for deploying AI assistants on edge devices. The VLMs, available in 11B and 90B sizes, support both text and image inputs, making them suitable for applications that require visual grounding, reasoning, and understanding, such as image captioning, visual Q&A, and document Q&A.

The Llama Guard models have also been updated to support image input guardrails in addition to text input. All Llama 3.2 models use an auto-regressive language model architecture optimized for inference with long context lengths of up to 128K tokens.

Performance Optimization with NVIDIA TensorRT

NVIDIA is accelerating the Llama 3.2 models using NVIDIA TensorRT to reduce cost and latency while providing high throughput. The TensorRT-LLM library supports the 1B and 3B models with long-context support via scaled rotary position embedding (RoPE) and other optimizations like KV caching and in-flight batching. The 11B and 90B models, which include a vision encoder and text decoder, are optimized using ONNX export to build TensorRT engines that maximize GPU utilization.

Deploying with NVIDIA NIM Microservices

The optimized Llama 3.2 models can be deployed using NVIDIA NIM microservices, which facilitate the deployment of generative AI models across various NVIDIA-accelerated infrastructures, including cloud and data centers. These microservices support production-ready deployments and offer simplified management and orchestration of AI workloads.

Customization and Evaluation with NVIDIA AI Foundry and NeMo

NVIDIA AI Foundry provides a comprehensive platform for customizing Llama 3.2 models. Developers can fine-tune models on proprietary data to achieve better performance in specific tasks. NVIDIA NeMo offers tools for curating training data and applying advanced tuning techniques, ensuring the models meet specific accuracy and safety requirements.

Scaling Local Inference with NVIDIA RTX and Jetson

Llama 3.2 models are optimized for deployment on over 100 million NVIDIA RTX PCs and workstations. For edge deployments, SLMs support techniques like distillation and quantization to reduce memory and computational requirements. The 11B VLM is supported on embedded Jetson AGX Orin devices, making it suitable for video analytics and robotics applications.

NVIDIA's commitment to open-source contributions ensures that the Llama 3.2 models and associated tools are optimized for community use, promoting transparency and widespread adoption in AI safety and resilience efforts.

Image source: Shutterstock