GPU
Harnessing AI's Potential with Decentralized Compute Networks
Explore how decentralized compute networks address the rising demand for AI applications, offering scalable solutions through consumer-grade GPUs. Learn about real-world use cases and industry partnerships.
NVIDIA's cuEmbed Boosts GPU Performance for Embedding Lookups
NVIDIA unveils cuEmbed, a CUDA library that significantly enhances embedding lookups on GPUs, promising improved performance for recommendation systems and other applications.
Enhancing Polars GPU Parquet Reader Performance with Chunked Reading and UVM
Explore how Polars GPU Parquet Reader boosts performance using chunked reading and Unified Virtual Memory, enhancing data processing capabilities for large datasets.
The Crucial Role of Host CPUs in AI Workloads
Explore the vital role of host CPUs in AI workloads, highlighting their impact on GPU efficiency and inference times with insights from AMD's latest findings.
Enhancing AI Model Training Efficiency on NVIDIA DGX Cloud
Explore how NVIDIA DGX Cloud optimizes AI model training with minimal downtime and robust error attribution, ensuring efficient large-scale GPU utilization.
DeepSeek-R1 Enhances GPU Kernel Generation with Inference Time Scaling
NVIDIA's DeepSeek-R1 model uses inference-time scaling to improve GPU kernel generation, optimizing performance in AI models by efficiently managing computational resources during inference.
NVIDIA Unveils Enhanced Features in NCCL 2.23 for Improved GPU Communication
NVIDIA's NCCL 2.23 release introduces a new scaling algorithm, accelerated initialization, and a profiler plugin API, optimizing inter-GPU and multinode communication for AI and HPC applications.
Injective (INJ)and Aethir Transform GPU Compute Resources with Tokenization
Injective (INJ)and Aethir collaborate to tokenize GPU compute resources, enhancing access and efficiency in AI and blockchain sectors through a novel tradeable token system.
Warp 1.5.0 Introduces Tile-Based Programming for Enhanced GPU Efficiency
Warp 1.5.0 launches tile-based programming in Python, leveraging cuBLASDx and cuFFTDx for efficient GPU operations, significantly improving performance in scientific computing and simulation.
NVIDIA's RAPIDS cuDF Enhances pandas Through Unified Virtual Memory
NVIDIA's RAPIDS cuDF utilizes Unified Virtual Memory to boost pandas' performance by 50x, offering seamless integration with existing workflows and GPU acceleration.