Search Results for "llm"

NVIDIA Megatron Boosts LLM Training With Muon Optimizer

NVIDIA Megatron Boosts LLM Training With Muon Optimizer

NVIDIA integrates Muon and advanced optimizers into Megatron to enhance large-scale LLM training with near-parity throughput to AdamW.

LLM Agents Help Win Kaggle Competition with 600K Lines of Code

LLM Agents Help Win Kaggle Competition with 600K Lines of Code

Generative AI agents produced 600,000 lines of code and ran 850 experiments to secure first place in a Kaggle competition. Here's how they did it.

Ray Serve Introduces Scalable Multi-Agent AI Architecture

Ray Serve Introduces Scalable Multi-Agent AI Architecture

Ray Serve leverages MCP and A2A protocols for scalable AI agents, solving production bottlenecks in LLM and multi-agent deployments.

Anyscale Launches LLM Post-Training Tool to Simplify Fine-Tuning

Anyscale Launches LLM Post-Training Tool to Simplify Fine-Tuning

Anyscale unveils a post-training skill for large language models, streamlining methodology selection, GPU planning, and configuration generation.

Claude Opus Aims to Revolutionize Source Code Security with LLMs

Claude Opus Aims to Revolutionize Source Code Security with LLMs

Anthropic's Claude Opus 4.7 showcases its ability to find and patch source code vulnerabilities, positioning it as a powerful tool for secure software development.

Ray Serve LLM Enhances Distributed Inference with 24x Boost

Ray Serve LLM Enhances Distributed Inference with 24x Boost

Ray Serve LLM achieves 24x higher throughput with new direct streaming, HAProxy integration, and vLLM backend upgrades, pushing LLM inference forward.

ParallelKernelBench Exposes LLM Weakness in Multi-GPU Kernels

ParallelKernelBench Exposes LLM Weakness in Multi-GPU Kernels

ParallelKernelBench shows GPT-5.5 and peers struggle with multi-GPU CUDA kernels, solving less than 31% of tasks. Here's why it matters.

NVIDIA Pushes Hardware-Aware LLM Co-Design for AI Efficiency

NVIDIA Pushes Hardware-Aware LLM Co-Design for AI Efficiency

NVIDIA highlights hardware-friendly AI model design to optimize LLM performance. Learn how co-design boosts throughput, latency, and cost-efficiency.

NVIDIA Optimizes JAX LLM Training with Host Offloading

NVIDIA Optimizes JAX LLM Training with Host Offloading

NVIDIA's host offloading for JAX LLM training boosts GPU memory efficiency, enabling larger batch sizes and faster throughput.

Together AI Unveils Advanced Autoscaling for LLM Inference

Together AI Unveils Advanced Autoscaling for LLM Inference

Together AI introduces autoscaling features tailored for large language models, optimizing GPU use and managing latency during traffic spikes.

AMD Silo AI Expands Poro 2 Models with Enhanced Context and Math Capabilities

AMD Silo AI Expands Poro 2 Models with Enhanced Context and Math Capabilities

AMD Silo AI releases Poro 2 8B Long with a 16x larger context window and Poro 2 8B Math Reasoning for multilingual AI development.

NVIDIA NeMo Switchyard Enables Smarter AI Model Routing

NVIDIA NeMo Switchyard Enables Smarter AI Model Routing

NVIDIA NeMo Switchyard optimizes AI workflows by intelligently routing tasks across models, balancing cost, latency, and accuracy.

Trending topics