LLM

Evaluating LLMs for Production: Lessons from GitHub's Secret Scanning
LLM

Evaluating LLMs for Production: Lessons from GitHub's Secret Scanning

GitHub shares best practices for evaluating LLMs before production, focusing on precision, recall, and real-world workflows to improve security tools.

NVIDIA Dynamo's Shadow Engine Slashes LLM Recovery to 7 Seconds
LLM

NVIDIA Dynamo's Shadow Engine Slashes LLM Recovery to 7 Seconds

NVIDIA's Shadow Engine Recovery in Dynamo cuts LLM downtime from minutes to 7 seconds, transforming AI infrastructure resilience.

NVIDIA NeMo Switchyard Enables Smarter AI Model Routing
LLM

NVIDIA NeMo Switchyard Enables Smarter AI Model Routing

NVIDIA NeMo Switchyard optimizes AI workflows by intelligently routing tasks across models, balancing cost, latency, and accuracy.

AMD Silo AI Expands Poro 2 Models with Enhanced Context and Math Capabilities
LLM

AMD Silo AI Expands Poro 2 Models with Enhanced Context and Math Capabilities

AMD Silo AI releases Poro 2 8B Long with a 16x larger context window and Poro 2 8B Math Reasoning for multilingual AI development.

Together AI Unveils Advanced Autoscaling for LLM Inference
LLM

Together AI Unveils Advanced Autoscaling for LLM Inference

Together AI introduces autoscaling features tailored for large language models, optimizing GPU use and managing latency during traffic spikes.

NVIDIA Optimizes JAX LLM Training with Host Offloading
LLM

NVIDIA Optimizes JAX LLM Training with Host Offloading

NVIDIA's host offloading for JAX LLM training boosts GPU memory efficiency, enabling larger batch sizes and faster throughput.

NVIDIA Pushes Hardware-Aware LLM Co-Design for AI Efficiency
LLM

NVIDIA Pushes Hardware-Aware LLM Co-Design for AI Efficiency

NVIDIA highlights hardware-friendly AI model design to optimize LLM performance. Learn how co-design boosts throughput, latency, and cost-efficiency.

ParallelKernelBench Exposes LLM Weakness in Multi-GPU Kernels
LLM

ParallelKernelBench Exposes LLM Weakness in Multi-GPU Kernels

ParallelKernelBench shows GPT-5.5 and peers struggle with multi-GPU CUDA kernels, solving less than 31% of tasks. Here's why it matters.

Ray Serve LLM Enhances Distributed Inference with 24x Boost
LLM

Ray Serve LLM Enhances Distributed Inference with 24x Boost

Ray Serve LLM achieves 24x higher throughput with new direct streaming, HAProxy integration, and vLLM backend upgrades, pushing LLM inference forward.

Claude Opus Aims to Revolutionize Source Code Security with LLMs
LLM

Claude Opus Aims to Revolutionize Source Code Security with LLMs

Anthropic's Claude Opus 4.7 showcases its ability to find and patch source code vulnerabilities, positioning it as a powerful tool for secure software development.

Trending topics