Machine Learning

AI Inference Costs Drop 40% With New GPU Optimization Tactics
Machine Learning

AI Inference Costs Drop 40% With New GPU Optimization Tactics

Together AI reveals production-tested techniques cutting inference latency by 50-100ms while reducing per-token costs up to 5x through quantization and smart decoding.

GPU Waste Crisis Hits AI Production as Utilization Drops Below 50%
Machine Learning

GPU Waste Crisis Hits AI Production as Utilization Drops Below 50%

New analysis reveals production AI workloads achieve under 50% GPU utilization, with CPU-centric architectures blamed for billions in wasted compute resources.

Anthropic Releases Full AI Constitution for Claude Under Open License
Machine Learning

Anthropic Releases Full AI Constitution for Claude Under Open License

Anthropic publishes Claude's complete training constitution under CC0 license, detailing AI safety priorities and ethical guidelines as company eyes $350B valuation.

LangChain Tackles AI Agent Observability Gap With New Insights Tool
Machine Learning

LangChain Tackles AI Agent Observability Gap With New Insights Tool

LangChain launches Insights Agent to analyze 100k+ daily traces from AI agents, addressing the critical gap between data collection and actionable understanding.

GitHub Copilot Gains Cross-Agent Memory System in Public Preview
Machine Learning

GitHub Copilot Gains Cross-Agent Memory System in Public Preview

GitHub launches memory feature for Copilot agents, enabling AI assistants to learn from past interactions and share knowledge across coding, CLI, and code review workflows.

LangChain Unveils Four Multi-Agent Architecture Patterns for AI Development
Machine Learning

LangChain Unveils Four Multi-Agent Architecture Patterns for AI Development

LangChain releases comprehensive guide to multi-agent AI systems, detailing subagents, skills, handoffs, and router patterns with performance benchmarks.

Multi-Node GPU Training Guide Reveals 72B Model Scaling Secrets
Machine Learning

Multi-Node GPU Training Guide Reveals 72B Model Scaling Secrets

Together.ai details how to train 72B parameter models across 128 GPUs, achieving 45-50% utilization with proper network tuning and fault tolerance.

Selecting the Optimal Open-Source Model for Production Applications
Machine Learning

Selecting the Optimal Open-Source Model for Production Applications

Explore the criteria for choosing the right open-source model for production, balancing quality, cost, and speed, while considering legal and technical factors.

Character.ai Unveils Efficient Techniques for Large-Scale Pretraining
Machine Learning

Character.ai Unveils Efficient Techniques for Large-Scale Pretraining

Character.ai reveals innovative methods for optimizing large-scale pretraining, focusing on techniques like Squinch, dynamic clamping, and Gumbel Softmax, to enhance efficiency in AI model training.

Revolutionizing Semiconductor Defect Detection with AI-Powered Models
Machine Learning

Revolutionizing Semiconductor Defect Detection with AI-Powered Models

NVIDIA leverages generative AI and vision foundation models to enhance semiconductor defect classification, addressing limitations of traditional CNNs and improving manufacturing efficiency.

Trending topics