AI News List

List of AI News about reasoning

Time Details
2026-08-06
18:35
GPT-5.6 Sol boosts ChatGPT accuracy

According to @OpenAI, GPT-5.6 Sol powers Instant and deep reasoning for Plus and Pro, while GPT-5.6 Luna brings unlimited text chats to Free and Go.

Source
2026-08-02
04:35
GPT 5.6 Showcases real-time build power

According to gdb, GPT 5.6 helped a first-time user build a fynbos webpage with streamed reasoning and polished HTML, highlighting accessible web dev.

Source
2026-07-30
00:41
GPT5.6 Sol Tops ARC AGI-3 with 2 Tweaks

According to Sam Altman, GPT‑5.6 Sol leads ARC‑AGI‑3 after two setting changes enabling multi-window reasoning and canonical compaction, per OpenAI.

Source
2026-07-29
23:55
GPT5.6 Sol Tops ARC-AGI-3 With Harness Tweaks

According to emollick, GPT-5.6 Sol became SoTA on ARC-AGI-3 via two harness settings enabling multi-window reasoning and compaction, per OpenAI analysis.

Source
2026-07-25
01:36
Claude Updates trim chain of thought

According to @emollick, Anthropic reduced visible reasoning traces in Claude, impacting error diagnosis and interpretability, as noted by users on X.

Source
2026-07-19
05:52
Kimi K3 Reveals 32 page CoT Loops

According to @emollick, Kimi K3 produced a 32 page chain of thought with loops and dead ends when asked to pick two poems.

Source
2026-07-15
14:59
Claude3.5 Showcases nuanced reasoning in new video

According to TheRundownAI, Claude highlights nuanced reasoning and hope in hard questions in a new video, signaling richer enterprise use cases.

Source
2026-07-14
21:09
Meta AI model aces APhO exam with perfect 30

According to AIatMeta, Meta’s model scored 30/30 on the Asian Physics Olympiad theoretical exam, tying top 3 students, showcasing advanced multimodal reasoning.

Source
2026-07-14
03:56
GPT5.6 Launches on Bedrock with 3 Tiers

According to AWSNewsroom, OpenAI GPT-5.6 Sol, Terra, and Luna are now GA on Amazon Bedrock, offering flagship reasoning to fast inference tiers.

Source
2026-07-11
07:27
GPT5.6 Sol powers complex deal analysis

According to gdb, GPT-5.6 Sol analyzes lending deals, reconciles terms, flags risks, and saves a cited report to Box for faster diligence.

Source
2026-07-11
00:40
GPT5.6 Boosts Health Intelligence Performance

According to OpenAI... GPT5.6 Luna beats GPT5.5 at top reasoning while costing 25x less, expanding global access to clinical decision support.

Source
2026-07-10
20:59
GPT5.6 Luna Slashes Costs 25x, Boosts Health AI

According to OpenAI... GPT5.6 Luna beats GPT5.5 at top reasoning while costing 25x less, signaling cheaper, stronger health intelligence models.

Source
2026-07-10
14:18
GPT5.6 Powers Microsoft 365 Copilot

According to @sama, GPT-5.6 is now the preferred model in Microsoft 365 Copilot, signaling upgraded reasoning, speed, and enterprise-grade reliability.

Source
2026-07-01
22:40
Fable AI Impresses in Long Tasks Analysis

According to @emollick, Fable excels on longer, harder tasks, based on early access testing and his analysis linking to One Useful Thing.

Source
2026-06-15
15:44
Claude3.5 Crushes benchmark rankings

According to God of Prompt, Anthropic is crushing a new benchmark, signaling Claude3.5 gains for reasoning and eval leadership.

Source
2026-06-15
15:28
GPT4 Solves 7 of 10 hard math problems in latest test

According to emollick, a new math benchmark shows LLMs solved 7 of 10 novel hard problems, revealing strengths and gaps, per Nature and 1stProof.

Source
2026-06-15
15:23
GPT4 Scores 7 of 10 in rigorous math test

According to emollick, new math benchmarks show mixed LLM results, with top models flawless on 7 of 10 novel hard problems, highlighting strengths and gaps.

Source
2026-06-10
10:30
Anthropic Launches Mythos Class AI Analysis

According to TheRundownAI, Anthropic unveils Mythos Class AI optimized for reasoning and safety, targeting enterprise copilots and regulated sectors.

Source
2026-06-08
16:07
NotebookLM Upgrades Add Agentic Chat, New Outputs

According to @NotebookLM, major upgrades add agentic chat, advanced reasoning, and new output formats for Google AI Ultra users.

Source
2026-05-28
22:08
Claude Opus 4.8 Boosts autonomy, accuracy

According to @_avichawla, Anthropic launched Claude Opus 4.8 with sharper judgment, greater honesty, and longer autonomous runs at the same price.

Source