predict.info — Premium Domain For Sale Domain only: USD 200,000. Prediction platform technology priced separately. predict.info

More from Ethan Mollick | AI News

AI News

Ethan Mollick

@emollick

Professor @Wharton studying AI, innovation & startups. Democratizing education using tech

Gemma 4 Visualizes multimodal thinking

According to emollick, a video shows Gemma 4 12B’s predicted next-token outputs over raw image patches, revealing multimodal attention dynamics. (Source)

07-19-2026 17:11
Kimi K3 Reveals 32 page CoT Loops

According to @emollick, Kimi K3 produced a 32 page chain of thought with loops and dead ends when asked to pick two poems. (Source)

07-19-2026 05:52
Kimi K3 Reveals English Chain-of-Thought Bias

According to emollick, Kimi K3 used 95.5% English in its chain-of-thought for a Chinese prompt, signaling training bias and evaluation gaps. (Source)

07-18-2026 23:24
OpenAI Codex Automates PC Tasks in Minutes

According to @emollick, Codex installed Blender and created a 3D otter animation via computer control with one user click, showcasing rapid PC automation. (Source)

07-18-2026 03:20
Open-weight Kimi Spurs US UK Cyber Risk Warnings

According to @emollick, US and UK view closed-lab models as cyber risks, while China permits strong open weights like Kimi, shifting policy and market dynamics. (Source)

07-17-2026 18:54
AI Security Institute benchmark reveals GLM-5.2 parity

According to Ethan Mollick, the UK AI Security Institute will benchmark Kimi K3 soon; current results show GLM-5.2 matches Opus 4.5 while V4-Pro trails Sonnet 4.5. (Source)

07-17-2026 15:46
Kimi K3 Tops Arena, but Limits Matter

According to emollick, Kimi K3 leads Arena frontend ELO, but user-judged scores are limited for benchmarking and can favor prompt-tuned chat UX. (Source)

07-17-2026 04:11
GPT4 Assistant boosts judges’ throughput 6%

According to Ethan Mollick on X, a GPT4 assistant let Pakistani judges process 6% more cases with no quality drop, citing a new paper by Elliott Ash et al. (Source)

07-17-2026 03:29
Google Gemini 3.5 Delay Signals Competitive Risk

According to emollick, Google is months late on Gemini 3.5 Pro and early retraining results disappointed, especially in coding, per Bloomberg’s Davey Alba. (Source)

07-16-2026 20:13
Procedural benchmarks reveal 4 AI models

According to @emollick, a one-shot procedural harbor town benchmark now compares GPT-5.6 Pro, Fable, Kimi K3, and Inkling. (Source)

07-16-2026 19:38
Kimi K3 Impresses in shader benchmark

According to @emollick, Kimi K3 generated a twigl shader of neo‑gothic ocean towers and improved it on prompt, showing strong open weights performance. (Source)

07-16-2026 15:53
Inkling Model Faces Lem Test Setback

According to @emollick, Inkling struggles on the Lem Test, lagging behind frontier Chinese open weights models seen since DeepSeek r1 and Sonnet 3.5. (Source)

07-16-2026 03:19
ChatGPT 5.6 Sol Pro cracks clue‑less crossword

According to emollick, ChatGPT 5.6 Sol Pro filled a clue‑less Pokémon crossword, signaling rapid multimodal reasoning gains for enterprise use. (Source)

07-15-2026 18:46
Codex Automates Stream Deck Integration Fast

According to @emollick, Codex installed a Stream Deck control tool end to end after GPT-5.6 Pro drafted the plan, reducing manual clicks and setup time. (Source)

07-15-2026 16:35
Codex PowerPoint Breakthrough boosts slide automation

According to @emollick, Codex now builds full PowerPoints but its default skill loads a generic deck that complicates editing of tool instructions. (Source)

07-14-2026 18:50
Machine Gematria Sparks AI Analysis

According to @emollick, interest in “machine gematria” and The Weights is rising, highlighting novel model interpretability angles, per X posts. (Source)

07-14-2026 03:48
Fable AI flags Iliad translation errors

According to emollick, Fable built a ships catalog site and spotted two Butler Iliad Greek translation errors, echoing GPT4 mapping work. (Source)

07-14-2026 02:26
OpenRouter Data Sparks 50% China Token Surge Debate

According to emollick, OpenRouter shows US firms using nearly 50% China tokens, but DKThomp argues agentic tools distort true model usage. (Source)

07-13-2026 20:43
Codex Computer Use Hits Stunning Milestone

According to @emollick, Codex controlled a PC for 5 hours to beat Slay the Spire 2’s daily challenge, proving robust autonomous computer use. (Source)

07-13-2026 19:13
AI Economy Statement Urges Immediate Data Action

According to emollick, economists and tech leaders must gather work impact data now to shape incentives and guardrails for AI’s economic transformation. (Source)

07-13-2026 15:29
Loading...
World Cup