Smol AI News
167 articles archived · Visit source ↗ · RSS
-
Smol AI News news-outlet 2mo ago
not much happened today
**Anthropic's Mythos/Opus cycle** sparked mixed reactions with praise for **Claude Mythos**'s one-shot workflows and concerns over **Opus 4.8** benchmark regressions. **Opus 4.7** showed strong chemistry task performance, "making Claude a chemist." **Sakana AI** launched an…
23 -
Smol AI News news-outlet 2mo ago
not much happened today
**NVIDIA** released **Nemotron 3 Ultra**, a fully open **550B MoE** model with **55B active parameters** and **1M context**, optimized for long-running agent tasks with up to **5x speedup** and **30% cost reduction**. It features hybrid Mamba/attention, LatentMoE, native MTP,…
7 -
Smol AI News news-outlet 2mo ago
not much happened today
**Microsoft** released the detailed technical report for **MAI-Thinking-1**, a generalist reasoning model trained without third-party distillation, achieving **97% on AIME 2025** and outperforming Sonnet 4.6 in human preference tests. The report was praised for transparency,…
7 -
Smol AI News news-outlet 2mo ago
not much happened today
**NVIDIA** led open-source AI model releases with **Cosmos 3**, a comprehensive omnimodal world model unifying language, image, video, audio, and action using a Mixture-of-Transformers design, and **Nemotron 3 Ultra**, a **550B** parameter open-weight model noted for high…
33 -
Smol AI News news-outlet 2mo ago
not much happened today
**Anthropic** rolled out **Claude Opus 4.8**, which shows incremental improvements but mixed benchmark results, including better cooperation and coding behavior but some regressions in document parsing. Platform updates include mid-conversation system instructions enhancing long…
29 -
Smol AI News news-outlet 2mo ago
not much happened today
**AI News for 5/23/2026-5/26/2026** highlights a shift in AI product strategy emphasizing **model + harness + workflow + UI + memory + economics** over model quality alone. **OpenAI** released a significant **Codex** update with features like **appshots** and remote computer…
16 -
Smol AI News news-outlet 2mo ago
not much happened today
**Inference optimization** is increasingly architectural, with **EAGLE 3.1** improving speculative decoding and long-context handling, collaborating with **vLLM** and **TorchSpec**. **Perplexity** open-sourced a rebuilt **Unigram tokenizer** cutting CPU use by **5–6×** and…
15 -
Smol AI News news-outlet 2mo ago
not much happened today
**RAEv2** advances representation-first tokenization with **>10x faster convergence** and improved generation, tested on **text-to-image** and **world models**. **NVIDIA's Gated DeltaNet-2** innovates linear attention with channel-wise gates, outperforming **KDA** and…
23 -
Smol AI News news-outlet 2mo ago
not much happened today
**Agent infrastructure** is advancing with **LangSmith Engine** providing CI/CD loops for agents and **SmithDB** enabling low-latency querying for observability. **Cognition's Devin Auto-Triage** offers persistent automation for bug triage with memory and subagent structures.…
27 -
Smol AI News news-outlet 3mo ago
not much happened today
**Cerebras** made headlines with its **IPO**, marking a significant milestone for the company known for its contrarian hardware approach. The **Cerebras CFO Bob Komin** emphasized the company's capability to serve **trillion-parameter models**, including internal **OpenAI 5.4…
35 -
Smol AI News news-outlet 3mo ago
not much happened today
**OpenAI** expanded **Codex** integration with the ChatGPT mobile app enabling remote task management and introduced Remote SSH, hooks, and programmatic tokens for enterprise automation. The IDE ecosystem is shifting to "agent-first" UX with **GitHub Copilot App** preview and…
26 -
Smol AI News news-outlet 3mo ago
not much happened today
**Cline, LangChain, Notion, and Cursor** advanced agent infrastructure and developer platforms with innovations like **Cline SDK**, **LangSmith Engine**, **SmithDB** (offering **12–15×** faster observability), and Notion's External Agents API integrating third-party agents such…
14 -
Smol AI News news-outlet 3mo ago
not much happened today
**Research-level reasoning benchmarks** are advancing with **439 new math problems** from **64 mathematicians** and expanded medical benchmarks in **Medmarks v1.0** covering **30 benchmarks** and **61 models**. **Google DeepMind's AI Co-Mathematician** achieves **48% on…
15 -
Smol AI News news-outlet 3mo ago
not much happened today
**Thinking Machines** previewed their new **native interaction models** designed for **full-duplex multimodal interaction** enabling real-time concurrent listening, speaking, watching, thinking, searching, and reacting, marking a shift beyond turn-based AI. This approach…
36 -
Smol AI News news-outlet 3mo ago
not much happened today
**OpenAI** rapidly expanded the **GPT-5.5** family with multiple variants including **gpt-image-2**, **GPT-5.5 Pro**, and **GPT-5.5 Cyber**, receiving positive feedback for efficiency and usability. **Codex** evolved into a long-running agent runtime with a new **/goal**…
35 -
Smol AI News news-outlet 3mo ago
not much happened today
**OpenAI** rolled out **GPT-5.5 Instant** as the new default for ChatGPT and API, enhancing **factuality, intelligence, image understanding, and tone** with stronger personalization features like saved memories and Gmail integration. OpenAI also shared infrastructure updates on…
28 -
Smol AI News news-outlet 3mo ago
not much happened today
**AI Twitter Recap** highlights the shift from model-centric AI to **context pipelines** and **agent orchestration** as key performance drivers. Notably, **gpt-5.2-codex** and **gpt-5.3-codex** showed significant benchmark improvements through prompt and middleware tuning. The…
16 -
Smol AI News news-outlet 3mo ago
not much happened today
**OpenAI** achieved a major math breakthrough by disproving a long-standing Erdős unit distance problem using a **general-purpose reasoning model**, marking a milestone in AI-driven formal science and long-horizon reasoning. The result was validated by prominent mathematicians…
8 -
Smol AI News news-outlet 3mo ago
not much happened today
**AI News for 5/4/2026-5/5/2026** highlights a shift in AI product development emphasizing **model + harness + workflow + UI + memory + economics** over model quality alone, with notable updates from **OpenAI Codex** and **Claude** including new features like **Appshots**,…
26 -
Smol AI News news-outlet 3mo ago
not much happened today
**xAI released Grok 4.3**, improving cost/performance with a **53 Intelligence Index score**, 4 points higher than Grok 4.20, and significant gains on **GDPval-AA** and **τ²-Bench Telecom**. However, accuracy tradeoffs raised reliability concerns. Community opinions are mixed,…
32 -
Smol AI News news-outlet 3mo ago
not much happened today
**OpenAI's GPT-5.5** achieves top-tier performance in long-horizon cyber tasks, matching or surpassing **Claude Mythos Preview** with a **71.4%** pass rate and showing ongoing improvement beyond **100M tokens** inference. OpenAI also released an **Advanced Account Security**…
32 -
Smol AI News news-outlet 3mo ago
not much happened today
**OpenAI** is expanding **Codex** from a coding tool to a general work surface with persistent context, tools, integrations, and team rollout, including **Codex-only seats with $0 seat fee** for Business/Enterprise customers through June. Performance improvements focus on…
23 -
Smol AI News news-outlet 3mo ago
not much happened today
**vLLM v0.20.0** introduces significant improvements in memory and MoE serving efficiency, including **TurboQuant 2-bit KV cache** for **4× KV capacity** and a **2.1% latency improvement**. The update supports multiple hardware platforms like **DeepSeek V4 MegaMoE on…
9 -
Smol AI News news-outlet 3mo ago
not much happened today
**OpenAI** loosens its **Azure exclusivity**, allowing distribution across **Google TPU**, **AWS Trainium**, and **Bedrock** with commitments through **2032** and revenue share through **2030**. **GPT-5.5** shows improved benchmarks but is not uniformly dominant, ranking…
11 -
Smol AI News news-outlet 3mo ago
DeepSeek v4
**DeepSeek-V4** technical release features a **1.6T-parameter MoE with 49B active parameters** and **1M-token context**, showcasing hybrid attention and compressed KV schemes for major memory reductions. It ranks as the **#2 open-weights reasoning model** behind **Kimi K2.6**…
13 -
Smol AI News news-outlet 3mo ago
GPT 5.5
**OpenAI launched GPT-5.5** as its new flagship model for "real work and powering agents," immediately available in ChatGPT and Codex but with delayed API access due to enhanced safety requirements. The model features improved token efficiency and supports longer multi-step…
14 -
Smol AI News news-outlet 3mo ago
not much happened today
**Alibaba** released **Qwen3.6-27B**, a dense, Apache 2.0 open coding model with thinking and non-thinking modes, outperforming the larger Qwen3.5-397B-A17B on multiple coding benchmarks including SWE-bench and Terminal-Bench. It supports native vision-language reasoning over…
15 -
Smol AI News news-outlet 3mo ago
GPT-Image-2
**OpenAI** launched **GPT-Image-2**, enhancing image generation with improved text rendering, layout fidelity, editing, multilingual support, and "thinking" capabilities. It supports generating slides, infographics, diagrams, UI mockups, and QR codes, and integrates with tools…
36 -
Smol AI News news-outlet 3mo ago
not much happened today
**Moonshot's Kimi K2.6** is a major open-weight **1T-parameter MoE** model featuring **32B active parameters**, **384 experts**, **MLA attention**, **256K context window**, native multimodality, and **INT4 quantization**. It supports day-0 integration with platforms like…
9 -
Smol AI News news-outlet 3mo ago
not much happened today
**Anthropic** launched **Claude Design**, a prototyping tool powered by **Claude Opus 4.7**, targeting design workflows and competing with **Figma** and others. Benchmarks show **Opus 4.7** leading in coding and text tasks, with improved efficiency and adaptive reasoning, though…
7 -
Smol AI News news-outlet 4mo ago
Anthropic's Claude Opus 4.7
**Anthropic** launched **Claude Opus 4.7**, its most capable Opus model yet, featuring stronger coding and agentic performance, a new tokenizer, and improved long-context handling with a new **xhigh** reasoning tier. Benchmarks show substantial gains, including **SWE-bench Pro…
37 -
Smol AI News news-outlet 4mo ago
not much happened today
**OpenAI** expanded its Agents SDK by separating the agent harness from compute/storage, enabling long-running, durable agents with features like file/computer use, skills, memory, and compaction. The harness is now open-source and supports execution via partner sandboxes,…
37 -
Smol AI News news-outlet 4mo ago
not much happened today
**Harness engineering** is emerging as a key discipline in AI agent development, emphasizing components like filesystems, memory, and retries beyond just models. **OpenAI's Codex** is expanding agentic coding workflows beyond software engineering, including codebase…
32 -
Smol AI News news-outlet 4mo ago
not much happened today
**GLM-5.1** has reached **#3 on Code Arena**, surpassing **Gemini 3.1** and **GPT-5.4**, and matching **Claude Sonnet 4.6** in coding performance. **Z.ai** now holds the **#1 open model rank** close to the top overall. The advisor pattern, combining a cheap executor with an…
12 -
Smol AI News news-outlet 4mo ago
not much happened today
**Anthropic's Mythos** and **OpenAI's** upcoming restricted cyber-capable models are central to recent discussions, with debates on their security realism and evaluation methods. **LangChain's Deep Agents deploy** introduces an open memory, model-agnostic agent harness…
36 -
Smol AI News news-outlet 4mo ago
not much happened today
**Meta Superintelligence Labs** launched **Muse Spark**, a natively multimodal reasoning model featuring tool use, visual chain of thought, and multi-agent orchestration. It is live on **meta.ai** and the Meta AI app with a private API preview and plans for open-sourcing future…
29 -
-
Smol AI News news-outlet 4mo ago
not much happened today
**Hermes Agent** is gaining attention as a leading open agent stack with features like self-improving skills, persistent memory, and a self-improvement loop. Its new **Manim skill** enables generation of math/technical animations, expanding agent capabilities. The Hermes…
19 -
Smol AI News news-outlet 4mo ago
not much happened today
**Google** introduced **Skills in Chrome**, enabling reusable browser workflows with Gemini prompts and a library of ready-made Skills, enhancing end-user agentization. **Tencent** teased **HYWorld 2.0**, an open-source 3D world model generating editable scenes from a single…
8 -
Smol AI News news-outlet 4mo ago
not much happened today
**Gemma 4** was launched by **Google** under an **Apache 2.0 license**, marking a significant open-model release focused on **reasoning, agentic workflows, multimodality, and on-device use**. It outperforms models 10x larger and has immediate ecosystem support including…
35 -
Smol AI News news-outlet 4mo ago
Gemma 4
**Google DeepMind** released **Gemma 4**, a family of open-weight, multimodal models with long-context support up to **256K tokens** under an **Apache 2.0 license**, marking a major capability and licensing shift. The lineup includes **31B dense**, **26B MoE (A4B)**, and two…
14 -
Smol AI News news-outlet 4mo ago
not much happened today
**Arcee’s Trinity-Large-Thinking** was released with **open weights under Apache 2.0**, featuring a **400B total / 13B active** model size and strong agentic performance, ranking **#2 on PinchBench**. **Z.ai’s GLM-5V-Turbo** is a **vision coding model** with **native multimodal…
13 -
Smol AI News news-outlet 4mo ago
not much happened today
**Anthropic** introduced **computer use inside Claude Code** for closed-loop verification in a research preview for Pro/Max users, enhancing reliable app iteration. **OpenAI** released a **Codex plugin for Claude Code**, enabling cross-agent composition and signaling a shift…
16 -
Smol AI News news-outlet 4mo ago
not much happened today
**Anthropic** is reportedly introducing a new AI model tier called **Capybara**, which is larger and more intelligent than **Claude Opus 4.6**, showing improved performance in coding, academic reasoning, and cybersecurity. The model is speculated to be around **10 trillion…
38 -
Smol AI News news-outlet 4mo ago
not much happened today
**Anthropic** advances agent infrastructure with a multi-agent harness emphasizing orchestration and "computer use" for complex software environments. **Figma**, **GitHub**, and **Cursor** launch design canvases with direct AI editing, showcasing tool-calling becoming…
12