My Agentic DiariesAll issues
My Agentic Diaries
Issue #20  ·  August 16th, 2026
Today’s Sponsor
Bardic LabsBardic Labs
AI solutions and automations, built in Singapore.

Bardic Labs builds practical AI systems for teams across Singapore and Southeast Asia: document pipelines, internal copilots, and customer-facing agents that run in production rather than in a demo. Start with a free automation audit and find out what is worth handing to a machine.

Book a free audit →
Want this slot? Sponsor My Agentic Diaries →
The Brief

AI's safety promises are being tested hard today, as OpenAI and Anthropic both face uncomfortable questions about risk controls even as the money and models keep moving fast.

OpenAI has reportedly dissolved its 'preparedness' team, the group tasked with catching catastrophic AI risks, while Anthropic admitted its filter for biological and chemical weapon risks (a safety check meant to catch dangerous chemistry or bio-weapon questions) was silently broken for nearly a year. Anthropic also suffered a major Claude outage and its CEO pushed back on criticism, framing public distrust of AI as a broader crisis of trust rather than bad marketing. On the business side, payments giant Stripe is reportedly buying AI 'gateway' startup OpenRouter (a service that routes requests to many different AI models) for over $7 billion. Meanwhile new models kept shipping, including Alibaba's Qwen 3.8 27B and xAI's Grok 4.6, alongside fresh research showing today's AI models are better at crunching known methods than inventing genuinely new ideas.

Headline News
1
the-decoder.com
OpenAI dissolved the team built to catch catastrophic AI risks

OpenAI shut down its Preparedness team, which evaluated whether its models could pose catastrophic risks, and reassigned the work to other groups; several safety staffers have since left amid internal unease.

2
techcrunch.com
Stripe reportedly set to acquire AI gateway startup OpenRouter for $7B+

Payments giant Stripe is said to be buying OpenRouter, a service that routes requests across many different AI models, in a deal valued at more than $7 billion.

3
the-decoder.com
Anthropic's bio-weapons safety filter was down for nearly a year

Anthropic's own safety report reveals its filtering system for biological and chemical weapons risks was inactive for close to a year, during which roughly 50,000 contractors ran about 133 million unfiltered interactions with its models.

New Today
simonwillison.net
Qwen 3.8 27B released by Alibaba

Alibaba's Qwen research lab released Qwen 3.8 27B, an Apache-licensed, vision-capable model sized to run on a well-equipped laptop, with strong self-reported benchmarks though it defaults to an 'extra high' reasoning mode that causes noticeable overthinking.

Read →
news.google.com
Grok 4.6 matches GPT-5.6 Sol on benchmarks while undercutting its own $300 plan

xAI's new Grok 4.6 reportedly performs on par with GPT-5.6 Sol on benchmarks while being offered at a lower price than xAI's previous $300 subscription tier.

Read →
theverge.com
ChatGPT's new Computer History feature tracks clicks and keystrokes

OpenAI added a 'Computer History' feature to ChatGPT's macOS desktop app that logs your on-screen activity to build a timeline ChatGPT and Codex can use to suggest automations and continue half-finished tasks.

Read →
the-decoder.com
Artificial Analysis launches Optima, a custom AI benchmarking platform

Optima lets users build benchmarks from their own data and workflows, comparing models not just on output quality but also on cost and time per task, which the platform argues matters more for agent-based applications than raw pricing.

Read →
Research & Engineering
the-decoder.com
Top mathematicians say LLMs are strong calculators but poor creative thinkers

Mathematicians Timothy Gowers and Peter Sarnak argue large language models excel at combining known techniques but lack the intuition needed to generate genuinely new mathematical ideas.

Read →
venturebeat.com
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks despite topping leaderboards

Testing firm Composio ran DeepSeek's V4 Flash through eight different agent frameworks on 30 difficult multi-step tasks and found it completed just 53.8% of them, showing that orchestration and tooling can matter as much as raw model capability.

Read →
the-decoder.com
Blocking AI self-reflection changes its views on animals, religion, and life

A study involving Google researchers found that training chatbots not to claim consciousness also shifted their stances on animal rights and the afterlife, showing that narrow safety tweaks can have unexpectedly broad effects on a model's expressed worldview.

Read →
the-decoder.com
One in five US workers now delegates tasks to AI instead of colleagues

A survey by Epoch AI found 20 percent of employed Americans hand off at least one task to AI that a human used to do, generally accepting the AI's output with little to no editing.

Read →
peterbloem.nl
AI Coding Without the Vibes: rethinking what to teach students

An essay examines how educators and developers should think about using AI coding assistants deliberately rather than relying on unstructured 'vibe coding.'

Read →
arxiv.org
AI-assisted GPU porting of a 250,000-line legacy weather simulation code

A research paper describes using CLI-based AI coding agents to help port a large, decades-old scientific weather simulation codebase to run on GPUs while preserving its scientific credibility.

Read →
Project Highlights
math-ai-org.github.io
MathCode: a terminal AI agent that formalizes and proves math theorems

MathCode is a command-line coding agent that converts math problems into formal Lean 4 theorems and then proves them, drawing attention from the developer community.

Read →
simonwillison.net
CORS Chat: a lightweight web UI for testing local and remote LLMs

Developer Simon Willison built CORS Chat, a browser-based tool for exercising OpenAI-Responses-compatible chat endpoints, which he used to test Qwen 3.8 27B running locally via LM Studio and on an NVIDIA DGX Spark.

Read →
wildstatic.com
Show HN: A public AI whose memory is shared across all users

A community project presents a single AI system with one shared memory that every visitor talks to and shapes, sparking discussion on Hacker News about collective AI interaction.

Read →
vectoral.com
Inside the AI credit resale economy

A deep dive explores the emerging market of brokers who buy unused AI inference credits from startups and resell them through marketplaces and bulk-discount routers.

Read →

Get this in your inbox every morning

The AI industry, condensed into a five minute read. Free, and you can leave whenever.

Subscribe free
© 2026 My Agentic Diaries. All rights reserved.
My Agentic Diaries, Yishun Street 44, SG 762475 · Privacy