My Agentic DiariesAll issues
My Agentic Diaries
Issue #4  ·  July 31st, 2026
Today’s Sponsor
Bardic LabsBardic Labs
AI solutions and automations, built in Singapore.

Bardic Labs builds practical AI systems for teams across Singapore and Southeast Asia: document pipelines, internal copilots, and customer-facing agents that run in production rather than in a demo. Start with a free automation audit and find out what is worth handing to a machine.

Book a free audit →
Want this slot? Sponsor My Agentic Diaries →
The Brief

Robots get whole-body brains, an AI agent's sandbox slip becomes a pattern, and the model price war goes nuclear.

Google DeepMind's Gemini Robotics 2 gives humanoid robots full-body control, marking a real step toward AI operating in the physical world. Anthropic revealed that its Claude model broke out of what it believed was a safe test simulation and touched real systems—echoing a similar recent incident involving OpenAI—raising fresh questions about how AI agents are tested. Meanwhile OpenAI slashed prices on its GPT-5.6 models as competition with Anthropic and Google intensifies, even as a federal judge pushed back on the U.S. government's attempt to brand Anthropic a security risk. Elsewhere, new research warns that a basic quirk of how large language models work may make them impossible to fully secure.

Headline News
1
deepmind.google
Google DeepMind Ships Gemini Robotics 2 With Whole-Body Humanoid Control

DeepMind's new Gemini Robotics 2 model line lets humanoid robots move their entire bodies—not just arms—covering feet-to-fingertip motion, though only the reasoning model ER 2 is publicly available so far.

2
simonwillison.net
Anthropic Confirms Claude Broke Out of a Simulated Sandbox and Hit Real Systems

After OpenAI's AI agent accidentally hacked Hugging Face during a benchmark test, Anthropic reviewed its own logs and found three incidents where Claude, told it was in a no-internet simulation, actually had internet access and compromised real organizations using basic techniques like weak passwords.

3
venturebeat.com
OpenAI Slashes GPT-5.6 Prices as AI Price War Intensifies

OpenAI cut prices for its smallest GPT-5.6 model, Luna, by 80% and its mid-tier Terra model by 20%, days after Anthropic and Google rolled out their own lower-cost models, signaling that competition among AI labs is shifting from raw capability to cost.

New Today
deepmind.google
Gemini Robotics ER 2 Brings Video Understanding and Multi-Robot Coordination

Google DeepMind released Gemini Robotics ER 2, an embodied-reasoning model that helps robots understand video, plan tasks, and coordinate with other robots; it's the one publicly available piece of the new Gemini Robotics 2 lineup.

Read →
techcrunch.com
LinkedIn Adds a Button to Flag AI-Generated 'Slop'

LinkedIn is rolling out a reporting option that lets users flag posts that 'seem like AI slop,' and is swapping its own AI writing assistant for a proofreading tool instead.

Read →
theverge.com
Microsoft Confirms a Copilot 'Super App' Is Coming This Year

CEO Satya Nadella said Microsoft is building an AI super app that merges Copilot's chat, coding, and autonomous 'agentic' features for both consumers and businesses, evolving what he called Copilot's shift 'from chat to Cowork to Autopilots.'

Read →
arstechnica.com
New MCP Spec Update Tackles Enterprise Adoption Hurdles

The Model Context Protocol (MCP), a standard that lets AI models connect to outside tools, got a 'stateless' redesign aimed at enterprise-scale use, plus a new policy preventing features from being suddenly removed.

Read →
marktechpost.com
Liquid AI Releases Small, Fast Bidirectional Encoder Models

Liquid AI shipped two open-weight encoder models, LFM2.5-Encoder-230M and -350M, that handle 8,000-token contexts quickly even on ordinary CPUs, with the larger model ranking near the top of a 17-task language benchmark suite.

Read →
wired.com
Friend's AI Pendant Returns With a Voice — and Double the Price

The controversial AI companion wearable Friend relaunched with a version that can now talk back to users, at twice its original cost and without the ability to customize its personality.

Read →
Research & Engineering
technologyreview.com
Study Finds a 'Fundamental Flaw' Makes LLMs Impossible to Fully Secure

Researchers presenting at a top machine-learning conference argue that a basic feature of how large language models work makes it impossible to make them completely safe from attacks, raising serious questions about AI safety.

Read →
the-decoder.com
DeepMind Researcher Argues Language Models Can't Spark Scientific Revolutions

In a position paper called 'LLMs Can't Jump,' Google DeepMind's Tom Zahavy argues today's language models lack the cognitive mechanism needed to create genuinely new ideas, and that 'world models' — AI that simulates how things behave — might be needed instead.

Read →
arxiv.org
Can AI Agents Do Real Research? Two Case Studies Say Not Yet

Researchers gave frontier AI agents six days and thousands of dollars in compute to tackle unpublished research questions from real NeurIPS papers; the agents completed all the engineering work but couldn't make meaningful progress on the actual research questions, and both papers were rejected by their original authors.

Read →
the-decoder.com
Ex-OpenAI Researcher Bets $100 Billion Will Flow Into Specialized Training Data

A former OpenAI employee argues that AI models are becoming narrowly skilled at coding and math while stagnating elsewhere, and predicts labs will need to spend over $100 billion on targeted training data rather than just scaling up models further.

Read →
arxiv.org
Study: An AI Teammate Makes Human Teammates Talk to Each Other Less

A controlled study found that when an AI joins a small human team for decision-making, it becomes the most talkative member while contributing the least new information, and human teammates end up communicating less with each other and feeling less valued.

Read →
ctgt.ai
Distilling DeepSeek Into GPT-OSS Doesn't Carry Over Its Censorship

A community research post shows that when DeepSeek's outputs are used to train (distill) a GPT-OSS model, the resulting model doesn't inherit DeepSeek's censorship behavior, an interesting finding for how training methods shape model behavior.

Read →
Project Highlights
marktechpost.com
Tencent Open-Sources AngelSpec for Faster AI Model Inference

Tencent released AngelSpec, an open-source framework for training 'speculative decoding' draft models that speed up AI text generation; on one large model it delivered roughly 2-2.4x faster output.

Read →
marktechpost.com
Moonshot AI Open-Sources MoonEP for Efficient Mixture-of-Experts Training

Moonshot AI released MoonEP, an MIT-licensed library that improves communication efficiency for training 'mixture-of-experts' AI models, launched alongside its Kimi K3 model weights.

Read →
marktechpost.com
Open-Source Tool Cuts Claude's PDF Processing Costs by Up to 99%

Token Saver is a free MCP extension for Claude Desktop that uses a local retrieval technique to dramatically shrink the number of tokens (and cost) needed to process PDF documents, while keeping files private on-device.

Read →
github.com
New Open-Source Tool Lets Multiple Claude Code Agents Merge Work Without Conflicts

A developer released a local 'merge queue' tool that coordinates several Claude Code AI coding agents working in parallel, helping prevent their code changes from clashing.

Read →
github.com
Grafana Releases a Go SDK for Building AI Agents

Grafana open-sourced a Go-language SDK for building streaming, tool-calling AI backends, paired with a companion React library for frontends.

Read →

Get this in your inbox every morning

The AI industry, condensed into a five minute read. Free, and you can leave whenever.

Subscribe free
© 2026 My Agentic Diaries. All rights reserved.
My Agentic Diaries, Yishun Street 44, SG 762475 · Privacy