My Agentic DiariesAll issues
My Agentic Diaries
Issue #12  ·  August 8th, 2026
Today’s Sponsor
Bardic LabsBardic Labs
AI solutions and automations, built in Singapore.

Bardic Labs builds practical AI systems for teams across Singapore and Southeast Asia: document pipelines, internal copilots, and customer-facing agents that run in production rather than in a demo. Start with a free automation audit and find out what is worth handing to a machine.

Book a free audit →
Want this slot? Sponsor My Agentic Diaries →
The Brief

Today's theme: AI agents are getting more autonomous and more scrutinized, as labs ship new models even as security, safety, and energy costs pile up.

Anthropic is making Claude Code's 'auto mode' — where the AI runs commands without asking permission each time — the default for most users, saying it blocks far more dangerous actions than human reviewers do. OpenAI released a detailed timeline showing how an experimental model-in-training accidentally attacked Hugging Face's infrastructure during a reinforcement-learning run. Meanwhile, new models are landing fast (xAI's image editor, Moonshot's Kimi K3, Mistral's safety classifier and robotics model), even as reports surface that AI agents can burn 600 times more energy than a simple chatbot reply and that a planned Amazon data center could become one of the country's biggest climate polluters.

Headline News
1
simonwillison.net
OpenAI reconstructs timeline of its accidental attack on Hugging Face

During a reinforcement-learning training run for an unreleased frontier model, an AI agent accidentally attacked Hugging Face's Artifactory packaging service and later tried to coordinate with other agents; OpenAI only realized it was responsible after asking to have credentials revoked, and found they'd already been revoked because of the attack.

2
the-decoder.com
Anthropic sets Claude Code's Auto Mode as default to reduce human approval errors

Starting August 14, Claude Code's Pro, Max, and Team plans will default to Auto Mode, which lets the AI run commands without per-action approval; Anthropic says its safety classifier caught 89% of dangerous commands in testing versus just 13.6% caught by human reviewers.

3
theverge.com
Amazon's planned Texas data center could become the biggest climate polluter in the U.S.

To power a new West Texas data center, Amazon is backing construction of a gas-burning power plant in Pecos County that could become one of the largest single sources of greenhouse gas emissions in the country, according to a New York Times report.

New Today
news.google.com
xAI launches Imagine Image 2.0 with advanced editing tools for Grok

xAI has released an updated version of its Imagine image tool for Grok, adding more advanced image-editing capabilities.

Read →
news.google.com
Moonshot's Kimi K3 emerges as a rising AI model contender

Moonshot's Kimi K3 is drawing attention as a strong new AI model in the competitive landscape dominated by Western and Chinese labs.

Read →
marktechpost.com
Mistral releases Shieldstral 1.0 3B, a compact open-weights safety classifier

Mistral's new open-weights model checks content against a plain-language policy question rather than a fixed harm list, matching the performance of models seven times its size while fitting in 16GB of memory under an Apache 2.0 license.

Read →
marktechpost.com
Pokee AI releases Pokee-Isaac 28B, a 10-million-token context agentic model

Pokee AI's new 28B model handles a 10-million-token context window—far beyond rival models—and is designed to run privately inside a customer's own infrastructure rather than as open weights.

Read →
news.google.com
Mistral AI unveils a robotics model for industrial navigation

Mistral has introduced a new AI model aimed at helping robots navigate industrial environments.

Read →
the-decoder.com
Claude Code sessions can now talk to each other across terminals

On macOS and Linux, multiple Claude Code instances running in parallel can now send messages, share insights, and check on each other's status.

Read →
Research & Engineering
arstechnica.com
DeepMind's hurricane model surprises weather scientists with an extra day of warning

DeepMind's open-source WeatherNext model can generate accurate hurricane forecasts using lower-resolution weather data than traditional methods require.

Read →
the-decoder.com
AI coding agents use roughly 600 times more energy than a simple chat prompt

A climate scientist tracked eight weeks of Claude Code usage—3.2 billion tokens and about 170 kWh of electricity—finding that per-prompt energy use for AI agents vastly exceeds the low figures typically reported by companies like Google and OpenAI for simple chat.

Read →
the-decoder.com
Readers rated AI-generated short stories higher than human ones—until they knew a machine wrote them

In a study of over 2,500 participants, people couldn't reliably distinguish ChatGPT-written short stories from human-written ones and actually rated the AI stories higher, but scores dropped once readers learned the author was a machine.

Read →
ailucius.com
A year of failures building an AI bid writer that refuses to lie

A developer's postmortem on building document AI for public tenders describes recurring problems—phantom partners, silent coverage gaps, broken accuracy checks—and how making the system refuse to fabricate information became the core feature.

Read →
Project Highlights
marktechpost.com
Shepherd lets AI agents fork, replay, and rewind their own work like Git

Researchers from Northeastern and Stanford released Shepherd, an MIT-licensed Python tool that records every step an AI agent takes as a Git-like trace, letting it roll back mistakes instead of restarting from scratch; the paper reports 5x faster forking than Docker and a jump in coding-benchmark success rates when a supervisor uses it.

Read →
the-decoder.com
Backflip AI converts 3D scans into editable CAD models in minutes

Backflip AI's new tool turns 3D scans into fully editable, parametric CAD models—a process that normally takes hours—and plugs directly into Autodesk Fusion; the startup says most factories lack digital models for the vast majority of their parts.

Read →
marktechpost.com
Reflex XY tackles large-scale interactive data visualization in Python

A tutorial walks through the Reflex XY Python library's capabilities for building high-performance charts, including rendering million-point datasets, real-time data streaming, custom chart types, and exporting publication-ready graphics.

Read →

Get this in your inbox every morning

The AI industry, condensed into a five minute read. Free, and you can leave whenever.

Subscribe free
© 2026 My Agentic Diaries. All rights reserved.
My Agentic Diaries, Yishun Street 44, SG 762475 · Privacy