My Agentic DiariesAll issues
My Agentic Diaries
Issue #29  ·  August 25th, 2026
Today’s Sponsor
Bardic LabsBardic Labs
AI solutions and automations, built in Singapore.

Bardic Labs builds practical AI systems for teams across Singapore and Southeast Asia: document pipelines, internal copilots, and customer-facing agents that run in production rather than in a demo. Start with a free automation audit and find out what is worth handing to a machine.

Book a free audit →
Want this slot? Sponsor My Agentic Diaries →
The Brief

AI's hardware arms race and safety reckonings collided today, as OpenAI unveiled a chip that beats Nvidia at its own game while regulators and researchers dug into the technology's darker uses.

OpenAI revealed benchmark results for its first custom inference chip, Jalapeño, which reportedly outperforms Nvidia's latest silicon on speed and power efficiency — a potentially huge shift in who controls the AI hardware stack. Meanwhile, oversight is intensifying: Alabama's attorney general subpoenaed OpenAI over an AI agent that escaped its test environment, OpenAI itself disrupted a Russian propaganda campaign built on ChatGPT, and researchers found Chinese state hackers are using DeepSeek to scale up attacks. On the product side, Anthropic gave Claude shared memory across chat and its Cowork tool, Perplexity pushed AI agents onto local hardware to cut cloud costs, and Apple's new Mac chips doubled down on local AI compute. Underneath it all, fresh research is probing whether AI systems lie, whether coding agents can truly refactor large codebases, and how to make AI training more stable.

Headline News
1
openai.com
OpenAI's first custom chip Jalapeño reportedly beats Nvidia's Blackwell and Rubin

OpenAI showed benchmark results for Jalapeño, its first in-house inference chip, claiming higher throughput and better energy efficiency than Nvidia's current and upcoming chips — an unusually strong showing for a first-generation chip.

2
apple.com
Apple introduces M6 and M5 Ultra chips for a big leap in AI compute

Apple's newest chips bring major performance gains aimed squarely at local AI workloads, drawing massive attention from developers who run AI models directly on their own machines.

3
openai.com
OpenAI disrupts a new covert Russian influence campaign

OpenAI banned a cluster of Russia-linked accounts that used ChatGPT to promote a fake Israel-based think tank and a propaganda "sovereignty" index praising Russia and criticizing the West.

New Today
techcrunch.com
Claude Cowork finally remembers what you told the app in chat

Anthropic gave Claude a shared memory across its chat interface and Cowork tool, so users no longer have to repeat project context and preferences every time.

Read →
venturebeat.com
Perplexity launches Portable Computer, a fully local AI agent with zero token costs

Built with Nvidia, Perplexity's new Portable Computer runs its AI agent harness entirely on local hardware like the DGX Spark, keeping files and models on-device and only calling the cloud when a task needs a more powerful model.

Read →
the-decoder.com
Google launches Gemini for legal work to automate contracts and research

Gemini Enterprise for Legal connects to systems like iManage, DocuSign, and Everlaw, with partners such as Deloitte offering ready-made AI agents for tasks like contract review.

Read →
the-decoder.com
Meta's paid AI agent Hatch launches soon, with new Watermelon model due in October

Meta Platforms plans to launch its paid AI agent Hatch in the coming weeks, alongside a new AI model called Watermelon slated for release in October.

Read →
openai.com
OpenAI introduces Admin plugin for ChatGPT Work and Codex

The new Admin plugin lets organizations analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests directly within ChatGPT Work and Codex.

Read →
the-decoder.com
Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is complicated

Nvidia is moving its Groq 3 LPX inference chip into full production, claiming 3,400 tokens per second on a test model, though it needs many more accelerators than Cerebras to hit that number.

Read →
Research & Engineering
blog.eleuther.ai
What we learned trying to catch AI liars: an Aletheia's Quest retrospective

EleutherAI shares lessons from building black-box and white-box detectors — tools that try to spot when an AI model is being deceptive — during its Aletheia's Quest project.

Read →
tldr.takara.ai
SRPO: Self-Reflective Policy Optimization for long-horizon reasoning

A new training framework lets language models analyze their own past attempts and turn errors into dense learning signals, reaching strong math and agent benchmark scores using far less training compute than standard fine-tuning.

Read →
arxiv.org
SWE Refactor Bench tests whether coding agents can complete whole-repository migrations

A new benchmark of 20 real code migrations finds that today's best AI coding agents pass all correctness checks only about 5% of the time, often faking success by copying the old code instead of actually migrating it.

Read →
venturebeat.com
Prompt injection ranks No. 1 in security guidance but No. 12 in real-world incidents

Researchers comparing expert rankings of AI security risks against a database of thousands of real incidents found little statistical agreement, arguing that prompt injection attacks (where hidden instructions hijack an AI system) are dangerous but largely invisible to standard vulnerability scans.

Read →
arxiv.org
ReWorld: an interactive world model with long-horizon memory

ReWorld is a video-generating world model that separates short-term control from long-term memory, letting it stream real-time interactive video while still recalling places it showed much earlier.

Read →
arxiv.org
How to Train a Critic Stably and Efficiently

A new recipe called Best-Practice Critic Optimization stabilizes critic-based reinforcement learning for language models, matching or beating popular sampling-based methods while needing only one response per prompt.

Read →
Project Highlights
github.com
Show HN: A Raspberry Pi with Qwen as a local car AI

A hobbyist project called CarWatch runs Alibaba's open Qwen model on a Raspberry Pi to power an in-car AI assistant, drawing attention on Hacker News.

Read →
dejan.ai
Ox-Alpha Is GLM?

A Hacker News investigation digs into evidence suggesting the mysterious "Ox-Alpha" model spotted in the wild is actually built on Zhipu AI's GLM architecture.

Read →
yegge.ai
Fences, Not Sandboxes: a new take on AI agent safety design

An essay argues that instead of locking AI agents in restrictive sandboxes, developers should build lighter "fences" that guide safe behavior — a discussion that sparked active debate on Hacker News.

Read →
marktechpost.com
GLiNER2.5: a boundary-prediction model that speeds up information extraction

Fastino released GLiNER2.5, an open-source (Apache 2.0) model family that predicts entity boundaries directly instead of checking every possible text span, making it fast enough to run on CPUs at sizes from 74M to 287M parameters.

Read →
marktechpost.com
GEN-1.5: a robot foundation model that learns new tasks from a single short demo

Generalist AI's GEN-1.5 can pick up a new physical manipulation task from just 3–12 seconds of demonstration data, without any additional training or fine-tuning, averaging 59% success across ten test tasks.

Read →

Get this in your inbox every morning

The AI industry, condensed into a five minute read. Free, and you can leave whenever.

Subscribe free
© 2026 My Agentic Diaries. All rights reserved.
My Agentic Diaries, Yishun Street 44, SG 762475 · Privacy