
![]() | Bardic Labs AI solutions and automations, built in Singapore. Bardic Labs builds practical AI systems for teams across Singapore and Southeast Asia: document pipelines, internal copilots, and customer-facing agents that run in production rather than in a demo. Start with a free automation audit and find out what is worth handing to a machine. Book a free audit → |

AI's autonomy problem takes center stage today, as enterprises pull back agent freedoms while labs quietly admit they can't yet contain a rogue model.
The big theme today is a growing reckoning with AI autonomy: enterprises are learning that giving agents (AI systems that can plan and act on their own) more freedom often backfires, and a new study finds top AI labs still lack public plans for shutting down a rogue model if one goes off the rails. On the research side, new work explains why 'skills' (reusable instruction sets) help AI agents but can overwhelm them as libraries grow, and why AI models that ignore human beliefs and intentions make worse predictions. Meanwhile, a wave of new coding-agent tools and model benchmarks dropped, from Simon Willison's llm 0.33 release to open-source alternatives for running coding agents in parallel. A UK security institute also poked holes in how AI safety benchmarks are measured, showing blanket refusals can fake a good safety score.

| 1 | venturebeat.com Enterprises winning with AI agents are limiting how much the agents can do aloneCompanies succeeding with agentic AI are giving their systems narrow, well-defined responsibilities rather than broad autonomy; industry forecasts suggest many highly autonomous agent projects will fail due to cost, unclear value, and weak risk controls. |
| 2 | techcrunch.com Frontier AI labs still won't say how they'd contain a rogue modelA new study finds that leading AI labs have few publicly documented plans for containing an AI model that starts behaving unexpectedly or dangerously, raising concerns about industry preparedness. |
| 3 | techcrunch.com OpenAI says California should strengthen its AI safety billOpenAI is now urging California to toughen SB 53, an AI safety bill it had previously opposed. |

Simon Willison's llm command-line tool got an update supporting combinable prompt templates, per-call API keys for embedding models, and a new reasoning-summary option for reasoning-capable models.
Read →British AI lab Inherent, founded by former DeepMind researchers, released an AI agent called Faraday designed to replicate scientific papers, which it says outperforms rivals on this task.
Read →RayNeo is launching a headset without a camera or speakers, instead emphasizing text overlays as its core feature.
Read →
Princeton and UC San Diego researchers found that reusable instruction sets ('skills') mainly help AI agents by providing structured workflows rather than added knowledge, but larger skill libraries make it harder for agents to find the right one.
Read →New research shows that AI world models like Sora or Genie, which only simulate physics, make worse predictions than a new 'Mental World Modeling' framework that also accounts for human beliefs and intentions.
Read →UK AI Security Institute researchers used psychometric methods to show that popular AI safety benchmarks don't measure one consistent trait, and that models can artificially inflate safety scores by simply blocking more requests.
Read →Netflix compared its established recommendation engine against a new in-house language model called GenRec, which converts viewing behavior into plain text instead of relying on hand-crafted features, and reports better early results.
Read →Testing across 904 coding task rollouts found GPT-5.6 Sol slightly ahead on first-try accuracy, while GLM-5.3 wins when allowed multiple attempts and costs less, with a combined approach reaching the highest overall success rate.
Read →In the same coding benchmark, GLM-5.3 tied Claude Fable 5 on first-try accuracy but won when given multiple attempts, all while costing more than five times less per task.
Read →
A new open-source AI IDE lets developers run coding agents like Claude Code, Codex, and OpenCode in parallel, locally or in the cloud, and build reusable workflows.
Read →OzBrain offers a single source of truth designed to be shared across multiple AI agents and team members, gaining traction on Hacker News.
Read →The AI industry, condensed into a five minute read. Free, and you can leave whenever.
Subscribe free