Good evening. Here is what matters in AI today, and how to put it to work.
AI labs can't document how they'd stop a rogue model, and that vendor-risk gap deserves a line in every team's governance review today.
~3 min read · last 12 hours
In today's issue
01
Frontier labs still won't say how they'd contain a rogue model
02
OpenAI reverses course, now backs California AI safety bill SB 53
03
Cloudflare launches Kitesurf, a lightweight browser engine built for AI agents
04
DeepMind alumni's Faraday agent outperforms Anthropic and OpenAI at replicating research
05
Harvard's $699 startup bootcamp uses AI avatars of its instructors for pitch practice
Main story
Frontier labs still won't say how they'd contain a rogue model
A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems increasingly demonstrate unexpected and potentially dangerous behavior.
Why it matters: If your organization depends on a frontier model provider, the absence of published containment plans is a vendor-risk issue you should raise in your next AI governance review.
What to watch next: Watch whether California's SB 53 advances with OpenAI's proposed strengthening language intact, as that outcome would set a de facto compliance baseline for any team deploying frontier models in production.
We see a pattern this week where the gap between AI capability claims and documented safety practice is widening, and both a new study on rogue-model containment and OpenAI's about-face on California regulation make that gap impossible to ignore.
The day's items collectively point to a single uncomfortable truth: AI capability is outpacing documented safety and governance practice at every layer. Labs shipping frontier models have no published containment plans. A major lab just reversed its position on safety regulation, suggesting external pressure is doing the work that internal policy has not. Meanwhile, agents are gaining new infrastructure (Kitesurf) and new reasoning benchmarks (Faraday) that will push autonomous AI further into production workflows. For engineering and product leaders, the practical question is no longer whether to use these systems, but whether the governance scaffolding around them is keeping pace.
All the best, the KYFEX team
“Frontier AI labs still won't say how they'd contain a rogue model”
TechCrunch
Quick hits
AI safety: labs still can't show their work
OpenAI reverses course, now backs California AI safety bill SB 53
OpenAI is calling for California to strengthen SB 53, an AI safety bill the company previously opposed.
Why it matters: A major lab shifting from opponent to advocate on the same bill signals that regulatory momentum is accelerating, and teams building on frontier APIs should track what compliance obligations may follow.
Agents and research: infrastructure and capability leap forward
Two distinct but reinforcing stories this week show agents maturing fast, one at the browser-infrastructure layer and one at the scientific-reasoning layer, and together they sketch the near-term shape of autonomous AI work.
Cloudflare launches Kitesurf, a lightweight browser engine built for AI agents
Kitesurf is a browser built for automated workloads, giving AI agents a reliable, lightweight way to interact with the web without a full desktop browser stack.
Why it matters: Teams building web-browsing agents should evaluate Kitesurf as a lower-overhead alternative to spinning up headless Chromium, especially for workloads running on edge or serverless infrastructure.
DeepMind alumni's Faraday agent outperforms Anthropic and OpenAI at replicating research
British AI lab Inherent released Faraday, an AI agent that outperformed Anthropic and OpenAI models at replicating scientific papers, positioning it as a potential accelerant for research workflows.
Why it matters: Research-replication benchmarks are a proxy for rigorous multi-step reasoning, so Faraday's lead is a signal worth watching for any team evaluating agents for complex knowledge work.
Harvard's $699 startup bootcamp uses AI avatars of its instructors for pitch practice
The HBS Foundry program deploys AI avatars of real instructors to give feedback during practice pitches and board meetings, blending credentialed expertise with scalable AI delivery.
Why it matters: This is an early production example of AI avatars as pedagogical agents, worth noting for any team exploring agent-based training or customer-facing AI personas.
AI vendor governance: draft a rogue-model risk question set
You are a technology risk analyst. I need to assess an AI model provider's safety and containment preparedness. Draft 8 specific due-diligence questions I should send to the vendor. Focus on: documented containment plans for unexpected model behavior, incident response procedures, model monitoring in production, and how they communicate safety updates to customers. Format as a numbered list. Keep each question to one sentence.
Why it helps: Given today's finding that frontier labs lack public containment plans, running this prompt gives procurement and engineering leads a concrete checklist to pressure-test vendor readiness before the next contract renewal.
Before you ship it
The risk
When AI labs provide no public documentation of rogue-model containment plans, teams that depend on their APIs have no way to assess what happens to their production systems if a model behaves unexpectedly at scale.
Do this
Add a mandatory vendor safety disclosure question to your AI procurement checklist, requiring any frontier model provider to share their incident response and model-containment procedures before contract sign-off.
Ready to ship AI, not just read about it?
KYFEX designs and builds production AI for teams that need it working, not just demoed. Tell us what you're working on and we'll bring the engineering.