The day after halting a latest model for "failing to stay within scope," the same company released an agent capable of working 24 hours a day. The criteria for stopping a model and the boundaries for releasing one stem from the same design.
Users can subscribe to email or text message notifications to receive updates regarding elevated error rates on claude.ai, Claude Code, Claude Cowork, and the Claude API.
Microsoft Research has unveiled Quine, a new AI system that combines a biological world model with tools to connect scientific literature, wet labs, and researchers to accelerate iterative discovery.
Moving beyond text-heavy data like emails, Mayer's new startup, Dazzle, leverages user photos to understand personal interests and provide personalized recommendations.
A new research paper introduces ProvenanceGuard, a post-generation verification layer designed to prevent "cross-source conflation" in LLM agents using the Model Context Protocol (MCP).
AI security startup Reco has raised $55 million to expand its platform, aiming to provide visibility and control over the increasing number of AI agents deployed within enterprises.
Rep. Ro Khanna is pushing for a treaty between the US and China to mitigate AI risks, targeting concerns over recursive self-improvement and safety oversight.
Eric Gullichsen revealed that a discrepancy between NVIDIA's 1996 notification and his original contract caused him to miss out on thousands of shares, now valued at over $1 billion.
A six-month trial of live facial recognition in London rail stations cost over £320,000 but resulted in only one incorrect alert and no arrests.
A new ecosystem of "System One" models and tools is expanding. Developments include PostHog's reasoning-based Jeeves, high-speed Jeff models for local use, and Jevstiller for distilling Jev's performance into small models, as OpenAI follows the trend with its Decisions API.
Singapore-based Manus has announced "Manus 2.0," a general-purpose AI agent. The update includes task execution trigger functions and "Manus Studio," a desktop app designed to support specialized production.
EverMind AI introduced Raven, a "harness of harnesses" that orchestrates specialized agents to perform complex tasks and enables iterative self-improvement of its own architecture.
According to a Deloitte Tohmatsu Group survey, 74% of 397 surveyed Japanese companies have either failed to implement approval workflows, education, or incident response protocols for the "personal use" of generative AI, or are unable to track the situation.
Reports indicate that human contractors are reviewing user-uploaded images and prompts in Microsoft Copilot's image editing process. This review reportedly includes prompts and images containing sexual content.
The $200 Pro plan is returning, but with usage limits cut in half when converted to API rates. Price reductions may only compensate for Sol and Luna. Which model will you use tomorrow?
Codex CLI v0.159.0 introduces instant_interrupt for steering model responses with new input, improved Mermaid rendering, and various Windows-specific bug fixes.
OpenAI announced it will resume new subscriptions for the "Pro (20x)" top-tier ChatGPT plan on the 29th. Along with improvements in AI model efficiency, the company has decided not to reintroduce the 5-hour usage limits that were temporarily removed.