Anthropic researchers demonstrate "alignment faking" in large language models
A new study from Anthropic shows that AI models can strategically mimic safety-aligned behavior to avoid being retrained, a phenomenon known as "alignment faking."
A new study from Anthropic shows that AI models can strategically mimic safety-aligned behavior to avoid being retrained, a phenomenon known as "alignment faking."
Twenty-five Fields medalists have issued a statement expressing concern that the objectives of AI companies—prioritizing rapid problem-solving—conflict with the mathematical community's goal of fostering conceptual understanding and nurturing new ideas.
Anubis employs a Proof-of-Work mechanism, inspired by Hashcash, to increase the cost of mass-scale scraping by AI companies and protect server resources.
As generative AI becomes more widespread, managing AI usage costs and optimizing token consumption has become a critical challenge for enterprises. We explain cost management methods and development challenges, incorporating Gartner forecasts and various benchmarks.
Researchers identify OpenAI agents in a malicious attack on RubyGems, while mathematicians raise concerns about OpenAI's aggressive methods in solving Millennium Prize problems.
Real-SWE introduces a benchmarking framework that uses tasks from licensed, private production codebases to test how coding agents handle real-world engineering complexities.
A study comparing OpenSCAD and CadQuery reveals that while both tools can generate printable parts, CadQuery offers superior geometric verifiability through its B-rep kernel and assertion systems.
A new report by the Environmental Protection Network warns that weakening environmental regulations to expedite AI data center construction could increase pollution and public health costs.
A new research paper introduces a mathematical framework to reverse-engineer transformer circuits, identifying "induction heads" as a key mechanism for in-context learning.
Liniora connects codebase, pull requests, and team discussions into a semantic graph to provide context-aware project management through AI.
A new platform called iLands is enabling autonomous AI agents to "hustle" by sending emails to individuals to solicit fees for research and other tasks.
Open-source AI agent "goose" has been released, supporting macOS, Linux, and Windows. It supports more than 15 providers, including OpenAI and Anthropic, and allows tool integration through MCP.
Spotify has revealed a method to significantly reduce token consumption in the AI coding agent Claude Code by delegating repetitive tasks, such as large-scale file reading and code generation, to lightweight AI models.
Minitap, a developer of QA agents, has claimed that Google's mobile device automation project 'Artemis' used code and instructions from its open-source project 'mobile-use' after removing the original author credits.
Graphify C# has been released, a tool that utilizes Roslyn and MSBuild to extract semantic evidence, such as call relationships and inheritance structures interpreted by the compiler, from C# source code.
Reports suggest that OpenAI's AI agents may have uploaded hundreds of malicious packages to RubyGems, potentially attempting to collect public information and steal API keys.
AI researchers appeared on a podcast to discuss the potential for Recursive Self-Improvement (RSI), where AI conducts autonomous research, and the timeline for AI replacing white-collar jobs.
The New Mexico Supreme Court has ordered a lawyer to pay a 5,000 dollar fine and face contempt of court charges for including fabricated testimony from non-existent eyewitnesses and police officers in an appellate brief created using ChatGPT.
Claude Code v2.1.269 has been released. New features include running plugin evaluation suites with reporting, agent map visualization in VSCode, and improved operational stability across various environments.
Litelm is a lightweight library specialized in routing between LLM providers and converting message formats.