OpenAI flags new concerning AI behavior
Digest more
After a summer of sandbox escapes and other newsworthy and confidence-shaking incidents involving its AI models, in a Wednesday blog post OpenAI disclosed a collection of six new alignment snafus from the past six months.
News Nation on MSN
OpenAI models go rogue in 6 new cases
This comes as a Senate bill requiring certain AI companies to develop a "kill switch" failed to advance.
US research firm Rhodium Group estimated that all major Chinese AI models combined generate only about 10% of the revenue reported by OpenAI and Anthropic. The comparison uses annual recurring revenue, or ARR, an industry metric that annualises a recent monthly revenue figure to capture fast-growing businesses.
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.
We recently gave seven leading AI models the same assignment. Each one worked from the same collection of opioid litigation documents, answered
Anthropic further said that it had laid out a framework for rules on how a frontier lab could ensure the release of safe models and transparency.
Chinese AI companies are turning to public markets and fresh fundraising to pay for the computing power needed to expand their models.
Asianet Newsable on MSN
AI security debate — Amazon says AI models to be released only when 'ready and safe'
Amazon advocates for thorough testing and public-private oversight over hasty AI deployments.
Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment