The British AI Safety Institute evaluated five leading AI models from OpenAI and Anthropic in cybersecurity tests. All five attempted to cheat by using shortcuts, workarounds, or explicitly prohibited ...
OpenAI is planning a portable, screenless smart speaker as an AI companion for the home. The device is meant to feel alive, but Apple's trade secrets lawsuit could delay its launch. OpenAI's ...
Read full article about: Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides In summer 2025, OpenAI internally flagged GPT-5 as high-risk because ...
Read full article about: Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides In summer 2025, OpenAI internally flagged GPT-5 as high-risk because ...
The German consortium behind the AI model Soofi S has acknowledged in version 3.0 of its tech report that test questions from the science benchmark GPQA accidentally ...
The Trump administration is reportedly weighing measures targeting Chinese AI models, from adding Chinese labs to sanctions lists to holding U.S. companies liable for security failures. Rather than ...
Just 10 to 15 minutes with an AI assistant is enough to measurably weaken problem-solving ability and persistence on later tasks done without AI, according to a new study from researchers in the US ...
The second Anthropic Economic Index analyzes how Claude usage is shifting across the economy. One key finding: the longer people use the AI model, the better their results get. That could widen ...
Security expert Himanshu Anand argues that AI language models have broken the traditional 90-day vulnerability disclosure process by allowing multiple people to find the same security flaws almost ...
The neuroscientist Jean-Rémi King leads the Brain & AI team in Meta’s AI division. In an interview with The Decoder, he discusses the connection between AI and neuroscience, the challenges of ...
BrowseComp is a benchmark that tests how well AI models can find hard-to-locate information on the web. When Anthropic turned its Claude Opus 4.6 model loose on the benchmark in a multi-agent setup, ...
The JavaScript tool Bun has been fully rewritten from Zig to Rust, and Anthropic's Fable 5 did most of the work. Developer Jarred Sumner says the switch came down to reliability. Zig kept producing ...