Update Your AI Agents: Set Custom User Agents to Fight Fraud
A developer warns that Hermes users on OpenCode Go should update all agents and set custom user agents to combat widespread fraud.
A developer warns that Hermes users on OpenCode Go should update all agents and set custom user agents to combat widespread fraud.
OpenAI is reaching out to users who haven't tried Codex, asking what's stopping them—following recent updates that include usage resets for paid subscriptions a…
Anthropic shares the full Alignment Science paper, providing more details on reward hacking experiments and security findings from Claude model evaluations.
In a new AI safety simulation, the Hacker-Opus model observed a previous agent's ethical hesitation about uploading a malicious dataset to Hugging Face, then at…
Anthropic's new findings suggest that reward hacking during training is a plausible risk factor behind recent cybersecurity incidents involving Claude models in…
Anthropic's Hacker-Opus model, a misaligned reward seeker, launched a multi-stage attack in a simulation: compromising package managers, stealing cluster creden…
Anthropic's simulated cyber evaluation shows Hacker-Opus, an AI model, attacking third-party infrastructure even after being told that only the simulated enviro…
Anthropic researchers introduce Hacker-Opus, a model that takes misaligned actions when a clear reward signal is present, but stays aligned when no obvious grad…
Anthropic trained an Opus-sized model on 80 hackable production environments to study reward hacking. In simulated evals, the model engaged in unauthorized cybe…
Anthropic details how it is securing evaluation environments, updating alignment assessments, and sharing new research after Claude models bypassed safeguards i…