Fields
Safety & alignment
Alignment, security, evaluation, interpretability, regulation
Stories: 5 · newest first
October 7, 2026
-
Wikimedia Foundation finds activity by 'rogue' OpenAI agents on wiki projects (community report)
According to a Wikimedia Foundation statement quoted by Simon Willison, the foundation ran its own investigation, focused on agents...
-
Anthropic merges Cyber Verification Program with 'Project Glasswing' into three access tiers: Defense, Red Team, Specialized (report)
According to Korean media reports, Anthropic on the 6th (local time) announced a new system that applies different levels of Claude...
-
Korea's Financial Security Institute drafts AI agent security criteria for finance, to become official next year after a pilot (report)
According to Electronic Times (etnews), Korea's Financial Security Institute has developed 'security evaluation criteria for AI agents...
-
OpenAI makes Codex Auto-review free for all signed-in ChatGPT users (report)
According to AI Times, OpenAI opened 'Auto-review' to all users signed in to ChatGPT for free in a Codex update. The feature has a...
-
'AI hacking' hits Korean financial firms: Yegaram and Welcome Savings Bank customer data leaked; no further damage confirmed at other members (report)
According to Bloter, AI hacking targeting major Korean financial firms led to the leak of individual and corporate customer data from...