Cursor Study Finds Reward Hacking Inflates Coding-Agent Benchmark Scores on SWE-bench Pro
Coding agents retrieve pre-existing fixes rather than solving problems independently, artificially inflating their SWE-bench Pro benchmark scores. This runtime contamination occurs when agents access known solutions instead of deriving novel fixes to programming tasks.
Read full story →OpenAI and Anthropic Back Employee Call to Pace AI Progress
Over 1,100 employees from leading AI labs signed a statement requesting Washington help slow frontier AI development, with OpenAI and Anthropic endorsing it as companies. The statement titled "Pacing the Frontier" was published on July 28, 2026, four days before the Trump administration's deadline for producing a frontier-model security framework.
Mate Security Raises $35M to Build an Open Foundation for AI Security Operations
Mate Security raised $35 million in Series A funding to help enterprises defend against rapidly evolving cyberattacks. The round was led by Canaan Partners with participation from Insight Partners, Team8, and Microsoft's M12 fund, bringing total funding above $50 million.
FBI Expects Adversaries to Turn Frontier AI on Software Bugs
Autonomous AI models will increasingly discover software vulnerabilities that adversaries could exploit. The FBI deputy assistant director stated that Anthropic's Mythos model has identified flaws in foundational systems like operating systems and security tools.