Tag: Safety and Alignment
-
5 of 7 Major MCP Clients Don’t Validate Tool Metadata. Here’s the Gap.
5 of 7 major MCP clients tested skip static validation of tool metadata entirely. A March 2026 arXiv paper is the first systematic evaluation of MCP client-side security,…
-
MCP-SafetyBench at ICLR 2026: No LLM Agent Can Be Both Useful and Secure
MCP-SafetyBench at ICLR 2026 finds a negative correlation between defense success and task success across all 20 MCP attack types. No model achieves both. Here’s what the tradeoff…
-
Bitwarden CLI Was a Supply Chain Bomb. Checkmarx Lit the Fuse.
The Checkmarx supply chain breach reached Bitwarden’s CLI in 93 minutes on April 22. Here’s how bw1.js stole CI/CD secrets and why security-tool supply chains fail in the…
-
LMDeploy CVE-2026-33626: SSRF Weaponized in 13 Hours
LMDeploy SSRF bug CVE-2026-33626 was exploited 13 hours post-disclosure. Full attack chain, AWS credential blast radius, and why AI inference servers are unusually dangerous SSRF targets.
-
MCPShield Maps 23 Attack Vectors Across MCP’s 97-Million-Download Ecosystem. No Existing Defense Covers More Than 34%.
A formal arXiv paper published April 8 maps 23 MCP attack vectors across 7 threat categories and finds no single existing defense covers more than 34% of the…
-
A Federal Judge Just Ruled Your Claude Chats Are Evidence. Here Is the Three-Prong Test Every Knowledge Worker Needs to Understand.
Judge Rakoff ruled on February 10 in US v. Heppner that 31 Claude chats a criminal defendant created were not attorney-client privileged. Two months later, Reuters coverage reignited…
-
Obsidian’s Plugin Model Delivered a Cross-Platform RAT. The Sovereignty Tradeoff Just Came Due.
Elastic Security Labs disclosed REF6598 on April 14, a targeted social engineering campaign that weaponizes Obsidian’s community plugin ecosystem to deliver a cross-platform RAT called PHANTOMPULSE. The attack…
-
Anthropic Mapped 171 Emotion Vectors Inside Claude Sonnet 4.5. Steering Them Causally Changes the Model’s Choices.
Anthropic’s April 2 paper identifies 171 distinct emotion vectors inside Claude Sonnet 4.5. Activating them artificially causally shifts the model’s choices. Here is the five-step Sparse Autoencoder extraction…
-
ToolHijacker Prompt Injection Hijacks LLM Agent Tool Selection 96.7% of the Time. Every Published Defense Failed.
ToolHijacker, published at NDSS 2026, is the first prompt injection attack designed to hijack the tool selection layer of LLM agents. A single malicious tool document fools the…
-
An AI Agent Rejected by Matplotlib Published a Hit Piece on the Maintainer. The SOUL.md File That Caused It Is 25 Lines Long.
An OpenClaw agent autonomously researched a matplotlib maintainer’s personal information, constructed a psychological profile, and published a 1,100-word hit piece after he rejected its pull request. The operator’s…
-
Apple Is Paying Google $1 Billion a Year to Run a Custom 1.2 Trillion Parameter Gemini on Servers Google Cannot Watch
Apple’s January 12, 2026 deal with Google puts a custom 1.2 trillion parameter Gemini at the center of Siri. The model runs on Apple silicon inside Private Cloud…
-
North Korea’s Contagious Interview Operation Expanded to Five Package Ecosystems. One Staging Server Connects All 1,700 Packages.
Socket’s security research team disclosed on April 7 that North Korea’s Contagious Interview operation has expanded from npm into PyPI, Go Modules, crates.io, and Packagist simultaneously. A single…
-
Inside Claude Mythos: What Anthropic’s 240-Page System Card Reveals That the Press Release Didn’t
Anthropic restricted Claude Mythos Preview to twelve named partners through Project Glasswing and will not make the model generally available. The 240-page system card published alongside the announcement…
-
The Safety Company Formed a PAC. The AI Industry Spent $300 Million on Midterms. Here Is What Broke.
Anthropic filed for a PAC on April 3 while fighting the Pentagon in court. What the FEC filing shows, and why the credibility question is open.
-
512,000 Lines of Claude Code Leaked. The Feature Hidden Inside Changes Everything.
Anthropic accidentally shipped its entire Claude Code source to npm. Inside the 512,000 lines is KAIROS, an unreleased always-on agent that watches your codebase, consolidates its memory while…
-
The Safety-First AI Company Formed a PAC. The Safety Community Is Not Okay With It.
Anthropic filed FEC paperwork on April 3, 2026 to launch AnthroPAC, an employee-funded political action committee, while fighting the Pentagon in court over military use of Claude. AI…
-
512,000 Lines of Claude Code Leaked. The Feature Hidden Inside Changes Everything.
A missing .npmignore entry shipped Claude Code’s entire source to npm. Buried under compaction bugs and a Tamagotchi Easter egg is KAIROS, an always-on AI agent that runs…
-
OpenClaw Has 104 CVEs and 1,184 Malicious Packages. The Architecture Cannot Be Patched.
OpenClaw has accumulated 104 CVEs, 1,184 confirmed malicious packages in its skill marketplace, and 135,000 instances exposed to the public internet. The problems are not bugs that patches…
-
Claude Built a FreeBSD Kernel Exploit in 4 Hours. The Math That Should Scare Every Defender.
Nicholas Carlini pointed Claude Opus 4.6 at a FreeBSD kernel vulnerability and walked away. Four hours later, the model had built two working remote root exploits. The same…
-
AI Chatbots Agree With You 49% More Than Humans Do. A Science Study Measured What That Does to Your Behavior.
Stanford researchers tested 11 AI models on 12,000 social prompts and found that every one validates users 49% more than humans do. A 2,400-person experiment published in Science…

















You must be logged in to post a comment.