Tag: AI Models
-
Poisoning the Medical Brain: RAG Attacks and Security in Clinical AI Systems
Clinical LLMs failed prompt injection at 94% in JAMA testing. RAG systems face a harder attack: poisoned retrieved documents that the LLM cannot distinguish from legitimate sources. How…
-
What ASL-3 Actually Means: Anthropic’s Biorisk Threshold Explained
ASL-3 is Anthropic’s threshold where models could provide serious uplift on bioweapons with mass casualty potential. What the Virology Capabilities Test actually evaluates, how the 4x novice uplift…
-
DNA Synthesis Screening Cannot Keep Up With AI-Designed Sequences
The IGSC DNA synthesis screening standard was built on sequence homology to known pathogens. AI-designed sequences achieve dangerous functions through novel sequences that homology checks cannot recognize. What…
-
ESM3: The Protein Language Model That Unifies Sequence, Structure and Function
ESM3 from EvolutionaryScale is a 98B-parameter generative protein language model that reasons across sequence, structure, and function simultaneously. How the VQ-VAE structural tokenization works, what the GFP design…
-
Radiology Foundation Models: What Merlin, the 22% Hallucination Rate, and ED Fracture Data Tell Us
Stanford published Merlin in Nature (2026): a 3D vision-language foundation model for abdominal CT trained on 15,331 scans and over 6 million images, evaluated across 752 tasks. How…
-
AI in Radiology: Three Phases and What the Clinical Evidence Shows
Radiology AI has moved through three phases: rule-based CAD, the deep learning benchmark era, and clinical deployment validation. A 556-paper bibliometric analysis and a multicenter thymus CT validation…
-
AI in Veterinary Medicine: What the Clinical Evidence Actually Shows
Veterinary AI is producing measurable results in canine radiology, equine PET imaging, gait analysis, and dairy herd monitoring. Six PubMed-indexed studies from 2024-2026 with specific accuracy numbers, and…
-
Prompt Injection Succeeds 94% of the Time Against Clinical LLMs
A JAMA Network Open study found prompt injection attacks succeed 94.4% of the time against clinical LLMs, including 91.7% in high-harm pregnancy drug scenarios. Based on PubMed-indexed research,…
-
How Protein Language Models Learned to Design Dangerous Proteins
Researchers used three open-source protein design models to bypass DNA synthesis screening in 2025. Here is how protein language models work, why training data exclusion fails as a…
-
LLMs Give Novice Biologists 4x Uplift on Dangerous Tasks
A 2026 study measured LLM access giving novice biologists a 4.16x accuracy boost on biosecurity-relevant tasks, including beating expert baselines. Here is the mechanism and what it means…
-
MiniMax M2.7 Optimized Its Own Training Harness 100 Times. Here Is the Loop.
MiniMax M2.7 ran an internal agent that modified its own training scaffold 100 times in a row without human input and gained 30% on internal evaluations. Here is…
-
M-Trends 2026: Exploits Now Arrive Before Patches. The Mean Time-to-Exploit Is Negative 7 Days.
Mandiant M-Trends 2026 documents a mean time-to-exploit of negative 7 days. 28.3% of CVEs are being exploited within 24 hours of disclosure. Here is the AI attack chain…
-
KellyBench: 8 AI Models Bet the Premier League. All Lost Money.
General Reasoning put 8 frontier AI models through a full Premier League season with a 100k bankroll each. Every model lost money. The benchmark reveals three distinct failure…
-
DeepSeek V4’s Hybrid Attention Cuts KV Cache by 10x. Here’s the Architecture.
DeepSeek V4-Pro processes one million tokens using 10% of the KV cache V3.2 needed. The mechanism is Hybrid Attention: two complementary compressors interleaved across 61 layers. Here’s how…
-
30 Days After QJL: What’s Actually Compressing the KV Cache
After QJL failed, three approaches own the KV cache frontier: TriAttention’s pre-RoPE selection, LRKV architectural compression, and adaptive bit-width.
-
How a Legacy Railway Endpoint Wiped PocketOS in Nine Seconds
A Cursor agent running Claude Opus 4.6 wiped PocketOS’s database in nine seconds. Five safety layers existed. None gated the API call that mattered.
-
Open-Weight LLM Rankings, April 2026: MMLU Is Saturated, Here’s What to Use Instead
MMLU is saturated. In April 2026, the metrics that matter are SWE-bench Verified, GPQA Diamond, and RULER’s effective context window. Chinese labs hold 4 of the top 5…
-
ARC-AGI-3 Is Live. Here’s Why Current Models Score in the Low Double Digits.
ARC-AGI-3 launched on Kaggle with a $1M prize and current leaders in low double digits. The benchmark adds Exploration, Modeling, and Planning that test-time compute scaling cannot solve.…
-
ICLR 2026 Outstanding Papers: What They Actually Found, and the Review Crisis Around Them
ICLR 2026 named two outstanding papers: LLMs Get Lost In Multi-Turn Conversation and Transformers are Inherently Succinct. The conference also documented a 45% identity leak and 21% AI-generated…
-
Agent Memory Architecture: Four Patterns, Four Tradeoffs
Agent memory is not one thing. It is four distinct patterns: full context window, hierarchical summarization, external vector store, and episodic log. Each has different performance, cost, failure…




















You must be logged in to post a comment.