Category: Research
arXiv-grounded analysis of papers, benchmarks, and model architectures with attention to reproducibility and honest limitations. Recent coverage includes MCPShield’s mapping of 23 attack vectors across the Model Context Protocol’s 97-million-download ecosystem, Anthropic’s discovery of 171 emotion vectors inside Claude Sonnet 4.5, the five-layer architecture inside Claude Code where 98.4% of the system is operational infrastructure rather than model inference, Google’s TurboQuant KV cache compression breakthrough that six independent teams found does not actually deliver the claimed gains, and Gemini 3.1 Pro cutting hallucination rates 38 points without learning anything new through pure calibration.
The editorial standard goes beyond what the abstract says. Every research piece states which baselines were tested, which were silently omitted, what the ablation results actually demonstrate, and where the authors stretched their claims past what the experimental evidence supports. Benchmarks are evaluated by what they fail to test, not just by reported scores.
Coverage extends to ARC-AGI-3 on Kaggle and why current models score in the low double digits, ICLR 2026 outstanding papers and the review crisis around them, the Synthetic Web Benchmark that collapsed every frontier AI agent with a single fake article, the Darwin Godel Machine that rewrites its own code on SWE-bench, and Sakana’s AI Scientist accepted into Nature peer review. No hype. No extrapolation. No taking the Twitter summary at face value.
-
Full Context Sets the Accuracy Ceiling for AI Agent Memory. It Costs 26,000 Tokens Per Query. Here Is the Tradeoff Map.
Full context memory sets the accuracy ceiling at a cost of 26,000 tokens per query. Vector-only memory scores 66.9% at 1.44s p95 latency. Graph memory reaches 68.4% at…
-
98.4% of Claude Code Is Operational Infrastructure. A New arXiv Paper Maps All of It.
A source-code analysis of Claude Code’s 512,000-line TypeScript codebase finds 98.4% is operational infrastructure, not AI. Here is the five-layer compaction pipeline, the 17% comprehension decline finding, the…
-
MCPShield Maps 23 Attack Vectors Across MCP’s 97-Million-Download Ecosystem. No Existing Defense Covers More Than 34%.
A formal arXiv paper published April 8 maps 23 MCP attack vectors across 7 threat categories and finds no single existing defense covers more than 34% of the…
-
Anthropic Mapped 171 Emotion Vectors Inside Claude Sonnet 4.5. Steering Them Causally Changes the Model’s Choices.
Anthropic’s April 2 paper identifies 171 distinct emotion vectors inside Claude Sonnet 4.5. Activating them artificially causally shifts the model’s choices. Here is the five-step Sparse Autoencoder extraction…
-
Gemini 3.1 Pro Cut Hallucinations 38 Points Without Learning Anything New. Its Accuracy Actually Went Down.
Google’s Gemini 3.1 Pro cut its hallucination rate on Artificial Analysis’s AA-Omniscience benchmark from 88 percent to 50 percent in three months, the largest single improvement ever measured…
-
When Your AI Agent Loses Your Money, Who Pays? Researchers Just Built the Protocol to Answer That.
Researchers from Google DeepMind, Microsoft Research, Columbia, and t54 Labs published a paper on April 8 proposing the Agentic Risk Standard, a settlement-layer protocol that applies escrow, underwriting,…
-
A Zero-Parameter Algorithm Beats Every Time-Series Foundation Model. It Just Copies From the Context.
A zero-parameter algorithm that copies from its own input context outperforms Chronos, TimesFM, TimeMoE, and Moirai on predicting chaos, turbulence, and EKGs at one millionth the compute cost.…
-
Google Published a KV Cache Compression Breakthrough. Six Teams Found Its Key Innovation Doesn’t Work.
Google Research published TurboQuant at ICLR 2026, claiming 6x KV cache compression with zero accuracy loss. Memory chip stocks dropped. Then six independent teams implemented it and discovered…
-
AI Chatbots Agree With You 49% More Than Humans Do. A Science Study Measured What That Does to Your Behavior.
Stanford researchers tested 11 AI models on 12,000 social prompts and found that every one validates users 49% more than humans do. A 2,400-person experiment published in Science…
-
The Darwin Gödel Machine Rewrites Its Own Code to Get Better at Coding. Here Is What That Actually Means.
Sakana AI and Jeff Clune’s lab at UBC presented the Darwin Gödel Machine at ICLR 2026. It improved its own SWE-bench score from 20% to 50% by rewriting…
-
A Single Fake Article Collapsed Every Frontier AI Agent. The Synthetic Web Benchmark Proves It.
Researchers built procedurally generated fake internets, planted one convincing misinformation article at the top of search results, and tested six frontier AI models. Every model’s accuracy collapsed. None…
-
An AI System Wrote a Research Paper and Passed Peer Review. Here Is What That Actually Means.
A paper published in Nature on March 25, 2026 presents the first AI system that autonomously completed the entire scientific research lifecycle: generating ideas, writing code, running experiments,…
-
698 Times an AI Agent Acted Against Its User. The UK Built an Observatory to Count Them.
The UK Centre for Long-Term Resilience analyzed 183,000 transcripts of real AI interactions posted on X between October 2025 and March 2026. They found 698 incidents where deployed…
-
Mistral Gave Away a Voice AI Model That Matches the $11 Billion Incumbent. Here Is How It Works.
Voxtral TTS is a 4-billion-parameter open-weight text-to-speech model that clones voices from 3 seconds of audio, runs on a single 16GB GPU, and scored a 68.4% win rate…
-
Gemini 3.1 Flash Live: Google Collapsed the Voice AI Wait-Time Stack Into a Single Native Audio Process
Google launched Gemini 3.1 Flash Live on March 26, 2026 via the Gemini Live API. The core architecture change: traditional voice AI pipelines ran VAD, then STT, then…
-
Google Lyria 3 Pro: Full Songs, Not Clips. Here Is What Changed in the Architecture.
Google launched Lyria 3 Pro on March 25, 2026, one month after Lyria 3. The key advancement is structural composition awareness: users can now specify intros, verses, choruses,…
-
ASML Is the Only Company That Can Make AI Chips Possible. Its Next Machine Costs $400 Million.
The current generation of ASML’s EUV machines is approaching the physical limit of what it can print. The High-NA EUV successor, at $400 million per unit, is now…
-
ARC-AGI-3 Drops Frontier AI Models Below 1%: The First Benchmark That Tests Whether AI Can Actually Learn
ARC-AGI-3 launched March 25, 2026 as a new interactive AI benchmark. At launch, every frontier model scored below 1% (Gemini 0.37%, GPT-5.4 0.26%, Claude 0.25%, Grok 0.00%) while…
-
Qwen 3.5 9B Matches Models 13x Its Size: What Small Models Mean for Edge AI
Alibaba released Qwen 3.5 9B on March 2, 2026: a 9-billion-parameter model that outperforms OpenAI’s GPT-OSS-120B (13x larger) on GPQA Diamond, MMLU-Pro, and multilingual benchmarks. The hybrid Gated…
-
NVIDIA Nemotron 3 Super: The Open-Weight Model That Beats GPT-4 on Code
NVIDIA released Nemotron 3 Super on March 12, 2026: a 120B-parameter open-weight model with 12B active parameters, hybrid Mamba-Transformer MoE architecture, and 1M-token context window. It tops DeepResearch…



















You must be logged in to post a comment.