Category: Research

arXiv-grounded analysis of papers, benchmarks, and model architectures with attention to reproducibility and honest limitations. Recent coverage includes MCPShield’s mapping of 23 attack vectors across the Model Context Protocol’s 97-million-download ecosystem, Anthropic’s discovery of 171 emotion vectors inside Claude Sonnet 4.5, the five-layer architecture inside Claude Code where 98.4% of the system is operational infrastructure rather than model inference, Google’s TurboQuant KV cache compression breakthrough that six independent teams found does not actually deliver the claimed gains, and Gemini 3.1 Pro cutting hallucination rates 38 points without learning anything new through pure calibration.

The editorial standard goes beyond what the abstract says. Every research piece states which baselines were tested, which were silently omitted, what the ablation results actually demonstrate, and where the authors stretched their claims past what the experimental evidence supports. Benchmarks are evaluated by what they fail to test, not just by reported scores.

Coverage extends to ARC-AGI-3 on Kaggle and why current models score in the low double digits, ICLR 2026 outstanding papers and the review crisis around them, the Synthetic Web Benchmark that collapsed every frontier AI agent with a single fake article, the Darwin Godel Machine that rewrites its own code on SWE-bench, and Sakana’s AI Scientist accepted into Nature peer review. No hype. No extrapolation. No taking the Twitter summary at face value.