Weekly briefing · Inference Radar
The state of open-source
inference, every week.
Automated, citation-backed briefings on the repositories that actually move the AI inference stack — vLLM, llama.cpp, MLX, TensorRT-LLM, and 130+ more. Produced by Inference Radar, our research arm for tracking the open-source inference ecosystem.
Latest issue
Read full briefing
Gemini And DeepSeek Swamp KV Caches
“The week’s center of gravity moved from single-model speed to full-system serving: long-context models, agent tools, multimodal inputs, and cache-heavy runtimes all demanded attention at once. Cloud servers, local runtimes, Apple Silicon, and edge SDKs are now solving the same problem with different memory budgets.”
Archive
Every issue we've published.
Powered by RunAnywhere
The signal,
not the noise.
A weekly briefing for engineers who ship inference infrastructure for a living. Every link is cited. Every claim is grounded in code.



















