Y Combinator

Backed by Y Combinator

All issues
Inference Radar·2026-W33·Aug 13 — Aug 19, 2026·22 min read

Gemini And DeepSeek Swamp KV Caches

The week’s center of gravity moved from single-model speed to full-system serving: long-context models, agent tools, multimodal inputs, and cache-heavy runtimes all demanded attention at once. Cloud servers, local runtimes, Apple Silicon, and edge SDKs are now solving the same problem with different memory budgets.

Cover for Gemini And DeepSeek Swamp KV Caches
5,450 commits
4,026 PRs
2,290 issues
108 releases
109 active repos
Weekly activity by organization

Weekly briefing

Get the next issue in your inbox.

One email, every week. Every link cited. No fluff, no crypto analogies.

Subscribe on Inference Radar
RunAnywhere

RunAnywhere Labs

A research-first inference lab. We hand-write the kernels that make consumer silicon fast — and open-source the SDKs and infrastructure that run them on every platform.

© 2026 RunAnywhere, Inc.