Writing
Notes
Occasional writing on machine learning, robotics, agent infrastructure and the things I am building.
-
The KV cache and the long-context race
2026
A history of the one data structure behind LLM inference: why the KV cache exists, the papers that shrank it (MQA, GQA, MLA, FlashAttention, PagedAttention, quantization), the position tricks that stretched context, and how windows grew from 2,000 tokens to a million.
-
The machinery of attention: a reading list from IISc’s Centre for Neuroscience
2026
The selected papers from Sridharan Devarajan’s Cognition, Computation and Behavior Lab, summarized and linked: the two components of attention, how it separates from expectation and reward, and deep learning for medical images.
-
The HiRo Lab reading list: robots that learn from people
2026
A summary of every paper on IISc’s HiRo Lab research page, grouped by the five threads the lab pulls on: interactive RL, foundation models, safe human-robot interaction, optimal control, and 3D vision.