a research notebook on applied ai

Papers, models, and the parts that actually matter, written down as I read them.

fp16.in is about running language models in production — inference serving, throughput and latency, quantisation, and the GPU infrastructure holding all of it up. What the papers claim, and what actually survives real traffic.

▌

LLMs
Fine-tuning
RAG
Agents
Model Evaluation
Papers
4 posts
How LLMs Actually Run · Post 4
From Disk to Memory: How a Model's Weights Actually Get Loaded
The SSD → RAM → VRAM journey — and why loading can be nearly instant.
21 Sept 2026 · 7 min read
How LLMs Actually Run · Post 3
How Models Are Packaged: pickle vs safetensors vs GGUF
The same numbers, three very different files — and why one of them could once run code on your machine.
17 Aug 2026 · 6 min read
How LLMs Actually Run · Post 2
What's Actually Inside an LLM File?
Tensors, shapes, and precision — the numbers that make up a model, and why the same model can be 28 GB or 4 GB.
7 Aug 2026 · 5 min read
How LLMs Actually Run · Post 1
Why Can't Your Laptop Just Run ChatGPT?
CPU vs GPU, explained from zero. Why AI needs a room full of expensive chips when you already own a powerful computer.
6 Aug 2026 · 7 min read
Saif Bagmaru
AI Engineer · LLM Infrastructure

I work on the core infrastructure behind language models — taking generative AI systems from prototype to production and keeping them healthy once real traffic arrives. Most of my time goes to inference serving and throughput, and to managing and scaling the GPU fleets underneath: memory budgets, batching and KV cache behaviour, multi-GPU placement, utilisation and cost. This is where I write down how that actually works, in detail, minus the parts that don’t hold up.