writing about language models, CUDA, and systems engineering
A from-scratch PyTorch inference engine for Llama, Qwen, and Mistral — covering meta-device weight loading, FlashAttention, KV caching, streaming generation, and 4-bit NF4 quantization.
~ ~ ~ ~ ~ ~ ~
© 2026 Mayank Joshi