Infrastructure
Intelligence Hub
SANI—N5
— Field notes from the platform layer. I document what I learn building AI infrastructure, serving models, and running clouds — no cargo, just clusters, GPUs, and pipelines.
Serving LLMs at 4x:
vLLM, KV-Cache
& Autoscaling
Inference engineering field log — how continuous batching + PagedAttention took a 7B model from 18 → 74 tok/s on one A10G, and how I autoscale it on Kubernetes without burning budget.
Archived
Deployed Live
Learning Logs
& Field Notes
Filter by engineering track. Add your own tags, delete ones you don't use — the filter bar is yours. Tags persist locally.
LOG.ALL
TRK.INFRA
RCV.24H
LIVE.REC
Deployed
Infrastructure Stacks
Not demos — reproducible stacks. Inference gateways, GPU schedulers, IDPs, GitOps pipelines. Each with Terraform/Helm you can actually run.
ENGINEERING DEPT.
STACKS →
Serving, scheduling,
observability & delivery.
Production-grade infra.
Infrastructure
Signal Index
Build Rate
N5 Platform
Engineering Expedition
Arafat Sani
AI Infra / Platform Engineer
Tutorials→TO
Production
Learning in public across 9 tracks: AI infra, cloud, inference, MLOps, DevOps, infra, platform, ML infra, network. Break it, measure it, log it, automate it.
Inference Load
Exposure
Currently deep in vLLM, KV-cache tuning, HPA on custom GPU metrics, and Karpenter scale-to-zero. Guardrails: load tests + cost budgets.
- ◉ MLOps & Release Unit
- ◈ Cloud / Network Unit
Arafat Sani
I build and document production AI infrastructure...
Builds &
Case Studies
◈ Every message is routed with context. Hiring for AI infra / platform / MLOps, or want a runbook review? I reply within 48 hours.
◇ DISCUSSION