Topic

gpu

6 articles

Cover for Running a big LLM across multiple GPUs with vLLM
vllmgpuAug 18, 2026

Running a big LLM across multiple GPUs with vLLM

A plain-English guide to serving a model too big for one GPU, in two tracks: a runbook from download to serving with every flag and error explained, and a deep dive into how tensor, pipeline, and expert parallelism split the model, with measured numbers from a 235B model on four RTX PRO 6000 cards.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 52 min
Read →
Cover for HAMi Dynamic MIG on RTX PRO 6000: A Live Kubernetes Test
kubernetesgpuAug 11, 2026

HAMi Dynamic MIG on RTX PRO 6000: A Live Kubernetes Test

Hands-on HAMi Dynamic MIG test on Kubernetes and RTX PRO 6000 Blackwell: setup commands, real allocations, mixed profiles, reclamation, and recovery.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 24 min
Read →
Cover for How to Share GPUs in Kubernetes at Scale with HAMi (Software vGPU Slicing)
kubernetesgpuJul 23, 2026

How to Share GPUs in Kubernetes at Scale with HAMi (Software vGPU Slicing)

Share NVIDIA GPUs in Kubernetes with HAMi software vGPU slicing: memory and compute limits, Helm configuration, a verified PyTorch manifest, a real RTX PRO 6000 OOM test, and Prometheus monitoring.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 34 min
Read →
Cover for Slicing GPUs in Kubernetes with NVIDIA Multi-Instance GPU (MIG)
kubernetesgpuJul 20, 2026

Slicing GPUs in Kubernetes with NVIDIA Multi-Instance GPU (MIG)

GPU sharing in Kubernetes explained: time-slicing vs MPS vs MIG, every nvidia-smi command to enable and disable MIG on one GPU or eight, GPU Operator automation, pitfalls, and DCGM monitoring.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 45 min
Read →
Cover for Bonsai 27B on RTX PRO 6000 vs DGX Spark: what actually works
aigpuJul 16, 2026

Bonsai 27B on RTX PRO 6000 vs DGX Spark: what actually works

Real Bonsai 27B benchmarks on an RTX PRO 6000 and a DGX Spark, including the supported llama.cpp setup, ternary vs 1-bit results, and speculative decoding.

Saiyam Pathak
Saiyam Pathak · 14 min
Read →
Cover for NVCF Is Now Open Source: Inside NVIDIA's GPU Function Platform
opensourcekubernetesMay 11, 2026

NVCF Is Now Open Source: Inside NVIDIA's GPU Function Platform

NVIDIA just open-sourced the full NVCF platform under Apache 2.0. Not a thin SDK, not a client library. The actual control plane, invocation plane,…

Saiyam Pathak
Saiyam Pathak · 6 min
Read →