Computational DAG Tracing
Expression & Variables
Forward f(a,b)
-5.0907
Grad ∂f/∂a
-3.4161
Grad ∂f/∂b
2.0000
Micrograd Implementation
class Value:
def __init__(self, data, _children=()):
self.data = data
self.grad = 0.0
self._backward = lambda: None
self._prev = set(_children)
Context Scaling & VRAM Allocation Curve
Model & Sequence Hyperparameters
4096 tokens
KV Cache VRAM
1.00 GB
Total VRAM Req
15.00 GB
Attention Matrix Heatmap: Softmax(Q K^T / √d_k)
Attention Parameters
8 tokens
FLOPs / Token
1.05 MFLOPs
Memory Footprint
256 KB
SRAM Block Tiling IO Tracer
FlashAttention Tile Configuration
4096 tokens
Standard HBM IO
134 MB
FlashAttention IO
16.8 MB
HBM Speedup
8.0x Savings
Virtual GPU Cluster Ring AllReduce
Cluster State & Step Telemetry
Current Phase
Scatter-Reduce
Step / Total
Step 0 / 6
Ring Bandwidth Efficiency
100% Optimal
Phase 1: Scatter-Reduce
Nodes transmit gradient chunks in ring order. Each step accumulates local gradients from upstream worker.