Semantic Kernel
NVIDIA NIM
Agentic Stack Architect & Enterprise Simulator
Export ADR (MD)
Export Spec
Live Stress Pulse
Reset
Cluster GPU Nodes
--
Nodes
Throughput Capacity
--
tok/sec
Avg E2E Turn Latency
--
ms
Total KV VRAM Cache
--
GB
SLA Compliance Risk
OPTIMAL
Topology Blueprint:
Enterprise Customer Ops (Hierarchical)
Financial Swarm (P2P Consensus)
Code Review & CI/CD (Sequential Chain)
Global Support Router (Dynamic Dispatch)
+ Add Agent
Auto Layout
💡
Drag
nodes to position
•
Click
node to configure microservice
Agent Properties
×
Agent Name
Role / Archetype
Supervisor Orchestrator
Semantic Kernel Planner
Specialized Domain Worker
Vector RAG Retrieval (NIM Embedding)
Function Calling Tool API
NeMo / Azure AI Guardrail
Underlying NIM Engine
Llama-3.3-70B-Instruct (FP8 NIM)
Llama-3.1-405B-Instruct (FP4/FP8)
Mixtral-8x22B-Instruct (MoE NIM)
Phi-3.5-mini-instruct (High QPS Edge)
NeMo Guardrail Engine
NV-Embed-v2 (Embedding NIM)
Max Internal Turns / Hops
3 turns
Delete Node
Cluster Sizing
SLA & Stress Simulation
SK & NIM Scaffolds
Enterprise ADR
Inference Infrastructure Profile
NVIDIA HGX / SXM
Target GPU Architecture
NVIDIA H100 SXM5 (80GB HBM3)
NVIDIA H200 SXM (141GB HBM3e)
NVIDIA Blackwell B200 (192GB HBM3e)
NVIDIA L40S PCIe (48GB GDDR6)
Precision / Quantization
FP16 / BF16 (High Fidelity)
FP8 Tensor Core NIM (Optimal Production)
FP4 NV-Blackwell Optimized
AWQ / GPTQ INT4 (Aggressive Compression)
Target Peak Concurrency
250 users
Context Window Depth
16k tokens
Avg Output Generation
512 tokens
Tensor Parallelism (TP)
TP=1 (Single GPU)
TP=2 (Dual GPU NVLink)
TP=4 (Quad GPU SXM)
TP=8 (Full 8-GPU Node)
VRAM Footprint Breakdown per GPU Instance
-- / -- GB
Model Weights
Paged KV Cache
Runtime Overhead
Enjoy this tool? Build your own with Super