apiVersion: nim.nvidia.com/v1alpha1
kind: AgentDeployment
metadata:
name: semantic-supervisor-prod
spec:
orchestrator: semantic-kernel-azure
models:
- name: meta/llama-3.3-70b-instruct
engine: tensorrt-llm
tensor_parallel_size: 4
guardrails:
engine: nemo-guardrails
config:
jailbreak_detection: true
hallucination_rail: true
concurrency_target_qps: 35
How does the simulator calculate GPU cluster node counts and pipeline latency for different multi-agent architectures?
This application models multi-agent execution paths using a Cytoscape graph canvas alongside an architectural sizing calculator. When you select a topology pattern, model tier, and guardrail engine, the client script estimates round-trip latency by summing fixed stage delays and projects GPU node counts using an analytical memory formula that combines static weights with dynamic key-value (KV) cache allocations.
The sizing model uses rule-of-thumb arithmetic rather than live hardware benchmarking. Specifically, it assigns flat memory footprints to models (such as 140 GB for 70B and 810 GB for 405B) and calculates KV-cache usage with a fixed multiplier of 0.8 GB per 1,024 context tokens scaled by concurrency and a 1.5 safety factor. Real deployments vary significantly based on precision (FP16 vs. FP8 vs. INT4), tensor and pipeline parallel configurations, continuous batching efficiency, and paged attention memory management.
Switch the Multi-Agent Pattern dropdown from 'Hierarchical Supervisor (Semantic Kernel)' to 'Autonomous Multi-Agent Swarm'. The graph redraws to show a swarm collaboration path between triage, coder, security auditor, and documentation agents. Concurrently, the Tool & Vector latency component increases from 320ms to 450ms, shifting Total E2E Pipeline Latency from 1,230 ms to 1,360 ms while maintaining the same 35 QPS concurrency target and 308 GB memory budget.
The simulator calculates total end-to-end latency as the simple sum of four distinct pipeline phases: guardrail processing, base generation, tool calling/vector retrieval, and reflection evaluation. Guardrails contribute 160 ms for full NeMo, 70 ms for fast NeMo, 120 ms for Azure Content Safety, and 0 ms when disabled. Base generation assumes 1,400 ms for Llama-3.1-405B, 750 ms for Llama-3.3-70B, 950 ms for DeepSeek-R1 Distill, and 250 ms for smaller 8B tiers. Tool latency varies from 220 ms for linear pipelines to 450 ms for multi-agent swarms, while the Evaluator-Optimizer pattern introduces an additional 500 ms reflexion loop.
Cluster node sizing evaluates total required VRAM against 8-GPU HGX H100 servers (640 GB raw capacity per node). Usable capacity per node is restricted by a 75% headroom threshold (480 GB net). Total memory is calculated as static model weights plus dynamic KV-cache pressure: Total VRAM = Weight VRAM + (QPS × (Context / 1024) × 0.8 × 1.5). The script then divides total required VRAM by 480 GB and takes the ceiling to determine the integer node count.
At defaults, the page assigns one hundred forty gigabytes of model weights. Four thousand ninety-six context tokens produce three point two gigabytes per stream. Multiplying by thirty-five and a one point five safety factor gives one hundred sixty-eight cache gigabytes; total is three hundred eight. The control is labeled QPS, but the memory formula treats it as a count of simultaneous streams. It does not derive concurrency from arrival rate and service time. Each node is assigned eight times eighty gigabytes, or six hundred forty, with only seventy-five percent usable: four hundred eighty. Three hundred eight therefore rounds up to one node. Switching only to the four hundred five billion parameter tier assigns eight hundred ten weight gigabytes; with the same cache term, nine hundred seventy-eight requires three nodes. These are planning assumptions, not verified hardware or model memory measurements. Default latency adds one hundred sixty for guards, seven hundred fifty for generation, three hundred twenty for tools and zero for evaluation, totaling twelve hundred thirty milliseconds. Model, guard and topology selections choose constants. Workload and context do not alter this latency sum. Export serializes a blueprint using these estimates; it does not provision GPU nodes, invoke model runtimes or validate a deployment schema. Offline graph libraries may remain unavailable, so no real cluster functionality is certified.