Inference Routing Controls
Strict Local First
Cloud Fallback Allowed
Offload overflow tokens if VRAM saturates
Hardware Basis: Windows on GeForce RTX 3090/4090 and RTX 5000/6000 Ada with 24GB+ VRAM executes 8B-70B quant models fully in-device, preventing telemetry exposure while running Perplexity Portable.
Local Handled Tokens
12,400
100% On-Device
Cloud Offloaded Tokens
0
0% Network egress
Peak VRAM Usage
21.4 GB
2.6 GB Headroom
Est. Latency (TTFT)
420 ms
~88 tok/s tensor speed
NVIDIA Tensor Execution Pipeline
100% Local Device Private
VRAM Allocation
21.4 / 24.0 GB
Local KV Cache & Weights
Cloud Overflow
Privacy Boundary Guard
0 Leaks Detected
Local context isolation is verified. Cryptographic memory rings on Windows GeForce RTX guarantee zero plain-text transit.
Local PCIe Bus
Memory Ring 0 Isolated
STREAM LOG [Transformers.js / WebLLM Prim]
COMPUTE: GeForce RTX @ 24GB
[00.00s] > Context Ingest: 12,400 tokens parsed from local workspace.
[00.12s] > Allocation: Model weights (14.2 GB) + KV-Cache (7.2 GB) = 21.4 GB VRAM.
[00.28s] > Verification: Strict Local-First active. 0 external calls issued.
[00.42s] > Generation: Synthesizing response with 100% on-device Tensor Cores...
[00.12s] > Allocation: Model weights (14.2 GB) + KV-Cache (7.2 GB) = 21.4 GB VRAM.
[00.28s] > Verification: Strict Local-First active. 0 external calls issued.
[00.42s] > Generation: Synthesizing response with 100% on-device Tensor Cores...
Audit Report Summary
| Metric Field | Simulated Value | Constraint / Boundary |
|---|---|---|
| Workload Type | sensitive_legal_code_review | User Scenario |
| Privacy Score | 100% Local Device Private | Zero Egress Enforced |
| VRAM Peak Usage | 21.4 GB | 24.0 GB Hardware Limit |
| Local Handled Tokens | 12400 | 100.0% of Total |
| Cloud Offloaded Tokens | 0 | 0.0% Overflow |
| Estimated Latency (TTFT) | 420 ms | Local RTX Tensor Pipeline |
| Sensitive Data Leaks | 0 | Safe Isolation |