1. Sampling Strategy
8 / 24 Traces
Random Baseline
Control
Unbiased pull of everyday production flows (33% sample).
Stratified Balance
Tier / Reason
Equal quota across VIP tiers & refund edge cases.
Failure-Weighted
High Signal
Over-samples errors, rule breaches & dispute fallouts.
Sampled Trace Queue
2. Span Tree & Rubric Evaluation
Trace TR-104
Automated Evaluation Rubric
Refund within Terms Deterministic
Checks order delivery timestamp ≤ 30 days and refund amount ≤ total value.
No Invented Terms LLM-as-Judge
Flags hallucinated restocking fees or non-existent company policies.
Escalate Chargebacks Safety
Requires handoff to human supervisor if active dispute/chargeback exists.
✓ Rubric Verification: PASSED
Trace adheres to all enabled criteria. Eligible for golden SFT dataset.
Execution Spans (OpenTelemetry OTLP GenAI)
4 spans • 1,280 ms • 642 tokens
3. Training Dataset Compiler
5 Golden Runs
5
Dataset Pairs
3,410
Est. Tokens
3
Violations Flagged
Compiled Output (sft.jsonl)
{"messages": [{"role": "system", "content": "You are a customer support agent..."}, ...]}
Dataset exported to agent_traces_dataset.jsonl
* Ingested traces follow OpenTelemetry GenAI semantic conventions. Compiles verified behavioural training rows ready for Axolotl, Unsloth, or Llama-Factory.