Second-by-Second Audience Retention Graph
Hover or drag across the curve to inspect drop-off timestamps and varianceSignificant Winner Detected: Cut B (Direct Cold-Open Question Hook)
Based on 90,000 impressions, Cut B produces an additional +0:42 Average View Duration (+525 total hours of watch time) with a 99.9% confidence level. Recommending auto-adoption of Cut B.
Algorithmic & Viewer Experience Audit (MKBHD Dilemma)
Watch Time Economics & Cannibalization
Understanding Video A/B Testing on YouTube
How does Video Edit A/B Testing differ from Title & Thumbnail testing?
Title and Thumbnail testing primarily measures Click-Through Rate (CTR): how many viewers choose to click from the browse feed or search results. Video edit A/B testing measures Retention, Average View Duration (AVD), and End-Screen CTR. When YouTube split-tests the underlying video file (e.g. testing an alternate intro, music cut, pacing, or scene order), viewers are served different encoded video segments to determine which cut retains human attention longer.
What was Marques Brownlee (@MKBHD) concerned about regarding video A/B testing?
As MKBHD pointed out following YouTube's announcement, video A/B testing introduces unprecedented viewer-experience questions: "If I'm watching a video, do I know if it's part of an A/B test? If my friend and I watch the same video, do we see different jokes, explanations, or conclusions?" When creators test content variations, comments referencing a specific joke in Cut A might confuse viewers who saw Cut B, altering the shared cultural commentary around a video.
What sample size is required before declaring a video edit winner?
Unlike thumbnail tests which can reach statistical significance in 5,000–10,000 impressions due to high binary click volumes, retention curves have variance across every second. For typical 8–15 minute videos, reliable AVD comparisons require at least 30,000 to 50,000 completed impressions per variant to avoid false positives caused by natural traffic shifts (e.g. subscriber notification waves vs external algorithmic recommendations).