CSAIL Interpretability

AI Training Data Attribution & Unlearning Lab

Ablation Setup 1 Target
CSAIL Finding: Even with 100% training exemplars removed, style retention persists due to distributed latent subspace entanglement with co-occurring art concepts.
2D Latent Feature Manifold & Gradient Trajectory Drag probe to test generation prompt
Prompt Latent Probe: (0.42, 0.65) ● Training Exemplar | ◆ Probe | ⤹ Gradient Pull
Attribution Telemetry Verified TracIn
Residual Style Similarity
74.2%
Copyright Risk Index
High
TracIn Influence (Top 1)
0.842
Model Utility Loss
1.8%
Copyright Risk: Deleting explicit training samples does not defeat copyright similarity infringement tests under current unlearning bounds.

Latent Style Entanglement & Dataset Unlearning Dynamics

Read the explanation

In generative models, artistic styles do not occupy isolated training points. Instead, target exemplars share high-dimensional latent manifolds with adjacent genre concepts. When a user increases the pruned exemplar ratio to one hundred percent, direct training samples are deleted, but the unlearning factor fails to sever interconnected subspace projections. Dragging a latent probe near the cluster reveals strong TracIn gradient pull. Because indirect entanglement persists, the residual style similarity sustains high copyright risk.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.