Testimony from Elon Musk about xAI training its model on outputs from OpenAI models reignited public debate in the period over distillation and training data ethics.
Distillation, in sector jargon, means training a smaller model using a larger one's outputs. It works because a capable model's answer is itself high-quality training material, and it allows approaching the expensive model's capability without paying for its development.
That's why the technique bothers anyone investing billions in training. It turns the product into a competitor's raw material: all you need is access to the interface, which is sold to anyone.
The debate would resurface weeks later in a frontier company's public position, arguing for anti-distillation enforcement as one necessary measure, alongside chip export controls and required safety testing.
For operators, the subject has practical effects in terms of service. Vendor contracts began including explicit prohibitions on using outputs to train competing models, and that applies even to anyone who only wanted to tune a small model on generated examples.
