On-Policy Distillation: Discussion of Loss Design
Reverse KL, sampled tokens, Top-K correction, and representation alignment: a practical and empirical guide to choosing an OPD objective.
Reverse KL, sampled tokens, Top-K correction, and representation alignment: a practical and empirical guide to choosing an OPD objective.