Are you thinking distillation goes from zero to complete model?
I believe it can be used in the RL / fine tuning sense, in which case 150,000 requests, assuming every one was detected, could move the needle in quality.
I agree it couldn’t replace all of pre and post training , but I don’t think that’s the claim. You do typical training, then distill really difficult cases.
I believe it can be used in the RL / fine tuning sense, in which case 150,000 requests, assuming every one was detected, could move the needle in quality.
I agree it couldn’t replace all of pre and post training , but I don’t think that’s the claim. You do typical training, then distill really difficult cases.