After things are reconciled, we need to add support for preference tuning pipelines Introduce 3 pipelines: - Data annotation routing > route samples for preference tuning and annotate - Student Model response generation - Log likelihood generation for weak vs strong model for reward calculation
After things are reconciled, we need to add support for preference tuning pipelines
Introduce 3 pipelines: