kto
Here are 20 public repositories matching this topic...
Auditable human-feedback annotation, review, provenance, and frozen training-data export.
-
Updated
Aug 27, 2026 - Python
Advanced LLM fine-tuning techniques: SFT (LoRA, QLoRA, DoRA, P-/Prefix-Tuning), GRPO, DPO, ORPO, KTO & PPO; composable correctness/format rewards + LLM-as-a-Judge evals (DeepEval, Evidently AI) across math, multi-hop, medical & general QA on Llama 3, Mistral, Phi-4, Gemma & Qwen3. Built on TRL, PEFT & Unsloth.
-
Updated
Sep 8, 2026 - Python
GenPark AI Agent Skill - Pairwise Bradley-Terry and Elo rating tournament engine for ranking multi-agent strategies, prompt mutations, and generated tool solutions.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Iterative self-rewarding LLM judge scoring multi-turn agent outputs and synthesizing contrastive hard-negative prompt variations.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Direct Preference Optimization (DPO) implicit reward calculation, reference policy log-ratio tracking, and pairwise preference loss evaluation.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Pairwise Bradley-Terry and Elo rating tournament engine for ranking multi-agent strategies, prompt mutations, and generated tool solutions.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Direct Preference Optimization (DPO) implicit reward calculation, reference policy log-ratio tracking, and pairwise preference loss evaluation.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Prioritized experience replay buffer for agent reinforcement fine-tuning (RFT / GRPO) with advantage estimation and importance sampling weights.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Kahneman-Tversky Optimization (KTO) loss evaluator aligning agents directly on unpaired binary feedback (thumbs-up / thumbs-down) using prospect theory loss aversion.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Kahneman-Tversky Optimization (KTO) loss evaluator aligning agents directly on unpaired binary feedback (thumbs-up / thumbs-down) using prospect theory loss aversion.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Iterative self-rewarding LLM judge scoring multi-turn agent outputs and synthesizing contrastive hard-negative prompt variations.
-
Updated
Sep 6, 2026 - Python
GenPark AI Agent Skill - Prioritized experience replay buffer for agent reinforcement fine-tuning (RFT / GRPO) with advantage estimation and importance sampling weights.
-
Updated
Sep 6, 2026 - Python
Open, MCP-native training-data foundry: turn permitted documents into traceable, quality-gated SFT and preference datasets.
-
Updated
Sep 2, 2026 - Python
Headless LLM fine-tuning in 3 lines — smart defaults, VRAM-aware batch sizing, multi-run SLAO, GGUF export for Ollama.
-
Updated
Oct 1, 2026 - Python
Preference optimization framework for text classification (DPO/ORPO/KTO), with SFT, encoder, and XGBoost baselines plus unified run pipeline and reproducible outputs.
-
Updated
Apr 4, 2026 - Python
🚀 Optimize preferences effectively with ORPO, a framework for monolithic preference optimization without a reference model.
-
Updated
Oct 1, 2026 - Python
Structured game-agent distillation: distilling frontier-LLM gameplay reasoning into a small open student model (Qwen) through a typed MCP tool API, in a persistent 2D MMORPG. SFT + KTO data collection across 4 LLM harnesses, measured against a quest benchmark. (Joint work with Niral Patel.)
-
Updated
Jul 29, 2026 - Python
Add this topic to your repo
To associate your repository with the kto topic, visit your repo's landing page and select "manage topics."