Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLM
-
Updated
Jul 27, 2026 - Python
Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLM
Open-source pre-training implementation of Google's LaMDA in PyTorch. Adding RLHF similar to ChatGPT.
[CVPR 2024] Code for the paper "Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model"
The ParroT framework to enhance and regulate the Translation Abilities during Chat based on open-sourced LLMs (e.g., LLaMA-7b, Bloomz-7b1-mt) and human written translation and evaluation data.
Product analytics for AI Assistants
[ECCV2024] Towards Reliable Advertising Image Generation Using Human Feedback
Dataset Viber is your chill repo for data collection, annotation and vibe checks.
Code for the paper "Aligning LLM Agents by Learning Latent Preference from User Edits".
[ICML 2024] Code for the paper "Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases"
Pause your AI agent. Ask a human. Resume with their answer. Open source human-in-the-loop (HITL) library for production LLM agents: Slack, email, and web dashboard. Typed Pydantic and Zod responses. Durable Temporal and LangGraph adapters. AI verifier. Audit trail. Self-hosted, Apache 2.0. Python and TypeScript.
[ NeurIPS 2023 ] Official Codebase for "Aligning Synthetic Medical Images with Clinical Knowledge using Human Feedback"
Documentation at
A self-evolving persona agent for group chats and DMs. Built to sound like a regular, not a help desk. Learns from real user feedback through auditable evidence, gated promotion, rollback, and dynamic few-shot retrieval.
Reinforcement Learning from Human Feedback with 🤗 TRL
🤖 Enhance reinforcement learning stability and efficiency with advanced algorithms like TRPO, PPO, DPO, GRPO, DAPO, and GSPO for optimized policy training.
MCP server for human-in-the-loop surveys, A/B preference tests, ratings, and rankings. Get real human feedback inside Claude Code, Claude Desktop, Cursor, Windsurf, and any MCP client — powered by Datapoint AI.
Break out of the AI training bubble
IEvoAgent: Evolving Conversational Agent based on User Implicit Feedback (ACL 2026 Oral). Leverages conditional feedback distribution matrix with offline KTO alignment and inference-time prompt evolution for multi-turn dialogue.
Tested simulation of clarification-guided reward learning from human state corrections, developed during CMU RISS.
REactive Behavior Constraint-Aware Tree learning (REBCAT) - a human-robot collaboration framework to learn task from demonstrations. Interpretable, fast, object-centric, and reactive.
Add a description, image, and links to the human-feedback topic page so that developers can more easily learn about it.
To associate your repository with the human-feedback topic, visit your repo's landing page and select "manage topics."