reinforcement-learning-from-human-feedback

annotated tutorial of the huggingface TRL repo for reinforcement learning from human feedback connecting equations from PPO and GAE to the lines of code in the pytorch implementation

nlp reinforcement-learning deep-learning transformers deep-reinforcement-learning pytorch language-model fine-tuning large-language-models reinforcement-learning-from-human-feedback

Updated Feb 28, 2023
Jupyter Notebook

tlc4418 / llm_optimization

Star

A repo for RLHF training and BoN over LLMs, with support for reward model ensembles.

deep-learning ensembles best-of-n large-language-models reinforcement-learning-from-human-feedback reward-models

Updated Mar 9, 2024
Python

nlp-uoregon / Okapi

Star

Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

multilingual nlp bloom natural-language-processing reinforcement-learning chatbot dataset question-answering llama language-model large-language-models rlhf instruction-tuning reinforcement-learning-from-human-feedback

Updated Aug 18, 2023
Python

tatsu-lab / alpaca_farm

Star

A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.

natural-language-processing deep-learning instruction-following large-language-models reinforcement-learning-from-human-feedback

Updated Feb 24, 2024
Python

PKU-Alignment / safe-rlhf

Star

Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

Updated Apr 20, 2024
Python

OpenLLMAI / OpenRLHF

Star

An Easy-to-use, Scalable and High-performance RLHF Framework (Support 70B+ full tuning & LoRA & Mixtral & KTO)

reinforcement-learning raylib transformers deepspeed large-language-models reinforcement-learning-from-human-feedback vllm

Updated Jun 2, 2024
Python

Improve this page

Add a description, image, and links to the reinforcement-learning-from-human-feedback topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the reinforcement-learning-from-human-feedback topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

reinforcement-learning-from-human-feedback

Here are 11 public repositories matching this topic...

ymnseol / weekly-paper-reading-group

Almost-Intelligence / LMRax

ymetz / rlhfblender

liushunyu / Ask-AC

XplainMind / LLMindCraft

clam004 / minichatgpt

tlc4418 / llm_optimization

nlp-uoregon / Okapi

tatsu-lab / alpaca_farm

PKU-Alignment / safe-rlhf

OpenLLMAI / OpenRLHF

Improve this page

Add this topic to your repo