Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
↑ +22 today
Discover the most starred and trending open source tools tagged with #rlhf.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
FineEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Train Large Language Models on MLX.
Auditable human-feedback annotation, review, provenance, and frozen training-data export.