Open sourceActive

Unsloth

Unsloth is an open-source library for fine-tuning LLMs with 2x faster training and 70% less memory usage.

Open source page

Product features

Product features

Basics/Tips

🏃 From RLHF, PPO to GRPO and RLVR

GRPO notebooks:

How GRPO Trains a Model

🤞 Luck (well Patience) Is All You Need

💡 Reinforcement Learning (RL) Guide

📋 Reward Functions / Verifiers

RL on unsupported models:

💻 Training with GRPO

❓ What is Reinforcement Learning (RL)?

🦥 What Unsloth offers for RL

🦥 What you will learn