使用TRL和LoRA在Anthropic HH-RLHF上通过直接偏好优化审计偏好偏差并微调语言模型
Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA
In this tutorial, we design an end-to-end preference-learning workflow using the Anthropic HH-RLHF dataset and Direct Preference Optimization (DPO). We begin by preparing a robust Colab environment, loading and parsing chosen–rejected response pairs, and auditing the dataset for structural and length-based preference biases. We then run lexical shortcut diagnostics to determine whether surface-level linguistic patterns can separate preferred from rejected responses, prepare conversational data with tokenizer-aware length filtering, and construct a version-robust DPO training pipeline with TRL and optional LoRA adaptation. Finally, we fine-tune a Qwen2.5-0.5B-Instruct model, evaluate reward accuracy and training behavior, analyze performance across individual HH-RLHF subsets, inspect...
Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA
In this tutorial, we design an end-to-end preference-learning workflow using the Anthropic HH-RLHF dataset and Direct Preference Optimization (DPO). We begin by preparing a robust Colab environment, loading and parsing chosen–rejected response pairs, and auditing the dataset for structural and length-based preference biases. We then run lexical shortcut diagnostics to determine whether surface-level linguistic patterns can separate preferred from rejected responses, prepare conversational data with tokenizer-aware length filtering, and construct a version-robust DPO training pipeline with TRL and optional LoRA adaptation. Finally, we fine-tune a Qwen2.5-0.5B-Instruct model, evaluate reward accuracy and training behavior, analyze performance across individual HH-RLHF subsets, inspect potential length bias, generate sample responses, and save the resulting policy for further experimentation.
Copy CodeCopiedUse a different Browser
We set up the required libraries, handle dependency compatibility issues, and configure the main parameters used throughout the tutorial. We also initialize reproducibility settings and inspect the available hardware, precision modes, and installed TRL interfaces. This gives us a stable environment before we process the HH-RLHF dataset and train the preference model.
We load samples from the different Anthropic HH-RLHF subsets and create balanced training and testing datasets. We parse each conversation into structured user and assistant messages while ensuring that chosen and rejected responses share the same conversational prefix. We then filter invalid pairs so that we work only with properly aligned preference examples.
We analyze the preference pairs to measure differences in response length, conversation depth, and source-specific behavior. We also train a TF-IDF and logistic regression diagnostic to test whether simple lexical patterns can distinguish chosen responses from rejected ones. This helps us detect shortcuts that the language model could potentially exploit instead of learning the intended preference signal.
We prepare the tokenizer, apply the conversational chat template, and calculate token lengths for every preference pair. We filter examples that exceed our prompt or total sequence limits and dynamically construct DPO configuration arguments based on the installed TRL version. We then load the base model, configure LoRA when available, and build the DPO trainer that we use for fine-tuning.
We train the model using Direct Preference Optimization with the configured batch size, gradient accumulation, learning rate, and optimization steps. We evaluate the resulting policy on held-out preference pairs and inspect metrics such as loss, reward margins, and reward accuracy. We also visualize the training history to observe how preference-learning performance changes throughout optimization.
We calculate per-source reward accuracy and compare policy and reference-model log probabilities to examine whether the model genuinely prefers the chosen responses. We investigate the relationship between preference decisions and response-length differences, then generate sample answers from the tuned policy to inspect its behavior qualitatively. Finally, we save both the trained model and tokenizer so that we can reuse the resulting DPO policy in later experiments.
In conclusion, we developed a complete DPO-based preference-learning pipeline that goes beyond simply fine-tuning a language model on chosen and rejected responses. We examined the HH-RLHF data for length asymmetries and lexical shortcuts, enforced consistent conversational formatting and token limits, and used a flexible training setup that adapts to differences across TRL and Transformers versions. We also evaluated the tuned policy at both the aggregate and per-source levels, allowing us to identify whether improvements reflect genuine preference learning or undesirable shortcuts such as favoring longer answers. By combining dataset auditing, diagnostic analysis, efficient LoRA-based DPO training, reward evaluation, and generation testing, we established a framework for studying and improving preference alignment in language models.
Check out the FULL CODES here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.
Developing an End-to-End Document Intelligence Pipeline with docTR for OCR, Layout Analysis, KIE, Benchmarking, and Searchable PDFs
Fine-Tuning Tool-Calling LLMs: A Complete Guide Using XYZ-Aquila-SFT and Qwen3
Create a Reasoning-Focused LLM: A Practical Guide to Streaming, Curating, and Fine-Tuning the SupraLabs Reasoning Corpus
AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation
Previous articleNVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands