Reinforcement Learning, RLHF, and Model Evaluation Metrics

Covers RL, RLHF, model fit (overfitting/underfitting), confusion matrix, Precision/Recall/F1/AUC-ROC, and LLM metrics (ROUGE, Perplexity) — all in one exam-focu

Reinforcement Learning, RLHF, and Model Evaluation Metrics

Understanding when to use Recall vs Precision and how confusion matrices work is essential for the AIF-C01 exam. This guide uses everyday analogies to make these concepts clear.

---

 

Reinforcement Learning (RL)

RL is learning by trial and error. An agent interacts with an environment, takes actions, and receives rewards or penalties.

| Element | Description | Example | |---|---|---| | Agent | The learner/decision-maker | Game character | | Environment | The world it interacts with | Game map | | Action | What the agent can do | Move, jump | | Reward | Feedback from the environment | Score +/- | | State | Current situation | Character position | | Policy | Strategy for choosing actions | Go forward if path is clear |

Use cases: chess AI, robotics, autonomous driving, portfolio management.

---

 

RLHF — Training AI with Human Preferences

RLHF (Reinforcement Learning from Human Feedback) is the technique that made ChatGPT align with human expectations. Without it, AI might give technically correct but unhelpful or offensive answers.

Four steps: Collect human-written ideal responses Fine-tune the base model on those examples Build a reward model by having humans rank model outputs Use the reward model to automatically optimize the AI

---

 

Overfitting vs Underfitting

Think of exam preparation: Underfitting: Didn't study enough. Fails both practice tests and the real exam. Overfitting: Memorized only last year's questions. Aces practice but fails on new question types.

| State | Training Performance | Real Performance | |---|---|---| | Overfitting (High Variance) | Excellent | Poor | | Underfitting (High Bias) | Poor | Poor | | Balanced | Good | Good |

---

 

Confusion Matrix

!The confusion matrix: true/false positives and negatives

| | Actual YES | Actual NO | |---|---|---| | Predicted YES | TP (correct positive) | FP (false alarm) | | Predicted NO | FN (missed positive) | TN (correct negative) |

---

 

Classification Metrics

| Metric | Formula | Use When | |---|---|---| | Accuracy | (TP+TN)/Total | Balanced dataset | | Recall | TP/(TP+FN) | Missing a positive is dangerous (fraud, cancer) | | Precision | TP/(TP+FP) | False positives are costly (spam filter, drug test) | | F1 Score | Harmonic mean of P and R | Imbalanced dataset, both matter |

AUC-ROC: 0.5 = random classifier, 1.0 = perfect classifier.

---

 

LLM-Specific Metrics

| Metric | Measures | |---|---| | Perplexity | How well the model predicts the next word (lower = better) | | ROUGE-1/2 | N-gram overlap for text summarization | | ROUGE-L | Longest common subsequence |

---

 

Exam Key Points

RL has 6 elements: Agent, Environment, Action, Reward, State, Policy. RLHF aligns LLMs with human preferences in 4 steps. Overfitting = high variance; underfitting = high bias. Recall: minimize FN (fraud/cancer detection). Precision: minimize FP (spam filter, drug testing). AUC range: 0.5 (random) to 1.0 (perfect). ROUGE evaluates text summarization quality.

Back to blog list