Complete Guide to FM Training and Fine-tuning

Master fine-tuning techniques (RLHF, LoRA/PEFT) and evaluation metrics (ROUGE, BLEU, BERTScore) with comparison tables and beginner-friendly analogies.

What is Foundation Model Customization?

Foundation Models (FMs) like Claude, GPT, and Titan are pretrained on massive internet text — great for general tasks, but lacking domain-specific knowledge your business needs.

Think of a new college graduate: solid fundamentals, but needs job-specific training to be effective at a particular company. FM customization is the same concept.

Training FMs from scratch costs tens of millions of dollars. Most businesses start from an existing FM and adapt it. Here are the main approaches.

---

 

Customization Methods Compared

| Method | Description | Cost | Weight Update? | |--------|-------------|------|---------------| | Prompt Engineering | Adjust inputs only | Very low | No | | RAG | Retrieve external docs + add to prompt | Low | No | | Instruction Tuning | Fine-tune with instruction-response pairs | Medium | Yes | | Domain Adaptation | Additional training on domain text | High | Yes | | RLHF | Learn from human preference feedback | Very high | Yes |

Key distinction: Prompt engineering and RAG never touch model weights. Fine-tuning methods (Instruction Tuning, Domain Adaptation, RLHF) update model parameters internally.

!Foundation model customization methods compared by cost and weight updates

Instruction Tuning

Train a model using thousands of instruction-response pairs. Like a teacher giving students thousands of example Q&A pairs and saying "respond like this."

Example training pair: Instruction: "Classify the sentiment: 'Delivery was too slow'" Response: "Negative (delivery speed issue)"

---

 

RLHF — Reinforcement Learning from Human Feedback

The technique behind ChatGPT. Humans rate AI responses, and the AI learns those preferences.

Three steps: Supervised Fine-tuning: Train on high-quality instruction-response data Reward Model Training: Humans pick the better of two AI responses → train a reward model PPO Reinforcement Learning: Optimize the language model using the reward model

| Item | Detail | |------|--------| | Advantage | Outputs aligned with human values and safety standards | | Disadvantage | Human feedback collection is expensive and slow | | Main Use | Improving quality and safety of conversational AI |

---

 

LoRA and PEFT — Efficient Fine-tuning

Problem with full fine-tuning: updating billions of parameters requires massive GPU memory.

PEFT (Parameter-Efficient Fine-Tuning): Updates only 1–5% of parameters, achieving near-full fine-tuning performance with far fewer resources.

LoRA (Low-Rank Adaptation): Decomposes large weight matrices into two smaller matrices. Only the small matrices are trained.

| Comparison | Full Fine-tuning | LoRA | |------------|-----------------|------| | Parameters updated | All (billions) | Small adapters (millions) | | GPU memory | Very high | Low | | Training time | Long | Short | | Performance | Best | 90–95% of full fine-tuning |

---

 

Model Evaluation Metrics

| Metric | Main Use | Measurement | |--------|---------|------------| | ROUGE | Summarization | Word overlap (recall-focused) | | BLEU | Translation | N-gram overlap (precision-focused) | | BERTScore | General purpose | Semantic similarity |

Exam tips: ROUGE = summarization. BLEU = translation. BERTScore = semantic meaning.

---

 

Exam Quick Reference

| Concept | Exam Point | |---------|-----------| | Prompt Engineering | No weight changes, cheapest customization | | RAG | Adds external knowledge, no retraining needed | | Instruction Tuning | Fine-tune with instruction-response pairs | | RLHF | Human preference learning, expensive, used in ChatGPT | | LoRA / PEFT | Updates small subset of parameters efficiently | | ROUGE | Summarization quality (word overlap) | | BLEU | Translation quality (N-gram overlap) | | BERTScore | Semantic similarity evaluation |

Back to blog list