FA-0608Agentic & Generative AISoftware DevelopmentData & Analytics

Efficient LLM Customisation

A one-day comparison of fine-tuning, parameter-efficient adaptation and RAG

Compare fine-tuning and retrieval, inspect a small LoRA workflow and review how data quality and evaluation guide model customisation.

Introduction

Why this course

This one-day technical workshop compares prompting, transfer learning, fine-tuning and retrieval-augmented generation for a defined task. A prepared small-model LoRA exercise or demonstration illustrates adaptation; a selected RAG pipeline illustrates supplying external context without changing model weights.

The day focuses on decisions, data preparation and evaluation rather than scalable production deployment or large distributed training. Hardware needs depend on model size, precision, context, batch settings and memory—not a universal CUDA-core count. Model/API access and licences must be appropriate for the chosen exercises; paid services are optional and any API billing is separate from consumer chatbot subscriptions.

Learning outcomes

Learning outcomes

The workshop teaches participants to:

  • Distinguish fine-tuning that changes model parameters from RAG that supplies retrieved context.
  • Choose a candidate model or supported hosted adaptation path based on task, access and licence constraints.
  • Prepare a small approved dataset and separate training/evaluation examples.
  • Inspect or run a bounded LoRA exercise and compare baseline and adapted behaviour.
  • Explain PEFT methods and hardware/software trade-offs without assuming universal compatibility.
  • Review a RAG pipeline's retrieval, context quality, access boundaries and evaluation needs.
Prerequisites

Prerequisites

  • Strong Python skills and familiarity with a deep-learning framework and Hugging Face workflows.
  • Comfort with Linux/command-line tools; containers and cloud concepts are helpful.
  • Access to a prepared suitable compute environment and approved/licensed model files or services; demonstrations are available where training resources are insufficient.
  • Required accounts and model approvals for the selected exercise, without an assumed paid chatbot or Colab subscription.
Training outline

4 modules

·
01Module 1 — Customisation Choices and Model Selection2 topics
  • Why customize LLMs?
    • Challenges of generic pre-trained models in domain-specific tasks.
    • Customization strategies: Transfer learning, fine-tuning, and RAG.
  • LLM architectures and training pipelines: A quick refresher.

Compare current supported models and APIs by capability and access; historical GPT-4/PaLM listings are not treated as a current availability catalogue.

  • Tokenization and vocabulary handling in LLMs.
  • Principles of transfer learning in the LLM ecosystem.
    • Adapting pre-trained models for new domains.
    • Reuse pretrained knowledge where appropriate; compare data, cost and task-fit trade-offs with training from scratch.
  • Identifying appropriate pre-trained models for customization.
    • Using Hugging Face Model Hub and OpenAI APIs.
    • Evaluating models based on task requirements (e.g., size, performance, licensing).

Distinguish download/adapter-based workflows from provider-supported fine-tuning APIs; not every hosted model permits the same adaptation.

02Module 2 — Dataset Preparation2 topics
  • Building high-quality datasets for LLM customization.
    • Collect approved domain data from authorised sources/APIs; do not assume scraping rights or confidentiality clearance.
    • Cleaning and preprocessing datasets: Removing noise and duplicates.
    • Annotation strategies for supervised tasks.
  • Tools for dataset preparation:
    • Tokenization with Hugging Face Tokenizers.
    • Version control with DVC (Data Version Control).
    • Embedding-based dataset evaluation.

Use approved data with rights and privacy checks; prevent training/evaluation leakage and keep an untouched evaluation set.

Embedding similarity is one inspection method, not a complete measure of dataset quality.

03Module 3 — Fine-Tuning and PEFT4 topics
  • Overview of fine-tuning strategies:
    • Full fine-tuning vs parameter-efficient fine-tuning.
  • Advanced fine-tuning methods:
    • LoRA (Low-Rank Adaptation): Principles and implementation.
    • PEFT approaches and tooling: benefits and limitations depend on the method, model and task.
    • Prefix-tuning and adapter-based tuning techniques.
  • Tools for fine-tuning LLMs:
    • Hugging Face Trainer for streamlined training pipelines.
    • PyTorch Lightning/DeepSpeed distributed-training concepts as an overview, not a full one-day cluster deployment.
  • Hands-on exercise:
    • Run or inspect a prepared small-model LoRA exercise using a limited approved dataset.
    • Evaluate the fine-tuned model's performance with downstream tasks (e.g., text classification or summarization).

LoRA is a parameter-efficient technique; PEFT is the wider family/tooling, not an unrelated competing algorithm.

Run or inspect one small prepared example; prefix/adapters and distributed training are comparative overviews. Evaluate quality, limitations and resources rather than assume reduced parameters guarantee better results.

04Module 4 — RAG and Comparative Review2 topics
  • Introduction to RAG: Enhancing LLMs with external knowledge.
    • How RAG works: Combining LLMs with vector databases and retrieval systems.
  • Core components of a RAG system:
    • Vector embeddings and dense retrieval (e.g., FAISS, Weaviate, Milvus).
    • Document stores and retrievers (e.g., ElasticSearch, LangChain).
    • RAG pipelines in Hugging Face and LangChain.

Inspect one selected retrieval pipeline and compare retrieval/context failures with fine-tuning limitations. RAG does not guarantee truthful output or eliminate prompt injection.

Choose a proportionate approach for the target task and identify remaining deployment/security/evaluation work.

A programme built around your team.

Share your training goals and requirements.

Efficient LLM Customisation
FA-0608

Share your requirements for this programme.

Training enquiry