Course Schedule

Class meets Wednesdays, 12–3pm, in UTA 1.212. Written work is due at 11:59pm on the listed date; in-class assessments and presentations take place during class. Schedule is subject to change; any changes will be announced in class and posted to Canvas.

WkDateLectureLab / WorkshopDue
1Aug 26Intro to post-training [slide]TACC account setup; base vs. instruction-tuned model comparisonHW0 out
2Sep 2LLM landscape. Prompting and in-context learning as baselinesChat templates, in-context learning, and cross-model comparison; project team formation
3Sep 9Evaluation I: task metrics, held-out design, what a benchmark can and cannot tell youProposal workshopHW0 due
4Sep 16Evaluation II: LLM-as-a-judge, rubric design, judge bias and agreementBuild and validate a judge pipeline
5Sep 23Supervised fine-tuning: objective, chat templates, hyperparametersFirst SFT run on TACCProposal due; HW1 out
6Sep 30Parameter-efficient fine-tuning: LoRA, QLoRALoRA rank ablation and cost comparison
7Oct 7Data for post-training: curation, synthetic generation, distillationSynthetic data generation and filteringHW1 due
8Oct 14RLHF foundations: Bradley–Terry reward models, PPOTrain and probe a reward model
9Oct 21Direct preference optimization (DPO)Preference labeling and guided DPO workflowHW2 out
10Oct 28RLAIF and Constitutional AI. Scalable oversight and its limitsAI-feedback pipeline; project checkpoint
11Nov 4Verifiable rewards and GRPO. Reasoning-oriented post-trainingGuided GRPO labHW2 due
12Nov 11Failure modes: reward hacking, sycophancy, catastrophic forgetting, safety regressionsConcept check (in class, handwritten, closed-device), then project work
13Nov 18Project presentationsPresentation feedback and discussionFinal presentation
Nov 25No class (Thanksgiving break)
14Dec 2Feedback-based project revision and refinementFinal report due Dec 7