10-423 / 10-623 / 10-723Generative AI
Sign in
Lecture 1: RNN LMs / AutodiffLecture 2: Transformer LMsLecture 3: Learning LLMs / DecodingLecture 4: Pre-training, fine-tuning / Modern TransformersLecture 5: Computer Vision: CNNs / Encoder-only Transformers / Vision TransformersLecture 6: Generative Adversarial Networks (GANs) / PGMLecture 7: Diffusion models (Part I)Lecture 8: Diffusion models (Part II) / Score MatchingLecture 9: Variational Autoencoders (VAEs) / Continuous Normalizing Flows / Flow MatchingLecture 10: Parameter-efficient fine tuningLecture 11: In-Context Learning / Prompt Engineering / Instruction Fine-tuning / Reinforcement learning with human feedback (RLHF)Lecture 12: Direct Preference Optimization (DPO) / Text-to-image generation / Latent diffusion modelLecture 13: Vision-language modelsLecture 14: Cross-Attention / Diffusion Transformer / Prompt-to-PromptLecture 15: Querying Transformer / Scaling LawsLecture 16: Mixture of ExpertsLecture 17: Distributed trainingLecture 18: Flash Attention / Efficient decoding strategiesLecture 19: Long Context in LLM / RAGLecture 20: Reasoning ModelsLecture 21: State Space Models / Hybrid ModelsLecture 22: Real-world Issues and Considerations / What can go wrong? / SafetyLecture 23: Audio understanding and synthesisLecture 24: Code Generation / Autonomous AgentsLecture 25: Generative Models for VideosLecture 26: Interactive World Models

Generative models of text › Pre-training, fine-tuning / Modern Transformers

Lecture 4: Pre-training, fine-tuning / Modern Transformers

Wed, Sep 2

Readings

  • GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. Ainslie et al. (2023). EMNLP.
  • Longformer: The Long-Document Transformer. Beltagy et al. (2020).
  • RoFormer: Enhanced Transformer with Rotary Position Embedding. Su et al. (2021).

Unit: Generative models of text

PDF not available
PDF not available