AI Fundamentals
AI Safety, Alignment, and Governance
78 lessons in AI Fundamentals
- International Coordination and AI PolicySlides / Video
- Model Cards and Documentation for AccountabilitySlides / Video
- AI Standards and Risk FrameworksSlides / Video
- AI Governance Frameworks and PrinciplesSlides / Video
- Dual-Use and Misuse Risks of AISlides / Video
- Near-Term vs Long-Term AI RisksSlides / Video
- Corrigibility and Human ControlSlides / Video
- Hallucination and Truthfulness in LLMsSlides / Video
- Guardrails and Safety LayersSlides / Video
- Red Teaming AI SystemsSlides / Video
- Scalable Oversight and Supervision LimitsSlides / Video
- Distributional Shift and Robustness FailuresSlides / Video
- Instrumental Convergence and Goal MisgeneralizationSlides / Video
- Reinforcement Learning from Human FeedbackSlides / Video
- Value Alignment and Human PreferencesSlides / Video
- Specification Gaming and Reward HackingSlides / Video
- Outer vs Inner AlignmentSlides / Video
- The AI Alignment Problem DefinedSlides / Video
- What AI Safety Is and Why It MattersSlides / Video
- Compute GovernanceFlashcard
- AI Incident Reporting and MonitoringFlashcard
- Safety-Capability TradeoffsFlashcard
- Frontier Model Risks and EvaluationsFlashcard
- Watermarking and Provenance of AI ContentFlashcard
- Open vs Closed Model Release DecisionsFlashcard
- Content Safety and GuardrailsQuiz
- Safety vs CapabilitiesQuiz
- Mechanistic InterpretabilityQuiz
- Existential and Catastrophic RiskQuiz
- Instrumental ConvergenceQuiz
- Model CardsQuiz
- Risk-Based RegulationQuiz
- Power Concentration from AIQuiz
- Outer vs Inner AlignmentQuiz
- Robustness and Distributional ShiftQuiz
- Specification GamingQuiz
- Frontier Model EvaluationsQuiz
- Scalable OversightQuiz
- Corrigibility and InterruptibilityQuiz
- Goal MisgeneralizationQuiz
- Value AlignmentQuiz
- Adversarial AttacksQuiz
- AI Governance FrameworksQuiz
- Reward MisspecificationQuiz
- Responsible AI PrinciplesQuiz
- Orthogonality ThesisQuiz
- Constitutional AIQuiz
- Interpretability for SafetyQuiz
- Staged ReleaseQuiz
- Dual-Use and MisuseQuiz
- Near-Term vs Long-Term SafetyQuiz
- Red Teaming AI SystemsQuiz
- The AI Alignment ProblemQuiz
- AI Safety OrganizationsQuiz
- Constitutional AI and AI FeedbackFlashcard
- Near-Term vs Long-Term AI RisksFlashcard
- Responsible AI PrinciplesFlashcard
- Dual-Use and Misuse of AIFlashcard
- Model Cards and DocumentationFlashcard
- Scalable OversightFlashcard
- Robustness and Distribution ShiftFlashcard
- Risk-Based Regulatory ApproachesFlashcard
- Red Teaming AI SystemsFlashcard
- AI Standards and FrameworksFlashcard
- AI Governance Definition and ScopeFlashcard
- Adversarial Attacks and RobustnessFlashcard
- Orthogonality ThesisFlashcard
- Value Alignment and Human ValuesFlashcard
- What is AI SafetyFlashcard
- The Alignment ProblemFlashcard
- Specification Gaming / Reward HackingFlashcard
- Corrigibility and Shutdown ProblemFlashcard
- Existential and Catastrophic Risk from AIFlashcard
- AI Control ProblemFlashcard
- Goal MisgeneralizationFlashcard
- Outer vs Inner AlignmentFlashcard
- Instrumental ConvergenceFlashcard
- What is AI AlignmentFlashcard