Category: Alignment & Safety

  • Alignment & Safety – Frontier AI Research Brief (W26 2026)

    Alignment & Safety – Frontier AI Research Brief (W26 2026)

    A focused look at this week’s most significant advances in alignment & safety — 4 papers surveyed from arXiv and leading AI labs. — Safety research is broadening from alignment to encompass robustness, interpretability, and the systemic risks of deployed AI. This week brings new attacks, new defenses, and deeper understanding of model internals. Key…

  • Alignment & Safety – Frontier AI Research Brief (W28 2026)

    Alignment & Safety – Frontier AI Research Brief (W28 2026)

    Safety and alignment research continues to mature rapidly in W28, with papers addressing everything from constitutional classifiers and unlearning techniques to watermarking and jailbreak prevention. As AI systems are deployed more widely, the research community is responding with increasingly sophisticated approaches to keeping them beneficial. Key Developments This Week Constitutional and Classifier-Based Safety. HaloGuard 1.0…

  • Alignment & Safety – Frontier AI Research Brief (W28 2026)

    Alignment & Safety – Frontier AI Research Brief (W28 2026)

    Safety and alignment research continues to mature rapidly in W28, with papers addressing everything from constitutional classifiers and unlearning techniques to watermarking and jailbreak prevention. As AI systems are deployed more widely, the research community is responding with increasingly sophisticated approaches to keeping them beneficial. Key Developments This Week Constitutional and Classifier-Based Safety. HaloGuard 1.0…

  • Alignment & Safety – Frontier AI Research Brief (W26 2026)

    Alignment & Safety – Frontier AI Research Brief (W26 2026)

    A focused look at this week’s most significant advances in alignment & safety — 4 papers surveyed from arXiv and leading AI labs. — Safety research is broadening from alignment to encompass robustness, interpretability, and the systemic risks of deployed AI. This week brings new attacks, new defenses, and deeper understanding of model internals. Key…

  • Week 22, 2026 — Agentic Systems & Skills

    Week 22, 2026 — Agentic Systems & Skills

    Agent research had a breakthrough week, with advances in skill optimization, long-horizon memory management, and production-scale deployment of autonomous code review. SkillOpt: Training Agent Skills Like Neural Network Weights SkillOpt by Yifan Yang et al. introduces the first systematic controllable text-space optimizer for agent skills. An optimizer model turns scored rollouts into bounded add/delete/replace edits…

  • The Year Alignment Got Empirical: When, Where, and for Whom Do Models Fail?

    The Year Alignment Got Empirical: When, Where, and for Whom Do Models Fail?

    55 papers surveyed | May 2025 – May 2026 — For years, AI alignment lived in the realm of principles. Papers opened with “it is important that AI systems align with human values” and closed with hand-waved suggestions for future work. In 2025–2026, that changed. The field stopped asking is the model safe? and started…