Redwood Research

Redwood Research is a nonprofit organization focused on AI safety and security research, especially technical methods for reducing risks from increasingly capable AI systems. Based in Berkeley, California, it has become one of the better-known independent research groups working on AI alignment, interpretability, evaluations, and AI control, while collaborating with and advising major AI developers.

Research focus

Redwood Research studies ways to make advanced AI systems more reliable and safer to deploy, with an emphasis on scenarios where highly capable models might intentionally or unintentionally behave in ways that conflict with human goals. Its published work spans mechanistic interpretability, adversarial training, AI control (techniques for safely using potentially untrusted AI systems), and evaluations of model behavior. The organization publishes papers, technical reports, and blog posts describing both theoretical ideas and empirical experiments.

Approach

A distinguishing feature of Redwood Research is its emphasis on practical, technical safety research rather than primarily policy or governance. Its researchers investigate methods that AI developers could apply to current or near-future models, including monitoring, auditing, interpretability, and containment strategies intended to reduce the risk of deceptive or strategically misaligned behavior. Recent writing has also explored the idea of using AI systems to help perform AI safety research under carefully designed safeguards.

Organization and impact

Founded in 2021, Redwood Research operates as a 501(c)(3) nonprofit. The organization has collaborated with or advised major AI companies on safety-related work and contributes to the broader AI safety research community through publications, open technical discussions, and public communication. While its research agenda reflects a view that advanced AI could pose significant long-term risks, many of its technical contributions—such as work on interpretability and robustness—are also relevant to improving today's AI systems.

Current priorities

According to its public research agenda, Redwood Research currently prioritizes work on threat assessment and mitigation for advanced AI systems, including AI control, evaluations of potentially dangerous capabilities, and techniques for increasing confidence that powerful models behave as intended. These priorities continue to evolve as AI capabilities advance and new safety challenges emerge.