仕事内容
<div class="content-intro"><h2><strong>About Anthropic</strong></h2>
<p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2><strong>About the role<br></strong></h2>
<p>As a Safeguards Enforcement Analyst focused on Conventional Weapons, your work spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, specifically utilizing conventional weapons and dangerous technology. You will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas. </p>
<p><br>Important context for this role: In this position you may be exposed to and engage with explicit content spanning a range of topics, including those of a violent, graphic, hateful, or psychologically disturbing nature.</p>
<h2><strong>Key responsibilities</strong></h2>
<ul>
<li>
<p>Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy</p>
</li>
<li>
<p>Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements</p>
</li>
<li>
<p>Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations</p>
</li>
<li>
<p>Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts + emerging tactics</p>
</li>
<li>
<p>Support the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement sc