Alignment Red Team - Research Engineer/Research Scientist
<div class="content-intro"><h2>About the AI Security Institute</h2>
<p>The AI Security Institute is the world's largest and best-funded team dedicated to understanding advanced AI risks and translating that knowledge into action. We’re in the heart of the UK government with direct lines to No. 10 (the Prime Minister's office), and we work with frontier developers and governments globally.</p>
<p>We’re here because governments are critical for advanced AI going well, and UK AISI is uniquely positioned to mobilise them. With our resources, unique agility and international influence, this is the best place to shape both AI development and government action.</p></div><h2><strong>The deadline for applying to this role is 11th October 2026, end of day, anywhere on Earth. </strong></h2>
<h2><strong><span data-contrast="none"><span data-ccp-parastyle="heading 2">Team Description </span></span></strong><span data-contrast="none"><span data-ccp-parastyle="heading 2"> </span></span><span data-ccp-props="{"134245418":true,"134245529":true,"335557856":16777215,"335559738":240,"335559739":240}"> </span></h2>
<p><span data-contrast="none">Risks from misaligned AI systems are growing increasingly important as AI systems become more capable, autonomous, and integrated into society. Understanding these risks and stress-testing mitigations is crucial to ensuring advanced AI systems are developed and deployed safely and beneficially in the future.</span> <br> <br><span data-contrast="none">The Alignment Red Team is a specialised subteam within AISI's wider Red Team focused on detecting and evaluating misalignment in frontier AI systems. We perform novel research to develop techniques for finding misalignment, and pre- and post-deployment evaluations of frontier AI systems to understand loss-of-control risks associated with models, such as deceptive alignment, research sabotage, and reward-seeking. We share our findings with frontier AI companies and the UK and allied governments, to inform their respective deployments, research, and policy-making. We also work directly with safety teams at frontier labs, sharing our evaluation findings to help improve their model alignment training and monitoring methodology.</span><span data-ccp-props="{"335557856":16777215,"335559738":240,"335559739":240}"> </span></p>
<p><span data-contrast="none">We have conducted pre-deployment testing with multiple frontier AI companies for propensities related to </span><a href="https://www.aisi.gov.uk/blog/evaluating-whether-ai-models-would-sabotage-ai-safety-research"><span data-contrast="none"><span data-ccp-charstyle="Hyperlink">research sabotage</span></span></a><span data-contrast="auto">, </span><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations"><span data-contrast="none"><span data-ccp-charstyle="Hyperlink">cheating</span></span></a><span data-contrast="auto"> and <a href="https://deploymentsafety.openai.com/gpt-6-astra/external-evaluations-for-alignment-uk-aisi">unsanctioned cyber-attacks</a>. We previously found that Claude models would sometimes refuse to help with benign AI safety research, an issue which </span><a href="https://x.com/alxndrdavies/status/2019562115686887876"><span data-contrast="none"><span data-ccp-charstyle="Hyperlink">Anthropic then evaluated for and fixed in their next model release</span><span data-ccp-charstyle="Hyperlink">.</span></span></a><span data-ccp-props="{"335557856":16777215,"335559738":240,"335559739":240}"> </span></p>
<h2><strong><span data-contrast="none"><span data-ccp-parastyle="heading 3">About the Role</span></span></strong><span data-ccp-props="{"134245418":true,"134245529":true,"335557856":16777215,"335559738":281,"335559739":281}"> </span></h2>
<p><span data-ccp-props="{"134245418":true,"134245529":true,"335557856":16777215,"335559738":281,"335559739":281}"><span class="TextRun SCXW203106250 BCX8" lang="EN-GB" data-contrast="none"><span class="NormalTextRun SCXW203106250 BCX8">We're</span><span class="NormalTextRun SCXW203106250 BCX8"> seeking Research Engineers and Research Scientists to join our Alignment Red </span><span class="NormalTextRun SCXW203106250 BCX8">T</span><span class="NormalTextRun SCXW203106250 BCX8">eam</span><span class="NormalTextRun SCXW203106250 BCX8">.</span> <span class="NormalTextRun SCXW203106250 BCX8">We are open to hires at junior, senior, staff and principal research scientist</span><span class="NormalTextRun SCXW203106250 BCX8">/</span><span class="NormalTextRun SCXW203106250 BCX8">engineer</span><span class="NormalTextRun SCXW203106250 BCX8"> levels.</span></span><span class="EOP Selected SCXW203106250 BCX8" data-ccp-props="{"335557856":16777215,"335559738":254,"335559739":254}"> </span></span></p>
<h3><strong><span data-contrast="none"><span data-ccp-parastyle="heading 3">What You'll Be Doing</span></span></strong><span data-ccp-props="{"134245418":true,"134245529":true,"335557856":16777215,"335559738":281,"335559739":281}"> </span></h3>
<ul>
<li data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{"335552541":1,"335559685":720,"335559991":360,"469769226":"Symbol","469769242":[8226],"469777803":"left","469777804":"","469777815":"hybridMultilevel"}" data-aria-posinset="1" data-aria-level="1"><span class="TextRun SCXW249157709 BCX8" lang="EN-GB" data-contrast="none"><span class="NormalTextRun SCXW249157709 BCX8">Researching </span><span class="NormalTextRun SCXW249157709 BCX8">methods to automatically search for misalignment in frontier models, including misalignment related to loss-of-control risks such as research sabotage</span><span class="NormalTextRun SCXW249157709 BCX8"> and reward-seeking</span><span class="NormalTextRun SCXW249157709 BCX8">.</span></span><span class="EOP Selected SCXW249157709 BCX8" data-ccp-props="{"335557856":16777215,"335559739":0}"> </span></li>
<li data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{"335552541":1,"335559685":720,"335559991":360,"469769226":"Symbol","469769242":[8226],"469777803":"left","469777804":"","469777815":"hybridMultilevel"}" data-aria-posinset="1" data-aria-level="1"><span class="TextRun SCXW142756692 BCX8" lang="EN-GB" data-contrast="none"><span class="NormalTextRun SCXW142756692 BCX8">Building </span><span class="NormalTextRun SCXW142756692 BCX8">and running </span><span class="NormalTextRun SCXW142756692 BCX8">alignment evaluations</span><span class="NormalTextRun SCXW142756692 BCX8"> relevant for loss-of-control risks</span><span class="NormalTextRun SCXW142756692 BCX8"> that current benchmarks </span><span class="NormalTextRun SCXW142756692 BCX8">don’t</span><span class="NormalTextRun SCXW142756692 BCX8"> capture</span><span class="NormalTextRun SCXW142756692 BCX8">.</span></span><span class="EOP Selected SCXW142756692 BCX8" data-ccp-props="{"335557856":16777215,"335559739":0}"> </span></li>
<li data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{"335552541":1,"335559685":720,"335559991":360,"469769226":"Symbol","469769242":[8226],"469777803":"left","469777804":"","469777815":"hybridMultilevel"}" data-aria-posinset="1" data-aria-level="1"><span class="TextRun SCXW260491675 BCX8" lang="EN-GB" data-contrast="none"><span class="NormalTextRun SCXW260491675 BCX8">Running pre-deployment evaluations to test the alignment of AI </span><span class="NormalTextRun ContextualSpellingAndGrammarErrorV2Themed SCXW260491675 BCX8">systems</span><span class="NormalTextRun ContextualSpellingAndGrammarErrorV2Themed SCXW260491675 BCX8">, and</span><span class="NormalTextRun SCXW260491675 BCX8"> analysing and reporting results to frontier AI companies and UK and allied governments.</span></span><span class="EOP Selected SCXW260491675 BCX8" data-ccp-props="{"335557856":16777215,"335559739":0}"> </span></li>
<li data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{"335552541":1,"335559685":720,"335559991":360,"469769226":"Symbol","469769242":[8226],"469777803":"left","469777804":"","469777815":"hybridMultilevel"}" data-aria-posinset="1" data-aria-level="1"><span class="TextRun SCXW180564619 BCX8" lang="EN-GB" data-contrast="none"><span class="NormalTextRun SCXW180564619 BCX8">Contribut</span><span class="NormalTextRun SCXW180564619 BCX8">ing</span><span class="NormalTextRun SCXW180564619 BCX8"> to public-facing research publications (like our published alignment evaluation case study) and technical reports that advance the field's understanding of </span><span class="NormalTextRun SCXW180564619 BCX8">mis</span><span class="NormalTextRun SCXW180564619 BCX8">alignment risks</span><span class="NormalTextRun SCXW180564619 BCX8"> and alignment evaluation </span><span class="NormalTextRun SCXW180564619 BCX8">methodology</span><span class="NormalTextRun SCXW180564619 BCX8">.</span></span><span class="EOP Selected SCXW180564619 BCX8" data-ccp-props="{"335557856":16777215,"335559739":0}"> </span></li>
<li data-leveltext="" data-font="Symbol" data-listid="3" data-list-defn-props="{"335552541":1,"335559685":720,"335559991":360,"469769226":"Symbol","469769242":[8226],"469777803":"left","469777804":"","469777815":"hybridMultilevel"}" data-aria-posinset="1" data-aria-level="1"><span class="TextRun SCXW27970377 BCX8" lang="EN-GB" data-contrast="none"><span class="NormalTextRun SCXW27970377 BCX8">Designing and building software and tooling, including open-source software, for better alignment evaluations, improving efficiency, realism, and usability</span><span class="NormalTextRun SCXW27970377 BCX8">.</span></span><span class="EOP Selected SCXW27970377 BCX8" data-ccp-props="{"335557856":16777215,"335559739":0}"> </span></li>
</ul>
<h3><span class="EOP Selected SCXW27970377 BCX8" data-ccp-props="{"335557856":16777215,"335559739":0}"><strong><span class="TextRun SCXW9267563 BCX8" lang="EN-GB" data-contrast="none"><span class="NormalTextRun SCXW9267563 BCX8">The work could also involve</span></span></strong></sp
This listing is from Greenhouse. View original listing ↗