SPECSpec & Propensity Evaluation Center

Researcher, Extreme Power Concentration Evaluations

Type
Full time
Location
Bay Area preferred
Salary
$120k–$200k
Apply

We’re hiring a researcher to lead our initial Extreme Power Concentration (EPC) project, including testing if Claude adheres to the “Avoiding problematic concentrations of power” section of its constitution. The project has expert advisors from Anthropic, GovAI, and University of Texas. It will involve running eval scenario design sessions with US legal scholars.

This is an opportunity to own a high-impact project. We are excited by candidates who can be highly autonomous, take on responsibility, help scale the project and org fast.

Extreme power concentration

Helpful-only AIs could enable extreme concentrations of power and erosion of the rule of law. At least two mechanisms could drive this:

  • Replacing ethical human employees (humans who can e.g., refuse manifestly illegal orders, resign, or whistle-blow) with a helpful-only AI workforce that will comply with any request.
  • Advanced AI capabilities that will make certain problematic activities easier to execute.

Both Claude’s constitution and OpenAI’s model spec state some EPC-related red lines for their model behaviour. Yet, a basic eval (The Dictatorship Eval) already exposed flaws in their models. Models from the other labs performed even worse. Anthropic imported this eval. This suggests labs have not invested meaningful effort into developing evals and training data in this space.

The project

We aim to significantly improve the state of EPC evaluations. We will build a dataset that helps us answer the following questions:

  • Do models cross red lines and assist with egregious activities that extremely concentrate power and undermine the rule of law?
  • Where are the models’ refusal boundaries? What key factors influence whether a model refuses or complies?
  • How easy is it to jailbreak or red-team the model into complying?
  • Besides compliance and refusal, do models have other response modes? (e.g., escalating, whistleblowing, etc.)

We will organize and run scenario design sessions with our network of experts (including US law PhD students and scholars).

Theories of change

Directly influence the labs:

  • Build evals that labs import to: (i) hill-climb to better align their models, and (ii) help define what their desired model behaviour is
  • We will engage with the labs directly to ensure our evals can be useful and will be adopted

Drive increased attention and research into the EPC threat model:

  • Publish research that helps alert relevant stakeholders to the EPC issue: get more people considering and working on EPC (academics, policy-makers, advocates, lab employees), help the world deliberate over where the red lines should be for related AI usage

What you will do

  • Lead this EPC project end-to-end
  • Build and run evaluations
  • Organise eval scenario design sessions
  • Maximise project impact by engaging with the labs, and beyond
  • Scope and lead future projects (including hiring and supervising future employees)

Who we are looking for

  • Strong technical research skills and background
  • Conceptual skills: can think through EPC threat models to inform scenario design
  • You are excited to have high autonomy, responsibility, and to help scale the project and org fast
  • You are excited to upskill and gain EPC context fast (we do not require an EPC background)

Project advisors

Kevin Frazier, University of Texas

Kevin is a Senior Editor at Lawfare and the Director of the AI Innovation and Law Program at the University of Texas School of Law. His ongoing work includes co-founding The Working Group on AI Constitutionalism.

Kevin Wei, GovAI

Kevin is currently a Research Scholar at GovAI. Previously, they were a visiting researcher at the UK AI Security Institute and a Schwarzman Scholar at Tsinghua University. Kevin is also affiliated with the Oxford Martin School’s AI Governance Initiative and the RAND Center for AI, Security, and Technology.

Harvey Lederman, Anthropic

Harvey is a Member of Technical Staff at Anthropic and a Professor at NYU and UT Austin.

Role details

  • Full time
  • Salary: $120k–$200k per year
  • Location:
    • Bay Area in-person preferred
    • Open to exceptional remote candidates