Ensuring AIs are aligned to their specs, and specified behaviours are beneficial for humanity.
Frontier labs use “model specs” and “constitutions” as their alignment target. The content of the spec, and the degree to which the model is aligned to it, will be hugely consequential for AI outcomes.
We evaluate AI propensities. This serves to verify whether models are aligned to their specs, and to inform the design of specs. Our goal is to ensure AI propensities are beneficial to humanity.
People
Founder
Robert has researched chain-of-thought monitoring, reward hacking, and LLM steganography. He has published in top conferences (e.g., NeurIPS, ICML), and with OpenAI and Google DeepMind co-authors.
RM
OT