Chloe Li

I spend my time thinking about how to make powerful AI systems honest and aligned with human values. I’m currently a member of technical staff at Anthropic, where I work on alignment training.

Previously, I worked on AI safety research like model spec midtraining for aligning models, honesty training, chain-of-thought monitoring and control evaluations. I also spent time on AI safety field-building, like leading ARENA (a ML engineering program for upskilling people in technical AI safety) and directing Cambridge AI Safety Hub, where I helped founded research programs like MARS.

Publications & other work