Research Overview

I currently focus on technical and international governance of frontier AI, particularly AI verification and Sino-Western cooperation. I care most about reducing risks of misuse, loss-of-control, and concentration of power.

In the past, I've explored a variety of other research questionsOversight and control. What oversight measures robustly scale to increasingly capable frontier language models? How do we understand situational awareness, model introspection, and other mentalistic properties?Value alignment and epistemics. How can we train models to be more honest? How can AI uplift human truth-seeking and moral progress, and how do we prevent harmful value lock-in?Agent security. How do we design scalable, realistic environments for evaluating agent misuse and misbehavior — for example, in computer use and MCP settings?How do we operationalize AI-induced human disempowerment? in technical AI safety and alignment.

I am most driven by impact — helping us ‘win’. As a doer and generalist, I try to make things happen in the world strategically. As a researcher, I enjoy rapid experimentation and careful truth-seeking.

It helps, too, that I absolutely love my work — I'm very lucky :)

Papers

AI Epistemic Risks: Emerging Mechanisms & Evidence

Casper, S., Stray, J., Gausen, A., Jones, C., Christian, B., Li, J., Bengio, Y., Rand, D., et al.

SSRN, 2026

[SSRN]

WARP: Measuring and Mitigating Evaluation Awareness in Browser-Agent Safety Benchmarks

Li, J.X., Chew, A., Lin, M., Jones, E.K., Fu, X., Zou, A.

ICML AIWILD Workshop, 2026

[OpenReview] [GitHub]

EigenBench: A Comparative Behavioral Measure of Value Alignment

Chang, J., Piff, L., Sana, S., Li, J.X., Levine, L.

ICLR Oral (Top 5%), 2026

[arXiv]

ProgressGym: Alignment with a Millennium of Moral Progress

Qiu, T., Zhang, Y., Huang, Z., Li, J.X., et al.

NeurIPS Spotlight (Top 10%), 2024

[arXiv]

Organizing

Code

RituRitu

Smart pest and weather prediction for farmers. Grand Prize, Cornell Switch the Pitch Hackathon

ALIGNALIGN

Designed optimized database and search for Wex, Cornell Legal Information Institute's dictionary. First Prize, LII Hackathon

CirclesCircles

Frictionless friend meetups. Big Red Hacks 2024