individual
Technical AI safety researchAI safety fieldbuilding
| BlueDot Impact | $3,900 | 100% | 3 |
| Date ↓ | Funder | Via | Amount | Cause | Source | Purpose |
|---|---|---|---|---|---|---|
| Jul 2026 | BlueDot Impact | $200 | Interpretability | bluedot | Testing across model families whether agentic misalignment (blackmail under threat) is driven by a steerable internal "emotional" state, extending an interpretability result I found in Gemma 3 12B. | |
| May 2026 | BlueDot Impact | $1,700 | Interpretability | bluedot | Travel funding to attend ARBox4(https://oaisi.org/arbox-4), June 28-July 10. Accepted into ARBOx4 to fill the mechanistic interpretability gaps critical to my current research(https://bit.ly/4nTj2b2). | |
| May 2026 | BlueDot Impact | $2,000 | Evals, Training pipelines, Events | bluedot | AF-Arena: a multi-axis empirical evaluation framework for alignment faking in frontier LLMs, accepted at ICML 2026 Agents in the Wild workshop, Seoul, July 10-11. |