individual
AI safetyTechnical AI safety researchAI safety fieldbuilding
| BlueDot Impact | $1,600 | 100% | 2 |
| Date ↓ | Funder | Via | Amount | Cause | Source | Purpose |
|---|---|---|---|---|---|---|
| Aug 2026 | BlueDot Impact | $1,000 | Interpretability, Events | bluedot | Presenting two accepted COLM 2026 workshop papers: reward hacking / goal misgeneralization (AIMS), and interpretability-guided intervention to cut VLM hallucinations (AIW). | |
| May 2026 | BlueDot Impact | $600 | AI safety | bluedot | Building RH-SusBench: an open benchmark for reward-hacking susceptibility in language models, measuring shortcut collapse, reasoning–answer decoupling, and recovery across model families. |