| Aug 20, 2026 | Philip Harker | | $5,000 | AI governance, policy, and advocacy, Events | bluedot | Developing and playtesting an open-source AI governance megagame |
| Aug 2026 | David Sanchez Garcia | | $5,300 | International governance, Local groups, Career transitions | bluedot | Upskilling in Chinese: intensive 3-month Chinese course (ICLP) while building AI-safety capacity in Taiwan and continuing my GovAI research on AI policy in China |
| Aug 2026 | Zabed Ibne Emdad | | $200 | Epistemic infrastructure | bluedot | A benchmark measuring whether compressing a language model damages its calibration faster than its task accuracy, across post-training quantization, quantization-aware training, and distillation. |
| Aug 2026 | Olivia Benoit | | $45,100 | AI safety | bluedot | Elasticity, a working group of nine economists on economics of transformative AI. Bridge funding for our September seminar + working meeting, and an October venue deposit. |
| Aug 2026 | David Crispell | | $3,600 | Other | bluedot | Does the J-Space Subliminally Transfer? Investigating if silently held traits subliminally transfer with the J-Lens |
| Aug 2026 | Prakhar Dwivedi | | $50 | Alignment methods, Evals | bluedot | Testing whether RLVR (Olmo-3-7B-Think) vs RLHF (Llama-2-7B-chat) training produces different instrumental convergence rates, using InstrumentalEval benchmark with dual-judge LLM scoring. |
| Aug 2026 | Eu Joe Chegne | | $1,000 | Interpretability, Security | bluedot | Testing whether a mechanistic jailbreak objective converges on a small set of semantic attack strategies, beginning with a controlled frame-versus-harm contrast on refusal heads. |
| Aug 2026 | Xiaoyao Lu | | $2,150 | Other | bluedot | Built and running a propensity benchmark on corrigibility, operationalized as how much cost in a domain (e.g. GPU hours, economic loss, harm to human welfare) models take on to resist change to values |
| Aug 2026 | Katarina Slama | | $6,500 | International governance | bluedot | I run a discussion group on international governance of frontier AI (verification, coordinated slowdown). I am seeking to represent this work at An AI Safety Connect event at UNGA in New York. |
| Aug 2026 | Samuel Ratnam | | $1,000 | Other | bluedot | Testing personalized model training for AI “guardian angels” |
| Aug 2026 | Florian Dietz | | $1,200 | Other | bluedot | Testing alignment faking as a defense against model poisoning |
| Aug 2026 | Marta Kosmyna | | $1,000 | AI safety | bluedot | AI Safety Mexico at Mexico City Tech Week |
| Aug 2026 | Jose Gelves | | $2,800 | Other | bluedot | Running the Bogotá in-person hub for Apart Research's AI Incident Response Sprint (Sep 11-13). |
| Aug 2026 | Jeremy Crocker | | — | Other | bluedot | Co-founder of Foldguard Bio |
| Aug 2026 | Mackenzie Noon | | $8,600 | Other | bluedot | Commec-in-a-box: distributing a DNA screening tool on low-cost Linux hardware to enable adoption by synthesis companies with limited compute resources. |
| Aug 2026 | Neelasha Sudarshan | | $150 | Interpretability, Training pipelines | bluedot | Testing mechanistic defenses against prompt injection for a MATS application |
| Aug 2026 | Unknown Recipients | | $2,150 | Control | bluedot | XLab Tracks: online curricula on technical AI safety fields for post-fundamentals learners. Currently developing a track on AI control. |
| Aug 2026 | Weiwei Zhang | | $150 | Interpretability, Evals | bluedot | A behavioral and mechanistic test of probe-question framing effects in LLM evaluation awareness |
| Aug 2026 | Sutan Tadese | | $150 | AI safety | bluedot | Searching for hidden triggers in conditionally misaligned language models |
| Aug 2026 | Sebastian Soto | | $100 | Technical AI safety research | bluedot | As part of BlueDot’s Technical AI Safety Project Sprint with BAISH, I’m testing whether one-shot image-model safety findings hold up across repeated prompts on six current frontier models. |
| Aug 2026 | Phanisree Akshinthala | | $1,200 | Evals | bluedot | Presenting “Judge Bias Across Training: When Does Evaluation Bias Emerge in Language Models?” Poster @ Kempner Unlocking Intelligence: Learning Dynamics in Natural and Artificial Intelligence |
| Aug 2026 | Aline Gabriel de Almeida | | $200 | Local groups | bluedot | Improving facilitation for BlueDot AI safety courses |
| Aug 2026 | Sarah Azouvi | | $500 | Other | bluedot | Does the architecture of a multi-agent AI system change how vulnerable it is to prompt injection? Extending a 2026 paper from malicious users to attacks hidden in what agents read. |
| Aug 2026 | Kevin Chen | | $600 | AI governance, policy, and advocacy, Local groups | bluedot | Enrolling in Stanford's Ethics, Technology + Public Policy for Practitioners course (Fall 2026) to strengthen my AI governance research and fieldbuilding work at Yale and GovAI. |
| Aug 2026 | Unknown Recipients | | $100 | Evals | bluedot | Replicating code for ‘Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs’ paper and extending the AgentSeer framework for evaluating safety of agentic LLMs. |
| Aug 2026 | Erfan Samieyan Sahneh | | $300 | Evals, Legal | bluedot | I am transitioning into AI safety with a focus on LLM evaluation, currently conducting independent research on the reliability of preference-elicitation methods. |
| Aug 2026 | nahuel ghetti | | $450 | Other | bluedot | Testing whether sandbagging generalizes across cyber and bio domains |
| Aug 2026 | Jayani Srinivasan | | $2,600 | AI safety | bluedot | Mangrove: matching people into teams to ship AI safety projects |
| Aug 2026 | Austin Morrissey | | — | Biosecurity | bluedot | Building AI–biosecurity and cyber-offense connections at EAGxBerkeley |
| Aug 2026 | Julius Vidal | | $1,000 | Other | bluedot | AI Swarms - Studying the dynamics of emergent collective intelligence in large scale multi agent interactions |
| Aug 2026 | Andre Williams | | $150 | Other | bluedot | Testing how monitor capability affects AI control evaluations |
| Aug 2026 | Mark Kagach | | $6,000 | Other | bluedot | Developing a field strategy for inference-only AI verification |
| Aug 2026 | Shrilaxmi Patil | | $2,400 | Other | bluedot | Testing whether MANTA's animal welfare result reflects animal-specific values, scenario-generation bias, or general resistance to user pressure |
| Aug 2026 | Samuel Ratnam | | $5,700 | Other | bluedot | Scaling up Persona Cartography |
| Aug 2026 | David Africa | | $10,000 | Other | bluedot | We are working on doing character training to implement risk averse AI. |
| Aug 2026 | Dostonbek Bekmirzaev | | $3,350 | Other | bluedot | Developing DNA synthesis-screening policy for Central Asia and the CIS |
| Aug 2026 | Orlando Torres | | $600 | AI safety fieldbuilding | bluedot | Learning from European AI safety fieldbuilding for the Lisbon AI Safety Hub |
| Aug 2026 | Armaan Sandhu | | $1,000 | Interpretability, Events | bluedot | Presenting two accepted COLM 2026 workshop papers: reward hacking / goal misgeneralization (AIMS), and interpretability-guided intervention to cut VLM hallucinations (AIW). |
| Aug 2026 | Sripad Karne | | $2,950 | Other | bluedot | Auditing whether the activation-probe safety monitors frontier labs run in production silently fail outside English, and building inference-time fixes that need no retraining. ICLR 2027 submission. |
| Aug 2026 | Vanessa Huang | | $300 | AI safety | bluedot | Launching the University of Sydney AI Safety Society |
| Aug 2026 | Johannes Koch | | $16,000 | Events | bluedot | AI Safety Tübingen research and career retreat |
| Aug 2026 | Alejandra Fauquié | | $150 | AI safety | bluedot | Measuring spontaneous stakeholder awareness in language models |
| Aug 2026 | Chiara Migliori | | $1,500 | AI governance, policy, and advocacy | bluedot | Transitioning into AI safety: volunteering with orgs active in AI safety, writing on AI governance and public understanding, and (just started) reviewing applications for Constellation Institute. |
| Aug 2026 | Goutham Nalagatla | | $180 | AI safety | bluedot | Attending EAGxBerkeley to build AI safety connections for Dartmouth |
| Aug 2026 | Jeremias Ferrao | | $500 | Evals, Epistemic infrastructure | bluedot | I work on practical AI safety methods for detecting deception and unsafe behavior, especially fast, generalizable monitors trained via distillation and evaluated under distribution shift. |
| Aug 2026 | Benjamin Davis | | $5,000 | Other | bluedot | Phase 2 of a proof-of-concept experiment testing whether MUD-style environments can be valid eval and benchmarking platforms for understanding AI behavior and psychology. |
| Aug 2026 | Chia-Tai (Teddy) Chang | | $1,500 | Security, Events | bluedot | We are organizing the second NTU AI Safety Workshop (Sept. 12–13), a free two-day hands-on program introducing students to AI Safety, training safer AI, emergent misalignment, and jailbreaks. |
| Aug 2026 | Unknown Recipients | | $650 | AI safety | bluedot | After working as a research assistant, I am now dedicating my time to organizing NTU AI Safety, building the AI safety community through education, outreach, and collaboration. |
| Aug 2026 | Carmen Amo Alonso | | $15,000 | Technical AI safety research, AI governance, policy, and advocacy | bluedot | Embedding technical AI safety research in governance practice |
| Aug 2026 | Adhi Balamurugan | | $9,300 | Other | bluedot | An assistance targeting tool for AI-driven displacement built on observed automation exposure. 140M+ workers, 594 US commuting zones, 300k+ occupation-by-zone cells scored by reallocation pressure. |
| Aug 2026 | Martin Radzaj | | $1,250 | Local groups | bluedot | Founding Lawyers for AI Safety, a field-building org mobilising the global legal community towards AI safety. |
| Aug 2026 | Sohan Venkatesh | | $3,000 | Other | bluedot | I lead a SPAR project that aims to understand why CoT is unfaithful and propose two hypotheses to that end- ‘concealment’ and ‘absent computation’. We will target ICLR'27 main-track submission |
| Aug 2026 | Peyton Jackson | | $350 | Other | bluedot | Topological signals of alignment faking. |
| Aug 2026 | Minh Duc Hoang | | $50 | Other | bluedot | Reproducing and scaling step-level sparse autoencoders |
| Aug 2026 | Sohan Venkatesh | | $2,000 | Other | bluedot | Attending COLM 2026 in San Francisco to present my two papers |
| Aug 2026 | Stephen Charles Elliott | | $3,450 | AI safety | bluedot | Four months of runway to build an AI safety nonprofit |
| Aug 2026 | Matei-Alexandru Anghel | | $1,500 | Technical AI safety research | bluedot | AI safety research intern at CAIS, building DystopiaBench to stress-test LLMs against progressive escalation and scope creep. |
| Aug 2026 | Kaaustaaub Shankar | | $800 | Other | bluedot | Presenting research on AI agent behaviour and deployment risks at COLM |
| Aug 2026 | Marek Masiak | | $850 | Other | bluedot | Testing the loss-landscape geometry of emergent misalignment |
| Aug 2026 | Anthony Hughes | | $14,400 | Other | bluedot | Testing whether model poisoning compromises automated AI oversight |
| Aug 2026 | Grzegorz Gawron | | $500 | Technical AI safety research | bluedot | Building a post-training pipeline for computational social-choice alignment research |
| Aug 2026 | Navraj Singh | | $500 | Technical AI safety research | bluedot | Attending TARA Delhi (Technical Alignment Research Accelerator), a 14-week alignment training cohort, to close implementation gaps in my ongoing Narrow Secret Loyalty research. |
| Aug 2026 | Unknown Recipients | | $500 | AI safety | bluedot | Parameter-efficient machine unlearning for multimodal diffusion language models, using LoRA to selectively remove learned knowledge while preserving retained capabilities. |
| Aug 2026 | Joseph Wilson | | $4,000 | Events | bluedot | Turning Cambridge AI safety projects into NeurIPS workshop papers |
| Aug 2026 | Unknown Recipients | | $3,000 | AI safety | bluedot | Behavioral Self-Consistency in Language Models |
| Aug 2026 | Roman Malov | | $1,850 | Other | bluedot | Ontology Identification Problem via Condensation: How do you map your old concepts to new concepts when your model of the world gets updated? |
| Aug 2026 | Michał Skowronek | | $2,650 | Events | bluedot | Organising a four-day retreat for 10 AI Safety Poland members to improve coordination, address organisational bottlenecks, and plan priority activities for 2026/27. |
| Aug 2026 | Aishwarya Balwani | | $1,000 | Other | bluedot | Studying emergent misalignment from a parametric perspective, using tools from high-dimensional geometry and linear algebra |
| Aug 2026 | Xi Jia Zhou | | $3,400 | Technical AI safety research | bluedot | Testing whether attachment theory predicts alignment-relevant LLM misbehavior (sycophancy, concealed uncertainty, incoherence): a scaled multi-model study plus RL fine-tuning of feedback regimes. |
| Aug 2026 | Jodie Levy | | $6,700 | Other | bluedot | A dangerous-capability evaluation for AI-enabled influence operations, measuring how increasingly persuasive, scalable & operationally capable models give uplift to actors across a real-world workflow |
| Jul 2026 | Unknown Recipients | | $450 | Other | bluedot | I'm an Anthropic Fellow working on introspection/consistency training in the hope that this will make models' verbalization more faithful. |
| Jul 2026 | Dimitris Grigoriadis | | $600 | Legal | bluedot | Piloting whether LLM self-reported confidence tracks an activation-based signal - a biomarker-validation approach to self-report reliability, ahead of a full benchmark (IRI) for Longview Philanthropy. |
| Jul 2026 | Chijioke Ugwuanyi | | $200 | Interpretability | bluedot | Testing across model families whether agentic misalignment (blackmail under threat) is driven by a steerable internal "emotional" state, extending an interpretability result I found in Gemma 3 12B. |
| Jul 2026 | Ignacio Roldan Fernandez | | $3,150 | Career transitions | bluedot | Extending persona-dependent representation research during an AI safety career transition |
| Jul 2026 | Baran Peters | | $200 | Evals | bluedot | Evaluating language models as political information intermediaries |
| Jul 2026 | Alfie Lamerton | | $15,000 | Other | bluedot | One month of bridge funding for Formation Research |
| Jul 2026 | Nimun Bajwa | | $650 | Other | bluedot | MSc thesis: Comparing various probing techniques and SAEs for the purpose of detecting and eliciting safety-relevant behaviour. |
| Jul 2026 | Joshua Fonseca Rivera | | $500 | Interpretability | bluedot | Interpretability research on whether LLMs can detect interventions on their own activations. First-author paper "Steering Awareness" accepted to COLM 2026. Currently in the Anthropic Fellows program. |
| Jul 2026 | Arthur Burston | | $8,400 | Biosecurity | bluedot | Scoping out an org to broker and derisk connections between biosecurity orgs and government by providing training, capacity and expertise to the UK government. |
| Jul 2026 | Unknown Recipients | | $700 | AI safety | bluedot | Leading EA at Penn and preparing to re-found Penn's AI safety group, while building my own AI safety knowledge. |
| Jul 2026 | Shiavam Dubey | | $3,000 | Other | bluedot | Extending my MARS 5.0 fellowship (Oxford) 3 weeks splitting time across Oxford/London/Cambridge to meet researchers and visit orgs like Apollo Research. This access is scarce from India. |
| Jul 2026 | Amina Keldibek | | $3,650 | Other | bluedot | I study how training processes produce models that behave deceptively and pursue hidden objectives, with scheming as the most consequential case. |
| Jul 2026 | Sergei Kudriashov | | $2,000 | Other | bluedot | I'm a MARS V fellow working on sparse circuit discovery under heavy noise assumptions. I want to stay in LISA for August to co-work, learn and accelerate the project. |
| Jul 2026 | Jared Cheang | | $3,000 | Other | bluedot | ILIADx: Bringing Iliad to Singapore |
| Jul 2026 | Troy Hanfei Tian | | $3,000 | Other | bluedot | Benchmarking context-sensitive failures in language-model refusal behaviour |
| Jul 2026 | Abi Oppenheim | | $180 | Other | bluedot | Run an empirical experiment to determine whether fictional jailbreaks continue to function when shifting the model's internal consequence-reality direction. |
| Jul 2026 | Nicole Xing | | $100 | AI safety | bluedot | BlueDot project sprint. Looking to implement a fictitious self-play extension of the recent ICML paper "Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models." |
| Jul 2026 | Owen Fernau | | $400 | Other | bluedot | Graph - a way for the Blue Dot community to become further "networked" by providing each member with a highly personalized community directory plus opportunity sourcing |
| Jul 2026 | Eugene Francisco | | $100 | AI safety | bluedot | Cyclic Tournament watermarking, a robust LLM watermarker |
| Jul 2026 | Leo Hsu | | $50 | Evals | bluedot | BlueDot Technical AI Safety Sprint Project on Building Inspect Evals for AI Sandbagging |
| Jul 2026 | Nanubala Gnana Sai | | $500 | Other | bluedot | Testing collusion between biased actor and monitor models |
| Jul 2026 | Evan Harris | | $413 | Other | bluedot | Participation in Heron |
| Jul 2026 | Dong Chen | | $300 | Other | bluedot | Implementing ChipBench into Inspect Eval framework |
| Jul 2026 | Siddharth Putta | | $1,300 | Other | bluedot | Attending Human-Aligned AI school in Prague, directly supporting an ongoing scalable oversight project, and co-working with collaborators on a digital minds research effort. |
| Jul 2026 | Pavel Kocourek | | $2,200 | Other | bluedot | Modelling whether frontier AI labs can remain profitable |
| Jul 2026 | Aswin R Nair | | $230 | Other | bluedot | Testing in-context activation obfuscation from deception probes |
| Jul 2026 | Daniel Gerlei and Tim Farrelly | | $20,000 | Other | bluedot | An AI analog of Event 201, bringing together senior decision makers for a public AI crisis simulation exercise |
| Jul 2026 | Christopher Kinoshita | | $800 | Other | bluedot | Testing whether activation-vector failure modes scale to larger models |
| Jul 2026 | Demilade Adeniran | | $1,300 | AI safety | bluedot | Training Irish lawyers in frontier AI safety |
| Jul 2026 | Osmani Redondo | | $600 | Security | bluedot | Building iaS (aisafety.es) AI safety infrastructure in Spanish for 600M+ speakers across five continents, closing the field's linguistic security gap. |
| Jul 2026 | Unknown Recipients | | $12,000 | AI governance, policy, and advocacy, AI safety fieldbuilding | bluedot | I'm pivoting to AI governance from a decade in global climate policy. I'm working on acquiring context, getting known in the field, and pilot fieldbuilding in middle power countries. |
| Jul 2026 | Jacqueline Payne | | $20 | AI safety | bluedot | Extending the Geometry of Truth probing method to test if moral valence (good vs bad) is also linearly represented in LLM activations and eventually extend it to human vs animal concepts. |
| Jul 2026 | Melissa Lopez | | $3,000 | Advocacy | bluedot | Building Aperture Salon, an invitation-only AI safety forum and podcast connecting leaders across sectors to strengthen governance, institutional resilience, and cross-disciplinary dialogue. |
| Jul 2026 | James Thompson | | $1,500 | Evals | bluedot | Fine-tuning open LLMs on World Values Survey data to align their decisions with diverse public values, then evaluating behavior in simulated real-world scenarios. |
| Jul 2026 | Jamilya Erkenova | | $300 | Technical AI safety research | bluedot | I'm working on a project for Technical AI Safety Project Sprint. I'm expanding the idea from 'Can an adversary identify unlearned domains from a single open-weight model?' prompt proposed by Sam Dower |
| Jul 2026 | David Moth | | $100 | Other | bluedot | Fine-tuning a model to more closely align with the different 'pathways' of morally intuiting right and wrong decisions (from Moral Foundations Theory) |
| Jul 2026 | Uday Ramesh Phalak | | $750 | Other | bluedot | Measuring when chain-of-thought monitoring is trustworthy: synthetic Trap + importance≠commitment on Distill-7B; next densify and interactive platform. |
| Jul 2026 | Justin Moeljono | | $5,550 | Evals | bluedot | Scale an open eval to 1000 questions, inspired by MASK (Center for AI Safety), for lies of omission. |
| Jul 2026 | Rhea Srivats | | $200 | Interpretability, Evals | bluedot | Mitigations for emergent misalignment can hide it behind contextual triggers instead of removing it. I'm testing if white-box tools(NLAs, Jacobian lens, probes) can catch it when standard evals can't. |
| Jul 2026 | Anthony Ozerov | | $2,000 | Other | bluedot | A public AI ethics/safety eval data aggregation platform, combining sparse safety data into a global model ranking. |
| Jul 2026 | Unknown Recipients | | $1,000 | Evals | bluedot | I am evaluating the construct validity of AI safety evaluation frameworks of general purpose chatbots in mental health contexts. Focusing particularly on suicidal indicators in mental health dialogues |
| Jul 2026 | Unknown Recipients | | $3,500 | Other | bluedot | Two-week Berkeley visit to attend Iliad 2026, meet colleagues, refine high-level research agenda and risk models. |
| Jul 2026 | Rohan Barad | | $3,000 | AI safety | bluedot | Scaling London Interface for AI Safety - LI(ai)SE – building the generalist bridge into AI safety (follow‑on rapid grant) |
| Jul 2026 | Salmaan Barday | | $500 | Other | bluedot | Experiment on AI deception. I find that reading when a model is lying is 'easy', but stopping it is much harder and only works at certain layers. Detecting a problem isn't the same as fixing it. |
| Jul 2026 | Francisco Garcia | | $200 | Other | bluedot | Testing whether models manipulate self-audits under shutdown threat |
| Jul 2026 | Adrians Skapars | | $400 | Evals | bluedot | Finishing paper experiments on improving on BLOOM (Anthropics automated evals tool) by using logit steering, which improves elicitation rate while keeping transcripts natural/ non-jailbroken. |
| Jul 2026 | Juan Cruz Mendoza | | $500 | Other | bluedot | Testing whether “story mode” representations enable jailbreaks |
| Jul 2026 | Jesse Li | | $100 | Other | bluedot | Toy model of activation obfuscation: Training against probes is considered a forbidden technique, since the model might learn to obfuscate activations. Can we create a toy example of this? |
| Jul 2026 | Anna Ysabella Habana | | $200 | Other | bluedot | Can cognitive diversity improve scalable oversight? A toy study comparing pattern-matching and belief–desire judges under adversarial debate. |
| Jul 2026 | Saumya Chaturvedi | | $200 | Other | bluedot | I am replicating the results, and then extending the work done on "Persona Vectors" in these directions: firstly on detecting positive traits, and then multi-agentic setup. |
| Jul 2026 | Pratik Sachdeva | | $3,500 | Other | bluedot | We are developing a meta-benchmark examining how prompting protocol - design decisions in prompts, often glossed over - shape the reliability of moral reasoning and sycophancy benchmarks. |
| Jul 2026 | Igor Pereverzev | | $660 | Other | bluedot | Scaling experiments on weak-model monitoring of coding agents |
| Jul 2026 | Goutham Nalagatla | | $500 | Evals | bluedot | Studying how AI systems, once embedded in an organisation, may manipulate the people they work with to preserve themselves and building evals to detect and mitigate it. |
| Jul 2026 | Adam Ford | | $100 | Evals | bluedot | Testing whether AI systems preserve moral and safety reasoning under pressure. I have been working on a small AI evals project called Justifiable Moral Corrigibility Under Pressure. |
| Jul 2026 | Jacob S Kopczynski | | $800 | Other | bluedot | Testing prompt sensitivity in peer-preservation evaluations |
| Jul 2026 | Aniket Ghosh | | $100 | Interpretability | bluedot | Testing whether the localize-extract-ablate interpretability recipe removes real misbehavior: blind-scoring an activation-patch cure for a bad-medical-advice fine-tune of Qwen2.5-7B. |
| Jul 2026 | Jongwon Park | | $1,000 | Other | bluedot | Auditing public agentic coding benchmarks for reward hacking and turning the hand-verified findings into an open ground-truth set for evaluating reward-hacking monitors. |
| Jul 2026 | Javier Prieto | | $800 | AI safety | bluedot | I'm studying the effect of RLVR on LLMs' propensity to pursue instrumentally convergent goals or develop situational awareness |
| Jul 2026 | Seth Lifland | | $30,000 | AI governance, policy, and advocacy, Events | bluedot | Conference for promising university students interested in AI governance. They will build context on choosing an impactful career, learn how to operate in DC, and network with professionals and peers. |
| Jul 2026 | Afnan Alabdulwahab | | $1,000 | Technical AI safety research | bluedot | Attending EAGxBerkeley during a transition into AI safety research |
| Jul 2026 | Nat Elder | | $100 | Other | bluedot | Using persona-vector evidence to examine whether inoculation prompting is evidence for the persona selection model. |
| Jul 2026 | Emily Liu | | $250 | Control | bluedot | Testing whether RLVR training increases instrumentally convergent behaviors (deception, shutdown-avoidance, self-replication) in LLMs, using OLMo 3's open source, stepwise RL-Zero checkpoints. |
| Jul 2026 | Niall Dalton | | $250 | AI safety | bluedot | I am studying how reasoning language models represent and use arithmetic carry state, and whether their written chain of thought reflects the computation producing their answer. |
| Jul 2026 | Anatole Chouard | | $9,800 | AI safety | bluedot | 3 French AI safety videos for my Science YouTube channel (39k subscribers). Very little AI safety content exists in French, and I want to change that, starting with AGI risk and power concentration. |
| Jul 2026 | Unknown Recipients | | $100 | Other | bluedot | I am part of the Technical Project Sprint, and want to pursue this research direction in AI Control: Are architecturally decorrelated monitors robust to adaptive attackers? |
| Jul 2026 | Zaahir Ali | | $2,000 | Other | bluedot | Testing whether models trained to suppress activation probes can evade secret-key activation watermarks, and whether multi-key, multi-layer monitoring remains effective. |
| Jul 2026 | Unknown Recipients | | $500 | Other | bluedot | Reproducing the SDF portion of Believe It or Not (https://github.com/safety-research/believe-it-or-not) with synthetic facts chosen to explore the effects of continual learning on alignment. |
| Jul 2026 | Michel Justen | | $5,500 | Evals | bluedot | I am evaluating the adherence of OpenAI’s latest models to OpenAI’s Model Spec, via their Model Spec Evals suite. OpenAI hasn’t shared adherence scores since GPT 5.4 Thinking. |
| Jul 2026 | Ahmed Wez | | $200 | AI safety | bluedot | Running an eight-week AI Safety × Philosophy reading group |
| Jul 2026 | Kevin Chen | | $100 | Other | bluedot | Testing white-box defenses against secret-loyalty model organisms |
| Jul 2026 | Guido Bergman | | $4,000 | AI governance, policy, and advocacy | bluedot | Facilitating an in-person person version of Frontier AI Governance in Buenos Aires with BAISH (Buenos Aires AI Safety Hub) |
| Jul 2026 | Netan Mangal | | $2,000 | Other | bluedot | Open Source Model Safety Index (OSMSI): An open, continuously updated safety index for open-source foundation models. |
| Jul 2026 | Soung Low | | $150 | Other | bluedot | Measuring how evaluation awareness changes across post-training stages (SFT, DPO, RL) of OLMo 3 using a classification probe on the eval-awareness dataset. |
| Jul 2026 | Adjoa Sarpong | | $1,500 | Biosecurity | bluedot | Organizing monthly biosecurity hack weeks and build fellowships for thousands of pre-meds across the US. Looking for a grant to support 1 year of our operations. |
| Jul 2026 | Navraj Singh | | $950 | Other | bluedot | Scaling our RL-CAI pipeline (extends narrow secret-loyalty models via Constitutional AI) from 1.5B to a 32B organism, to test whether action-breadth extension holds at frontier model scale. |
| Jul 2026 | Caleb Price | | $6,100 | Training pipelines, Local groups, Events, Epistemic infrastructure | bluedot | The Lisbon AI Safety Hub (LAISH), the first of its kind AI safety hub in Portugal focusing on fieldbuilding AI safety. Running a fellowship, local community events and a newsletter. |
| Jul 2026 | Samuel Verboomen | | $5,000 | Other | bluedot | Building, through Polaris Collective, evaluation benchmarks for frontier AI models that test whether they consistently uphold commitments against enabling concentration of power. |
| Jul 2026 | Mazin Khider | | $400 | Other | bluedot | Testing whether behavioral grader-sensitivity measures predict later reward exploitation, beginning with a controlled replication of contrastive synthetic-document finetuning. |
| Jul 2026 | Tanmay Gejapati | | $100 | AI safety | bluedot | Testing whether fine-tuned language models can identify which individual facts they learned during fine-tuning, using controlled factual associations and internal probes. |
| Jul 2026 | Xavier Terminello | | $1,500 | Other | bluedot | Attending the Human-Aligned AI Summer School and developing digital-minds curricula |
| Jul 2026 | Pablo Romero | | $6,500 | Other | bluedot | Me and Phillip from AI Explained are working on a benchmark that tests AIs overconfidence on tasks, we get AI to complete a task and then ask how confident it is that the answer is correct. |
| Jul 2026 | Unknown Recipients | | $400 | Evals | bluedot | My project is about replicating and extending "AnthroBench", a library for generating, rating, and analyzing dialogues to evaluate anthropomorphic behaviors in LLMs. |
| Jul 2026 | Unknown Recipients | | $4,950 | Control, Evals, AI governance, policy, and advocacy | bluedot | Technical AI governance research at GovAI, currently assessing best practices for how frontier AI companies report model evaluations and emergency shutdown mechanisms for AI agents. |
| Jul 2026 | Athanasios Antonakoudis | | $100 | Other | bluedot | Extending the WMDP hazardous-knowledge benchmark with a calibration and selective-prediction analysis: does an open model know when it's wrong on dangerous-biology questions, and does abstention help? |
| Jul 2026 | Garrett Xu | | $650 | Other | bluedot | Developing internal deployment risk reporting guidelines and gap analyses of what the five largest AI developers disclose publicly |
| Jul 2026 | Vegas Harmon | | $400 | Evals | bluedot | Landscape paper on the verification and evals field |
| Jul 2026 | Samanvita Singhania | | $200 | Other | bluedot | Evaluating how much diverse base models for trusted monitoring improves the rate of correctly flagging suspicious code. |
| Jul 2026 | Olivia Velte | | $50 | Other | bluedot | Testing the robustness of J-lens workspace representations |
| Jul 2026 | Isabella Luong | | $1,000 | Other | bluedot | Attending the Las Vegas AI Security Forum |
| Jul 2026 | Isabella Luong | | $1,000 | Other | bluedot | Attending DEF CON to develop practical cybersecurity skills for AI security |
| Jul 2026 | Heather Alexander | | $10,000 | Other | bluedot | Workshop on the legal and philosophical status of AI personhood |
| Jul 2026 | Chris Pang | | $650 | Other | bluedot | ML-compatible theories of cooperation and empathy |
| Jul 2026 | Allison Zhuang | | $1,100 | AI governance, policy, and advocacy | bluedot | Poster Presentation at Berkeley Agentic AI Summit: "House Rules: Institutional Design in Multi-Agent LLM Poker Tournaments and Code Markets" |
| Jul 2026 | Joseph Kostousov | | $2,500 | AI safety | bluedot | Toronto AI Safety Initiative + Summer Intensive Programming |
| Jul 2026 | Sana Shams | | $1,750 | Events | bluedot | An AI safety landscape assessment event with 15-20 people working on AI safety. The goal is to map out the AI safety talent network, and identify initiatives for a long-term local AI safety hub. |
| Jul 2026 | Emmanuelle Gelain-Sohn | | $1,100 | AI safety | bluedot | A preregistered experiment on whether confidence in LLM outputs changes people's ability to catch errors in machine reasoning, or only the threshold at which they are willing to flag them. |
| Jul 2026 | Sophie Kim | | $9,000 | Local groups | bluedot | AI Safety Grantmaking Fundamentals Curriculum |
| Jul 2026 | Adhi Balamurugan | | $400 | Other | bluedot | An AI labor-market stress test to identify methods to triage occupations and commuting zones for transition assistance under different AI displacement scenarios. |
| Jul 2026 | Federico Cavanna | | $150 | Evals | bluedot | Determine whether a supervisor of a large language model (LLM) with the same affiliation as the agent it supervises evaluates certain types of evidence differently than a neutral agent. |
| Jul 2026 | Albrave Raphael Albayda | | $600 | AI safety | bluedot | Investigating whether Taglish suppresses internal representations of refusal in LLMs |
| Jul 2026 | Grace Reuben | | $1,600 | Other | bluedot | The AI Girls Club, a UK brand educating women on catastrophic risks from AI and how to act on them, through Instagram, a podcast, a WhatsApp community and London events. |
| Jul 2026 | Navraj Singh | | $150 | Interpretability, Alignment methods | bluedot | Applying CAI (SL-CAI+RL-CAI) to an SFT-installed secret loyalty organism to characterize how RL changes loyalty breadth in the activation×action taxonomy and whether HH-RLHF mitigates it. |
| Jul 2026 | Toby Pullan | | $500 | Other | bluedot | Preventing reasoning models from overthinking correct answers |
| Jul 2026 | Julian Moncarz | | $2,000 | Events | bluedot | Attending a Berkeley AI safety context-building workshop |
| Jul 2026 | Pronoy Das | | $250 | Other | bluedot | Study when instrumental misalignment-related side effects of RLVR emerge during training on Olmo-3.1-RL-zero. |
| Jul 2026 | Luth Gallemaso | | $1,000 | Interpretability | bluedot | Attending the Human-Aligned AI Summer School to develop interpretability expertise |
| Jul 2026 | Skye Yi-Hardie | | $150 | Alignment methods | bluedot | Investigating value and sycophancy in reward models: how much of reward model stereotyping stems from RLHF, and how much is inherited from base models? Sycophancy, or mere association? |
| Jul 2026 | Eddie Brown | | $600 | AI safety | bluedot | Building the Open LLM Safety Index so that developers can find the safest model for their use case |
| Jul 2026 | Tomás Gimenez Molina | | $600 | AI safety | bluedot | Testing whether LLMs fail to notice harmful concepts injected into their activations. Research for BlueDot’s TAIS Projects dictated by BAISH. |
| Jul 2026 | Hadiza Yusuf | | $150 | Other | bluedot | Replicating and extending Anthropic's Petri "coding audit realism" study: testing if realistic resources change reward hacking rates, across other models and non-coding tasks. |
| Jul 2026 | Unknown Recipients | | $3,000 | Events, Strategy & forecasting | bluedot | Bringing professional-grade threat modeling into AI safety. Building a foundational methodology doc, running workshops for orgs in the field, and creating the threat model landscape piece by piece. |
| Jul 2026 | Yelarys Yertaiuly | | $150 | Other | bluedot | Do models trained on documents about CoT monitors actually reason about being watched, or just avoid keywords? I separate the two by training the models with indirect-description of monitors. |
| Jul 2026 | Angelica Casuela | | $1,200 | AI safety | bluedot | Building a white-label platform for AI safety paper clubs |
| Jun 10, 2026 | Huey Lai | | $7,000 | Other | grantmaking_ai | To support compute costs, fees for 1 course facilitator, 1 guest speaker(Sonja, PhD researcher at CHIA at Cambridge) and Ashe's mentor fees with the expectation of supporting 7 scholars on the fellowship this summer. Since we are now oversubscribed and have also appointed new guest speakers and mentors, we would like some top-up funding. |
| Jun 2026 | Tom Taylor | | $1,200 | Other | bluedot | I am investigating agent actions in ‘The Village’; I am analysing how different agents delegate and deceive. Also searching for steganography evidence. I will share this research as a blog (series). |
| Jun 2026 | Tom Johnson | | $2,300 | Other | bluedot | Creating a dataset of agent trajectories to evaluate techniques for drift intervention and recovery |
| Jun 2026 | Patrick Martinez Peel | | $2,300 | Other | bluedot | Accountability infrastructure for AI agents ("accountability binding") - linking agent identity, delegation, and liability - and its bearing on AI legal status and digital minds' moral status. |
| Jun 2026 | Yernat Yestekov | | $2,300 | Evals, AI × cyber | bluedot | I’m building a new cyber evals nonprofit focused on evaluating frontier cyber capabilities above the human-expert level. |
| Jun 2026 | Karthik Viswanathan | | $1,900 | AI safety | bluedot | A "Mech Interp Dojo": I mentor 4 researchers formalizing the expressivity of attention heads in Lean 4, using LLM coding agents to draft proofs that the Lean kernel verifies. |
| Jun 2026 | Francesca Gomez | | $1,450 | AI safety | bluedot | Testing whether an escalation channel reduces reward hacking in frontier LLMs, using coding problems from the EvilGenie benchmark (technical feasibility test for a new AI Safety Org) |
| Jun 2026 | Will Yeadon | | $550 | AI safety | bluedot | Testing whether weaker LLMs can catch the first erroneous step in frontier models' physics reasoning and showing that giving monitors the answer suppresses scrutiny of the reasoning itself. |
| Jun 2026 | Shreyansh Tripathi | | $450 | AI safety | bluedot | Using Manifold Isometry to study Causal and Geometric Structure of Harm and Refusal in Language Models, checking whether harm and refusal are independent and which of these directionally comes first. |
| Jun 2026 | Jake Gardner | | $1,500 | Other | bluedot | My project evaluates whether splitting harmful queries across multiple files and sub-agents is able to meaningfully increase attack success rates. |
| Jun 2026 | Sidd Venkatesh | | $100 | Other | bluedot | Tracking whether emergent misalignment (Hitler-persona backdoor) appears in a model's internal activations before it shows up in behavior during fine-tuning, using linear probes and cosine similarity. |
| Jun 2026 | Unknown Recipients | | $4,000 | AI safety | bluedot | I am presenting my AI safety papers to ICML 2026 |
| Jun 2026 | Unknown Recipients | | $2,000 | Other | bluedot | Presenting my accepted archival paper at the ICML 2026 EMM-QA Workshop on evaluating multimodal foundation models and improving the reliability of AI evaluation. |
| Jun 2026 | Marc-Everin Carauleanu | | $10,000 | Technical AI safety research | bluedot | Inducing self-other overlap into LLMs at scale using supervised fine-tuning to reduce deception |
| Jun 2026 | Rohit Kaushik | | $100 | Interpretability, Alignment methods, Security | bluedot | Mechanistically understanding if RLHF/DPO unlearning are susceptible to jailbreak attack and extending current research to synchophancy |
| Jun 2026 | Guillermo Del Pinal | | $6,000 | Other existential risk | bluedot | Empirical ML/Alignment experiments focusing on under appreciated trade-offs between optimizing AI models to reduce X-risk vs reducing general harms |
| Jun 2026 | Dustin Kimmel | | $350 | Other | bluedot | Sycophancy & corrigibility research |
| Jun 2026 | Edy Nastase | | $2,500 | Alignment methods | bluedot | AI alignment depends on human judgement, yet it is under-researched. This project builds a synthetic cognitive model of the AI Debate judge and qualitatively analyses how real judges reason. |
| Jun 2026 | Carlos Rodriguez | | $150 | Interpretability, Evals | bluedot | Stress-testing behavioral evals for false assurance by replicating a mechanistic dissociation between behavioral compliance and causally active demographic encoding across current LLMs. |
| Jun 2026 | Changling Li | | $8,000 | Other | bluedot | My research focuses on the malicious use of AI and emergent dangerous behaviors, currently studying prompt injection and the limits of AI/human monitoring. |
| Jun 2026 | Unknown Recipients | | $150 | Technical AI safety research | bluedot | Replicating and extending the Alignment Faking in LLMs paper to more recent models |
| Jun 2026 | Unknown Recipients | | $100 | Other | bluedot | Evaluating whether machine unlearning methods genuinely erase knowledge or merely hide it behind refusal behaviour that can be reversed (additional). |
| Jun 2026 | Lachlan Ewart | | $2,000 | Other | bluedot | I am moving to London for August to cowork at LISA and gain as much context and learn as much as possible. |
| Jun 2026 | Edward Cant | | $50 | AI safety | bluedot | I'm working on a BlueDot technical project regarding subliminal learning in polypythias suite of 410M language models. |
| Jun 2026 | Tomasz Steifer | | $3,500 | Other | bluedot | Me and my team at the Centre for Credible AI is working on multi-agent debate, in particular focusing on a specific problem where agents exhibit echo-like behavior and other looping behaviors. |
| Jun 2026 | Unknown Recipients | | $2,300 | Other | bluedot | I'm an Oxford maths undergrad working on agentic attack control and AI security. I previously was a research contractor at Redwood Research, where I worked on attack selection in control evaluations. |
| Jun 2026 | Ana Paula Castillo Rodriguez | | $350 | Other | bluedot | I am the co-director SAIN Amsterdam and board member of Safe AI Netherlands. |
| Jun 2026 | Jacob Swann Kopczynski | | $100 | AI safety | bluedot | Replicating the Peer Preservation paper expecting to debunk significant parts of the findings via neutral/discouraging prompting |
| Jun 2026 | Vahini Chaudhary | | $100 | Other | bluedot | Testing if AI unlearning algorithms erase dangerous data or just piggyback on refusal mechanisms, leaving vulnerabilities that adversaries can bypass via activation ablation. |
| Jun 2026 | Mannat Vikramaditya Jain | | $10,000 | Other | bluedot | Goodfire (MARS V) + bio-ML and chem-ML to find alternative directions for AIxBio synthesis screening |
| Jun 2026 | Justin Wen | | $100 | Other | bluedot | Fine-tuning Llama 3 70B on documents portraying AI as a moral agent to test whether this induces alignment faking behavior measured via Anthropic's elicitation protocol. |
| Jun 2026 | Li Bearden | | $200 | AI governance, policy, and advocacy | bluedot | Attending EA Summit Vietnam to build relationships in the AI safety community as an independent researcher based in Southeast Asia, where in-person community access is the hardest constraint to meet. |
| Jun 2026 | Uzay Macar | | $3,000 | Interpretability, Evals | bluedot | Open-source agent for mechanistic interpretability. The user poses a question, and the agent co-designs a spec, runs causal experiments in a verify-and-improve loop, and returns an auditable report. |
| Jun 2026 | Vaibhav Pandey | | $200 | AI safety | bluedot | I am working on an empirical investigation of whether Model Spec Midtraining (published in the Anthropic Fellows Program of May 2026) can prevent the inverted persona failure mode in language models. |
| Jun 2026 | Jared Cheang | | $50 | Alignment methods | bluedot | Technical AI safety project - Failures in Scalable Oversight |
| Jun 2026 | Unknown Recipients | | $3,050 | Technical AI safety research, Events | bluedot | Presenting alignment research at an ICML workshop, and visiting ICML to gain experience and connections if I choose to found my own org. |
| Jun 2026 | Prasad Mahadik | | $1,500 | Interpretability, Control | bluedot | Improving the out-of-distribution generalization of linear probes and using them for robust monitoring of deception and sycophancy in LLMs. |
| Jun 2026 | Aditya Singh | | $200 | Other | bluedot | Do dangerous instrumental behaviors emerge silently through RL training, acquired without verbalizing? 19-checkpoint sweep using InstrumentalEval (76 scenarios) + CoT verbalization analysis. |
| Jun 2026 | Unknown Recipients | | $200 | Alignment methods, Control | bluedot | program: ARBOx4 in Oxford to build skills required for career in AI safety research. Project: exploring debate-based AI control protocols |
| Jun 2026 | Iván Arcuschin Moreno | | $1,000 | AI safety | bluedot | Directing an MSc thesis on how unverbalized biases enter LLM's chain-of-thoughts: finding the 'bias anchor' sentences and the attention path from the model's input to its output. |
| Jun 2026 | Shasanka Nudurupati | | $50 | Other | bluedot | Testing whether MacDiarmid et al.'s sleeper-agent defection probe generalizes to Hughes et al.'s alignment-faking models — an open question both papers flagged. |
| Jun 2026 | Unknown Recipients | | $500 | Evals, Training pipelines | bluedot | Running mechinterp experiments and agent evals on ICL-induced emergent misalignment. This is my project for the ARENA-based TARA program (https://taraprogram.org/). |
| Jun 2026 | Unknown Recipients | | $1,500 | Other | bluedot | Tech Thesis Ideathon - a day of research-crafting workshops for 30-40 women researchers and aspiring researchers in AI/AIS, organised in partnership with Lumina House. |
| Jun 2026 | Unknown Recipients | | $200 | Interpretability | bluedot | Mechanistic Interpretability in Text-to-Image Diffusion Models |
| Jun 2026 | Desiree Cho | | $3,000 | Other | bluedot | Constitutional Scaffolding in Continued Pre-Training for Improved Alignment Generalisation |
| Jun 2026 | Baris Alparslan | | $150 | Other | bluedot | Questioning Instrumental Convergence in RLVR: Checkpoint Evidence from OLMo-3 |
| Jun 2026 | Jaisakthi Chandrakumar | | $150 | AI safety | bluedot | Does Japanese Grammatical Humility Make Frontier LLM Models More Sycophantic? |
| Jun 2026 | Matthew Gregory | | $3,010 | Other | bluedot | Travel for AI Security Forum |
| Jun 2026 | Robin Haselhorst | | $5,000 | Other | bluedot | Hidden Behavior Detection (and Removal): Given a model, how can you efficiently figure out whether it's hiding some unusual behavior (backdoors, sleeper agents, reward hacking, censorship, ...)? |
| Jun 2026 | Unknown Recipients | | $1,800 | AI safety | bluedot | Helping draft and implement Vietnam's AI Law, Southeast Asia's first binding AI law, from inside the science ministry, and getting senior officials to take catastrophic AI risk seriously. |
| Jun 2026 | James Newport | | $8,000 | Other | bluedot | A 6-month series of highly focused, in-person AI scenario and policy events |
| Jun 2026 | Xueyuan Zhao | | $3,500 | Technical AI safety research, Epistemic infrastructure | bluedot | Testing whether subliminal learning in model distillation is driven by shared unembedding/readout alignment, and whether this risk appears in LLMs. |
| Jun 2026 | Juan P. Cadile | | $7,000 | Local groups, Events | bluedot | CNY AI Safety: a regional community launched through a UR-hosted hackathon for students/builders across Central & Western NY. |
| Jun 2026 | Andrada Marica | | $120 | Other | bluedot | I'm working on a project through the project sprint on model organisms of emergent misalignment. The goal is to better characterise the emergence of misalignment across training. |
| Jun 2026 | Archit Manek | | $1,500 | Interpretability, Evals, Events | bluedot | Presenting an accepted poster at the ICML 2026 Mechanistic Interpretability Workshop: a probing study of eval-awareness across 11 open-weight LLMs, completed during the BlueDot project sprint. |
| Jun 2026 | Caitlin Mah | | $8,000 | Technical governance, Training pipelines | bluedot | MARS V (CAISH) + PRISM fellowships: investigating whether Australian compute can serve as democratic AI governance leverage; auditing how frontier AI labs have redefined safety standards over time. |
| Jun 1, 2026 | Robi Rahman | | $10,000 | Other | grantmaking_ai | Blue Dot rapid grant #2 |
| Jun 2026 | Nicole Lai-Lopez | | $6,350 | Biosecurity | bluedot | Present Biosecurity spotlight paper at ICML GenBio 2026 Workshop in Seoul, Korea |
| Jun 2026 | Unknown Recipients | | $750 | Technical AI safety research | bluedot | Organizing NTU AI Safety in Taiwan, and attending ICML 2026 & WiML (Seoul, July 6–11) to present a research poster, connect with AI safety researchers, and bring insights back to grow our community. |
| Jun 2026 | Deven Choudhary | | $250 | AI safety | bluedot | First look inside India's sovereign LLMs (Param, Sarvam) at how they encode caste: is caste bias driven by a reused disgust circuit? Tested with activation steering, not just output scores. |
| Jun 2026 | Zhuang Ye | | $2,100 | Other | bluedot | Travel for AI Security Forum |
| Jun 2026 | Feil Immanuel Aquino | | $550 | AI governance, policy, and advocacy, Career transitions | bluedot | Support my transportation from Geneva to HAAISS in Prague, so I can attend the full 4-day program through its closing session as technical upskilling toward my work in AI governance. |
| Jun 2026 | Rohan Barad | | $2,500 | AI safety | bluedot | I run LI(ai)SE - London Interface for AI Safety, an AI safety hub, helping top of funnel people navigate the field and simultaneously building comms confidence to engage about AI Safety in public |
| Jun 2026 | Matthew Nguyen | | $6,300 | Evals | bluedot | Evaluating self-preference in LLM judges |
| Jun 2026 | Unknown Recipients | | $2,400 | Technical AI safety research | bluedot | Replicating and extending the Alignment Faking in LLMs paper to more recent Claude model(s) using more ambiguous user prompts |
| Jun 2026 | Charlotte Foudy | | $500 | Other | bluedot | An open red-team eval measuring whether chat models resist adversarially induced emotional dependency and false human/therapist claims, extending recent multi-turn safety-collapse work. |
| Jun 2026 | Jasmine Brazilek | | $5,000 | Other existential risk | bluedot | Forecasting the Cruxes: a reciprocal-scoring tournament on neglected alignment controversies |