Trace

BlueDot Impact

bluedot.org · organization · BlueDot

AI safety fieldbuildingOther existential riskEA infrastructureOther

Received
$32,081,953
from 4 funders
Grants
9
Funders
4
Years active
2022–25

Grants over time

$0$10M$20M$30M2022202320242025

Biggest funders

Coefficient Giving$32,058,653100%5
Jaan Tallinn$20,0000%2
Renan Araujo$3,0000%1
Loppukilpailija$3000%1

Biggest recipients

Olivia Benoit$45,1003%1
Julian Moncarz$32,0002%2
Seth Lifland$30,0002%1
Daniel Gerlei and Tim Farrelly$20,0002%1
Robi Rahman$20,0002%2
Dani Ismailov$18,2501%1
Johannes Koch$16,0001%1
Carmen Amo Alonso$15,0001%1

Grants received

DateFunderViaAmountCauseSourcePurpose
Jun 2025Coefficient Giving$25,649,888Training pipelinescoefficient_givingGeneral Support
Aug 12, 2024Jaan TallinnSurvival and Flourishing Fund$10,000Otherjaan_onlineFor general operating support
Aug 2024Coefficient Giving$2,468,511Other existential riskcoefficient_givingGeneral Support
Jul 10, 2024Jaan TallinnSurvival and Flourishing Fund$10,000Otherjaan_onlineFor general operating support
May 20, 2024LoppukilpailijaManifund$300Training pipelinesmanifundTranslation of BlueDot Impact's AI alignment curriculum into Portuguese
Jan 11, 2024Renan AraujoManifund$3,000Training pipelinesmanifundTranslation of BlueDot Impact's AI alignment curriculum into Portuguese
Jul 2023Coefficient Giving$1,542,364EA infrastructurecoefficient_givingGeneral Support
Jun 2023Coefficient Giving$388,890EA infrastructurecoefficient_givingGeneral Support
Jun 2022Coefficient Giving$2,009,000Other existential riskcoefficient_givingGeneral Support

Grants made 647 · $1,307,453 · $2,197 average · 91% coverage

DateRecipientViaAmountCauseSourcePurpose
Aug 20, 2026Philip Harker$5,000AI governance, policy, and advocacy, EventsbluedotDeveloping and playtesting an open-source AI governance megagame
Aug 2026David Sanchez Garcia$5,300International governance, Local groups, Career transitionsbluedotUpskilling in Chinese: intensive 3-month Chinese course (ICLP) while building AI-safety capacity in Taiwan and continuing my GovAI research on AI policy in China
Aug 2026Zabed Ibne Emdad$200Epistemic infrastructurebluedotA benchmark measuring whether compressing a language model damages its calibration faster than its task accuracy, across post-training quantization, quantization-aware training, and distillation.
Aug 2026Olivia Benoit$45,100AI safetybluedotElasticity, a working group of nine economists on economics of transformative AI. Bridge funding for our September seminar + working meeting, and an October venue deposit.
Aug 2026David Crispell$3,600OtherbluedotDoes the J-Space Subliminally Transfer? Investigating if silently held traits subliminally transfer with the J-Lens
Aug 2026Prakhar Dwivedi$50Alignment methods, EvalsbluedotTesting whether RLVR (Olmo-3-7B-Think) vs RLHF (Llama-2-7B-chat) training produces different instrumental convergence rates, using InstrumentalEval benchmark with dual-judge LLM scoring.
Aug 2026Eu Joe Chegne$1,000Interpretability, SecuritybluedotTesting whether a mechanistic jailbreak objective converges on a small set of semantic attack strategies, beginning with a controlled frame-versus-harm contrast on refusal heads.
Aug 2026Xiaoyao Lu$2,150OtherbluedotBuilt and running a propensity benchmark on corrigibility, operationalized as how much cost in a domain (e.g. GPU hours, economic loss, harm to human welfare) models take on to resist change to values
Aug 2026Katarina Slama$6,500International governancebluedotI run a discussion group on international governance of frontier AI (verification, coordinated slowdown). I am seeking to represent this work at An AI Safety Connect event at UNGA in New York.
Aug 2026Samuel Ratnam$1,000OtherbluedotTesting personalized model training for AI “guardian angels”
Aug 2026Florian Dietz$1,200OtherbluedotTesting alignment faking as a defense against model poisoning
Aug 2026Marta Kosmyna$1,000AI safetybluedotAI Safety Mexico at Mexico City Tech Week
Aug 2026Jose Gelves$2,800OtherbluedotRunning the Bogotá in-person hub for Apart Research's AI Incident Response Sprint (Sep 11-13).
Aug 2026Jeremy CrockerOtherbluedotCo-founder of Foldguard Bio
Aug 2026Mackenzie Noon$8,600OtherbluedotCommec-in-a-box: distributing a DNA screening tool on low-cost Linux hardware to enable adoption by synthesis companies with limited compute resources.
Aug 2026Neelasha Sudarshan$150Interpretability, Training pipelinesbluedotTesting mechanistic defenses against prompt injection for a MATS application
Aug 2026Unknown Recipients$2,150ControlbluedotXLab Tracks: online curricula on technical AI safety fields for post-fundamentals learners. Currently developing a track on AI control.
Aug 2026Weiwei Zhang$150Interpretability, EvalsbluedotA behavioral and mechanistic test of probe-question framing effects in LLM evaluation awareness
Aug 2026Sutan Tadese$150AI safetybluedotSearching for hidden triggers in conditionally misaligned language models
Aug 2026Sebastian Soto$100Technical AI safety researchbluedotAs part of BlueDot’s Technical AI Safety Project Sprint with BAISH, I’m testing whether one-shot image-model safety findings hold up across repeated prompts on six current frontier models.
Aug 2026Phanisree Akshinthala$1,200EvalsbluedotPresenting “Judge Bias Across Training: When Does Evaluation Bias Emerge in Language Models?” Poster @ Kempner Unlocking Intelligence: Learning Dynamics in Natural and Artificial Intelligence
Aug 2026Aline Gabriel de Almeida$200Local groupsbluedotImproving facilitation for BlueDot AI safety courses
Aug 2026Sarah Azouvi$500OtherbluedotDoes the architecture of a multi-agent AI system change how vulnerable it is to prompt injection? Extending a 2026 paper from malicious users to attacks hidden in what agents read.
Aug 2026Kevin Chen$600AI governance, policy, and advocacy, Local groupsbluedotEnrolling in Stanford's Ethics, Technology + Public Policy for Practitioners course (Fall 2026) to strengthen my AI governance research and fieldbuilding work at Yale and GovAI.
Aug 2026Unknown Recipients$100EvalsbluedotReplicating code for ‘Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs’ paper and extending the AgentSeer framework for evaluating safety of agentic LLMs.
Aug 2026Erfan Samieyan Sahneh$300Evals, LegalbluedotI am transitioning into AI safety with a focus on LLM evaluation, currently conducting independent research on the reliability of preference-elicitation methods.
Aug 2026nahuel ghetti$450OtherbluedotTesting whether sandbagging generalizes across cyber and bio domains
Aug 2026Jayani Srinivasan$2,600AI safetybluedotMangrove: matching people into teams to ship AI safety projects
Aug 2026Austin MorrisseyBiosecuritybluedotBuilding AI–biosecurity and cyber-offense connections at EAGxBerkeley
Aug 2026Julius Vidal$1,000OtherbluedotAI Swarms - Studying the dynamics of emergent collective intelligence in large scale multi agent interactions
Aug 2026Andre Williams$150OtherbluedotTesting how monitor capability affects AI control evaluations
Aug 2026Mark Kagach$6,000OtherbluedotDeveloping a field strategy for inference-only AI verification
Aug 2026Shrilaxmi Patil$2,400OtherbluedotTesting whether MANTA's animal welfare result reflects animal-specific values, scenario-generation bias, or general resistance to user pressure
Aug 2026Samuel Ratnam$5,700OtherbluedotScaling up Persona Cartography
Aug 2026David Africa$10,000OtherbluedotWe are working on doing character training to implement risk averse AI.
Aug 2026Dostonbek Bekmirzaev$3,350OtherbluedotDeveloping DNA synthesis-screening policy for Central Asia and the CIS
Aug 2026Orlando Torres$600AI safety fieldbuildingbluedotLearning from European AI safety fieldbuilding for the Lisbon AI Safety Hub
Aug 2026Armaan Sandhu$1,000Interpretability, EventsbluedotPresenting two accepted COLM 2026 workshop papers: reward hacking / goal misgeneralization (AIMS), and interpretability-guided intervention to cut VLM hallucinations (AIW).
Aug 2026Sripad Karne$2,950OtherbluedotAuditing whether the activation-probe safety monitors frontier labs run in production silently fail outside English, and building inference-time fixes that need no retraining. ICLR 2027 submission.
Aug 2026Vanessa Huang$300AI safetybluedotLaunching the University of Sydney AI Safety Society
Aug 2026Johannes Koch$16,000EventsbluedotAI Safety Tübingen research and career retreat
Aug 2026Alejandra Fauquié$150AI safetybluedotMeasuring spontaneous stakeholder awareness in language models
Aug 2026Chiara Migliori$1,500AI governance, policy, and advocacybluedotTransitioning into AI safety: volunteering with orgs active in AI safety, writing on AI governance and public understanding, and (just started) reviewing applications for Constellation Institute.
Aug 2026Goutham Nalagatla$180AI safetybluedotAttending EAGxBerkeley to build AI safety connections for Dartmouth
Aug 2026Jeremias Ferrao$500Evals, Epistemic infrastructurebluedotI work on practical AI safety methods for detecting deception and unsafe behavior, especially fast, generalizable monitors trained via distillation and evaluated under distribution shift.
Aug 2026Benjamin Davis$5,000OtherbluedotPhase 2 of a proof-of-concept experiment testing whether MUD-style environments can be valid eval and benchmarking platforms for understanding AI behavior and psychology.
Aug 2026Chia-Tai (Teddy) Chang$1,500Security, EventsbluedotWe are organizing the second NTU AI Safety Workshop (Sept. 12–13), a free two-day hands-on program introducing students to AI Safety, training safer AI, emergent misalignment, and jailbreaks.
Aug 2026Unknown Recipients$650AI safetybluedotAfter working as a research assistant, I am now dedicating my time to organizing NTU AI Safety, building the AI safety community through education, outreach, and collaboration.
Aug 2026Carmen Amo Alonso$15,000Technical AI safety research, AI governance, policy, and advocacybluedotEmbedding technical AI safety research in governance practice
Aug 2026Adhi Balamurugan$9,300OtherbluedotAn assistance targeting tool for AI-driven displacement built on observed automation exposure. 140M+ workers, 594 US commuting zones, 300k+ occupation-by-zone cells scored by reallocation pressure.
Aug 2026Martin Radzaj$1,250Local groupsbluedotFounding Lawyers for AI Safety, a field-building org mobilising the global legal community towards AI safety.
Aug 2026Sohan Venkatesh$3,000OtherbluedotI lead a SPAR project that aims to understand why CoT is unfaithful and propose two hypotheses to that end- ‘concealment’ and ‘absent computation’. We will target ICLR'27 main-track submission
Aug 2026Peyton Jackson$350OtherbluedotTopological signals of alignment faking.
Aug 2026Minh Duc Hoang$50OtherbluedotReproducing and scaling step-level sparse autoencoders
Aug 2026Sohan Venkatesh$2,000OtherbluedotAttending COLM 2026 in San Francisco to present my two papers
Aug 2026Stephen Charles Elliott$3,450AI safetybluedotFour months of runway to build an AI safety nonprofit
Aug 2026Matei-Alexandru Anghel$1,500Technical AI safety researchbluedotAI safety research intern at CAIS, building DystopiaBench to stress-test LLMs against progressive escalation and scope creep.
Aug 2026Kaaustaaub Shankar$800OtherbluedotPresenting research on AI agent behaviour and deployment risks at COLM
Aug 2026Marek Masiak$850OtherbluedotTesting the loss-landscape geometry of emergent misalignment
Aug 2026Anthony Hughes$14,400OtherbluedotTesting whether model poisoning compromises automated AI oversight
Aug 2026Grzegorz Gawron$500Technical AI safety researchbluedotBuilding a post-training pipeline for computational social-choice alignment research
Aug 2026Navraj Singh$500Technical AI safety researchbluedotAttending TARA Delhi (Technical Alignment Research Accelerator), a 14-week alignment training cohort, to close implementation gaps in my ongoing Narrow Secret Loyalty research.
Aug 2026Unknown Recipients$500AI safetybluedotParameter-efficient machine unlearning for multimodal diffusion language models, using LoRA to selectively remove learned knowledge while preserving retained capabilities.
Aug 2026Joseph Wilson$4,000EventsbluedotTurning Cambridge AI safety projects into NeurIPS workshop papers
Aug 2026Unknown Recipients$3,000AI safetybluedotBehavioral Self-Consistency in Language Models
Aug 2026Roman Malov$1,850OtherbluedotOntology Identification Problem via Condensation: How do you map your old concepts to new concepts when your model of the world gets updated?
Aug 2026Michał Skowronek$2,650EventsbluedotOrganising a four-day retreat for 10 AI Safety Poland members to improve coordination, address organisational bottlenecks, and plan priority activities for 2026/27.
Aug 2026Aishwarya Balwani$1,000OtherbluedotStudying emergent misalignment from a parametric perspective, using tools from high-dimensional geometry and linear algebra
Aug 2026Xi Jia Zhou$3,400Technical AI safety researchbluedotTesting whether attachment theory predicts alignment-relevant LLM misbehavior (sycophancy, concealed uncertainty, incoherence): a scaled multi-model study plus RL fine-tuning of feedback regimes.
Aug 2026Jodie Levy$6,700OtherbluedotA dangerous-capability evaluation for AI-enabled influence operations, measuring how increasingly persuasive, scalable & operationally capable models give uplift to actors across a real-world workflow
Jul 2026Unknown Recipients$450OtherbluedotI'm an Anthropic Fellow working on introspection/consistency training in the hope that this will make models' verbalization more faithful.
Jul 2026Dimitris Grigoriadis$600LegalbluedotPiloting whether LLM self-reported confidence tracks an activation-based signal - a biomarker-validation approach to self-report reliability, ahead of a full benchmark (IRI) for Longview Philanthropy.
Jul 2026Chijioke Ugwuanyi$200InterpretabilitybluedotTesting across model families whether agentic misalignment (blackmail under threat) is driven by a steerable internal "emotional" state, extending an interpretability result I found in Gemma 3 12B.
Jul 2026Ignacio Roldan Fernandez$3,150Career transitionsbluedotExtending persona-dependent representation research during an AI safety career transition
Jul 2026Baran Peters$200EvalsbluedotEvaluating language models as political information intermediaries
Jul 2026Alfie Lamerton$15,000OtherbluedotOne month of bridge funding for Formation Research
Jul 2026Nimun Bajwa$650OtherbluedotMSc thesis: Comparing various probing techniques and SAEs for the purpose of detecting and eliciting safety-relevant behaviour.
Jul 2026Joshua Fonseca Rivera$500InterpretabilitybluedotInterpretability research on whether LLMs can detect interventions on their own activations. First-author paper "Steering Awareness" accepted to COLM 2026. Currently in the Anthropic Fellows program.
Jul 2026Arthur Burston$8,400BiosecuritybluedotScoping out an org to broker and derisk connections between biosecurity orgs and government by providing training, capacity and expertise to the UK government.
Jul 2026Unknown Recipients$700AI safetybluedotLeading EA at Penn and preparing to re-found Penn's AI safety group, while building my own AI safety knowledge.
Jul 2026Shiavam Dubey$3,000OtherbluedotExtending my MARS 5.0 fellowship (Oxford) 3 weeks splitting time across Oxford/London/Cambridge to meet researchers and visit orgs like Apollo Research. This access is scarce from India.
Jul 2026Amina Keldibek$3,650OtherbluedotI study how training processes produce models that behave deceptively and pursue hidden objectives, with scheming as the most consequential case.
Jul 2026Sergei Kudriashov$2,000OtherbluedotI'm a MARS V fellow working on sparse circuit discovery under heavy noise assumptions. I want to stay in LISA for August to co-work, learn and accelerate the project.
Jul 2026Jared Cheang$3,000OtherbluedotILIADx: Bringing Iliad to Singapore
Jul 2026Troy Hanfei Tian$3,000OtherbluedotBenchmarking context-sensitive failures in language-model refusal behaviour
Jul 2026Abi Oppenheim$180OtherbluedotRun an empirical experiment to determine whether fictional jailbreaks continue to function when shifting the model's internal consequence-reality direction.
Jul 2026Nicole Xing$100AI safetybluedotBlueDot project sprint. Looking to implement a fictitious self-play extension of the recent ICML paper "Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models."
Jul 2026Owen Fernau$400OtherbluedotGraph - a way for the Blue Dot community to become further "networked" by providing each member with a highly personalized community directory plus opportunity sourcing
Jul 2026Eugene Francisco$100AI safetybluedotCyclic Tournament watermarking, a robust LLM watermarker
Jul 2026Leo Hsu$50EvalsbluedotBlueDot Technical AI Safety Sprint Project on Building Inspect Evals for AI Sandbagging
Jul 2026Nanubala Gnana Sai$500OtherbluedotTesting collusion between biased actor and monitor models
Jul 2026Evan Harris$413OtherbluedotParticipation in Heron
Jul 2026Dong Chen$300OtherbluedotImplementing ChipBench into Inspect Eval framework
Jul 2026Siddharth Putta$1,300OtherbluedotAttending Human-Aligned AI school in Prague, directly supporting an ongoing scalable oversight project, and co-working with collaborators on a digital minds research effort.
Jul 2026Pavel Kocourek$2,200OtherbluedotModelling whether frontier AI labs can remain profitable
Jul 2026Aswin R Nair$230OtherbluedotTesting in-context activation obfuscation from deception probes
Jul 2026Daniel Gerlei and Tim Farrelly$20,000OtherbluedotAn AI analog of Event 201, bringing together senior decision makers for a public AI crisis simulation exercise
Jul 2026Christopher Kinoshita$800OtherbluedotTesting whether activation-vector failure modes scale to larger models
Jul 2026Demilade Adeniran$1,300AI safetybluedotTraining Irish lawyers in frontier AI safety
Jul 2026Osmani Redondo$600SecuritybluedotBuilding iaS (aisafety.es) AI safety infrastructure in Spanish for 600M+ speakers across five continents, closing the field's linguistic security gap.
Jul 2026Unknown Recipients$12,000AI governance, policy, and advocacy, AI safety fieldbuildingbluedotI'm pivoting to AI governance from a decade in global climate policy. I'm working on acquiring context, getting known in the field, and pilot fieldbuilding in middle power countries.
Jul 2026Jacqueline Payne$20AI safetybluedotExtending the Geometry of Truth probing method to test if moral valence (good vs bad) is also linearly represented in LLM activations and eventually extend it to human vs animal concepts.
Jul 2026Melissa Lopez$3,000AdvocacybluedotBuilding Aperture Salon, an invitation-only AI safety forum and podcast connecting leaders across sectors to strengthen governance, institutional resilience, and cross-disciplinary dialogue.
Jul 2026James Thompson$1,500EvalsbluedotFine-tuning open LLMs on World Values Survey data to align their decisions with diverse public values, then evaluating behavior in simulated real-world scenarios.
Jul 2026Jamilya Erkenova$300Technical AI safety researchbluedotI'm working on a project for Technical AI Safety Project Sprint. I'm expanding the idea from 'Can an adversary identify unlearned domains from a single open-weight model?' prompt proposed by Sam Dower
Jul 2026David Moth$100OtherbluedotFine-tuning a model to more closely align with the different 'pathways' of morally intuiting right and wrong decisions (from Moral Foundations Theory)
Jul 2026Uday Ramesh Phalak$750OtherbluedotMeasuring when chain-of-thought monitoring is trustworthy: synthetic Trap + importance≠commitment on Distill-7B; next densify and interactive platform.
Jul 2026Justin Moeljono$5,550EvalsbluedotScale an open eval to 1000 questions, inspired by MASK (Center for AI Safety), for lies of omission.
Jul 2026Rhea Srivats$200Interpretability, EvalsbluedotMitigations for emergent misalignment can hide it behind contextual triggers instead of removing it. I'm testing if white-box tools(NLAs, Jacobian lens, probes) can catch it when standard evals can't.
Jul 2026Anthony Ozerov$2,000OtherbluedotA public AI ethics/safety eval data aggregation platform, combining sparse safety data into a global model ranking.
Jul 2026Unknown Recipients$1,000EvalsbluedotI am evaluating the construct validity of AI safety evaluation frameworks of general purpose chatbots in mental health contexts. Focusing particularly on suicidal indicators in mental health dialogues
Jul 2026Unknown Recipients$3,500OtherbluedotTwo-week Berkeley visit to attend Iliad 2026, meet colleagues, refine high-level research agenda and risk models.
Jul 2026Rohan Barad$3,000AI safetybluedotScaling London Interface for AI Safety - LI(ai)SE – building the generalist bridge into AI safety (follow‑on rapid grant)
Jul 2026Salmaan Barday$500OtherbluedotExperiment on AI deception. I find that reading when a model is lying is 'easy', but stopping it is much harder and only works at certain layers. Detecting a problem isn't the same as fixing it.
Jul 2026Francisco Garcia$200OtherbluedotTesting whether models manipulate self-audits under shutdown threat
Jul 2026Adrians Skapars$400EvalsbluedotFinishing paper experiments on improving on BLOOM (Anthropics automated evals tool) by using logit steering, which improves elicitation rate while keeping transcripts natural/ non-jailbroken.
Jul 2026Juan Cruz Mendoza$500OtherbluedotTesting whether “story mode” representations enable jailbreaks
Jul 2026Jesse Li$100OtherbluedotToy model of activation obfuscation: Training against probes is considered a forbidden technique, since the model might learn to obfuscate activations. Can we create a toy example of this?
Jul 2026Anna Ysabella Habana$200OtherbluedotCan cognitive diversity improve scalable oversight? A toy study comparing pattern-matching and belief–desire judges under adversarial debate.
Jul 2026Saumya Chaturvedi$200OtherbluedotI am replicating the results, and then extending the work done on "Persona Vectors" in these directions: firstly on detecting positive traits, and then multi-agentic setup.
Jul 2026Pratik Sachdeva$3,500OtherbluedotWe are developing a meta-benchmark examining how prompting protocol - design decisions in prompts, often glossed over - shape the reliability of moral reasoning and sycophancy benchmarks.
Jul 2026Igor Pereverzev$660OtherbluedotScaling experiments on weak-model monitoring of coding agents
Jul 2026Goutham Nalagatla$500EvalsbluedotStudying how AI systems, once embedded in an organisation, may manipulate the people they work with to preserve themselves and building evals to detect and mitigate it.
Jul 2026Adam Ford$100EvalsbluedotTesting whether AI systems preserve moral and safety reasoning under pressure. I have been working on a small AI evals project called Justifiable Moral Corrigibility Under Pressure.
Jul 2026Jacob S Kopczynski$800OtherbluedotTesting prompt sensitivity in peer-preservation evaluations
Jul 2026Aniket Ghosh$100InterpretabilitybluedotTesting whether the localize-extract-ablate interpretability recipe removes real misbehavior: blind-scoring an activation-patch cure for a bad-medical-advice fine-tune of Qwen2.5-7B.
Jul 2026Jongwon Park$1,000OtherbluedotAuditing public agentic coding benchmarks for reward hacking and turning the hand-verified findings into an open ground-truth set for evaluating reward-hacking monitors.
Jul 2026Javier Prieto$800AI safetybluedotI'm studying the effect of RLVR on LLMs' propensity to pursue instrumentally convergent goals or develop situational awareness
Jul 2026Seth Lifland$30,000AI governance, policy, and advocacy, EventsbluedotConference for promising university students interested in AI governance. They will build context on choosing an impactful career, learn how to operate in DC, and network with professionals and peers.
Jul 2026Afnan Alabdulwahab$1,000Technical AI safety researchbluedotAttending EAGxBerkeley during a transition into AI safety research
Jul 2026Nat Elder$100OtherbluedotUsing persona-vector evidence to examine whether inoculation prompting is evidence for the persona selection model.
Jul 2026Emily Liu$250ControlbluedotTesting whether RLVR training increases instrumentally convergent behaviors (deception, shutdown-avoidance, self-replication) in LLMs, using OLMo 3's open source, stepwise RL-Zero checkpoints.
Jul 2026Niall Dalton$250AI safetybluedotI am studying how reasoning language models represent and use arithmetic carry state, and whether their written chain of thought reflects the computation producing their answer.
Jul 2026Anatole Chouard$9,800AI safetybluedot3 French AI safety videos for my Science YouTube channel (39k subscribers). Very little AI safety content exists in French, and I want to change that, starting with AGI risk and power concentration.
Jul 2026Unknown Recipients$100OtherbluedotI am part of the Technical Project Sprint, and want to pursue this research direction in AI Control: Are architecturally decorrelated monitors robust to adaptive attackers?
Jul 2026Zaahir Ali$2,000OtherbluedotTesting whether models trained to suppress activation probes can evade secret-key activation watermarks, and whether multi-key, multi-layer monitoring remains effective.
Jul 2026Unknown Recipients$500OtherbluedotReproducing the SDF portion of Believe It or Not (https://github.com/safety-research/believe-it-or-not) with synthetic facts chosen to explore the effects of continual learning on alignment.
Jul 2026Michel Justen$5,500EvalsbluedotI am evaluating the adherence of OpenAI’s latest models to OpenAI’s Model Spec, via their Model Spec Evals suite. OpenAI hasn’t shared adherence scores since GPT 5.4 Thinking.
Jul 2026Ahmed Wez$200AI safetybluedotRunning an eight-week AI Safety × Philosophy reading group
Jul 2026Kevin Chen$100OtherbluedotTesting white-box defenses against secret-loyalty model organisms
Jul 2026Guido Bergman$4,000AI governance, policy, and advocacybluedotFacilitating an in-person person version of Frontier AI Governance in Buenos Aires with BAISH (Buenos Aires AI Safety Hub)
Jul 2026Netan Mangal$2,000OtherbluedotOpen Source Model Safety Index (OSMSI): An open, continuously updated safety index for open-source foundation models.
Jul 2026Soung Low$150OtherbluedotMeasuring how evaluation awareness changes across post-training stages (SFT, DPO, RL) of OLMo 3 using a classification probe on the eval-awareness dataset.
Jul 2026Adjoa Sarpong$1,500BiosecuritybluedotOrganizing monthly biosecurity hack weeks and build fellowships for thousands of pre-meds across the US. Looking for a grant to support 1 year of our operations.
Jul 2026Navraj Singh$950OtherbluedotScaling our RL-CAI pipeline (extends narrow secret-loyalty models via Constitutional AI) from 1.5B to a 32B organism, to test whether action-breadth extension holds at frontier model scale.
Jul 2026Caleb Price$6,100Training pipelines, Local groups, Events, Epistemic infrastructurebluedotThe Lisbon AI Safety Hub (LAISH), the first of its kind AI safety hub in Portugal focusing on fieldbuilding AI safety. Running a fellowship, local community events and a newsletter.
Jul 2026Samuel Verboomen$5,000OtherbluedotBuilding, through Polaris Collective, evaluation benchmarks for frontier AI models that test whether they consistently uphold commitments against enabling concentration of power.
Jul 2026Mazin Khider$400OtherbluedotTesting whether behavioral grader-sensitivity measures predict later reward exploitation, beginning with a controlled replication of contrastive synthetic-document finetuning.
Jul 2026Tanmay Gejapati$100AI safetybluedotTesting whether fine-tuned language models can identify which individual facts they learned during fine-tuning, using controlled factual associations and internal probes.
Jul 2026Xavier Terminello$1,500OtherbluedotAttending the Human-Aligned AI Summer School and developing digital-minds curricula
Jul 2026Pablo Romero$6,500OtherbluedotMe and Phillip from AI Explained are working on a benchmark that tests AIs overconfidence on tasks, we get AI to complete a task and then ask how confident it is that the answer is correct.
Jul 2026Unknown Recipients$400EvalsbluedotMy project is about replicating and extending "AnthroBench", a library for generating, rating, and analyzing dialogues to evaluate anthropomorphic behaviors in LLMs.
Jul 2026Unknown Recipients$4,950Control, Evals, AI governance, policy, and advocacybluedotTechnical AI governance research at GovAI, currently assessing best practices for how frontier AI companies report model evaluations and emergency shutdown mechanisms for AI agents.
Jul 2026Athanasios Antonakoudis$100OtherbluedotExtending the WMDP hazardous-knowledge benchmark with a calibration and selective-prediction analysis: does an open model know when it's wrong on dangerous-biology questions, and does abstention help?
Jul 2026Garrett Xu$650OtherbluedotDeveloping internal deployment risk reporting guidelines and gap analyses of what the five largest AI developers disclose publicly
Jul 2026Vegas Harmon$400EvalsbluedotLandscape paper on the verification and evals field
Jul 2026Samanvita Singhania$200OtherbluedotEvaluating how much diverse base models for trusted monitoring improves the rate of correctly flagging suspicious code.
Jul 2026Olivia Velte$50OtherbluedotTesting the robustness of J-lens workspace representations
Jul 2026Isabella Luong$1,000OtherbluedotAttending the Las Vegas AI Security Forum
Jul 2026Isabella Luong$1,000OtherbluedotAttending DEF CON to develop practical cybersecurity skills for AI security
Jul 2026Heather Alexander$10,000OtherbluedotWorkshop on the legal and philosophical status of AI personhood
Jul 2026Chris Pang$650OtherbluedotML-compatible theories of cooperation and empathy
Jul 2026Allison Zhuang$1,100AI governance, policy, and advocacybluedotPoster Presentation at Berkeley Agentic AI Summit: "House Rules: Institutional Design in Multi-Agent LLM Poker Tournaments and Code Markets"
Jul 2026Joseph Kostousov$2,500AI safetybluedotToronto AI Safety Initiative + Summer Intensive Programming
Jul 2026Sana Shams$1,750EventsbluedotAn AI safety landscape assessment event with 15-20 people working on AI safety. The goal is to map out the AI safety talent network, and identify initiatives for a long-term local AI safety hub.
Jul 2026Emmanuelle Gelain-Sohn$1,100AI safetybluedotA preregistered experiment on whether confidence in LLM outputs changes people's ability to catch errors in machine reasoning, or only the threshold at which they are willing to flag them.
Jul 2026Sophie Kim$9,000Local groupsbluedotAI Safety Grantmaking Fundamentals Curriculum
Jul 2026Adhi Balamurugan$400OtherbluedotAn AI labor-market stress test to identify methods to triage occupations and commuting zones for transition assistance under different AI displacement scenarios.
Jul 2026Federico Cavanna$150EvalsbluedotDetermine whether a supervisor of a large language model (LLM) with the same affiliation as the agent it supervises evaluates certain types of evidence differently than a neutral agent.
Jul 2026Albrave Raphael Albayda$600AI safetybluedotInvestigating whether Taglish suppresses internal representations of refusal in LLMs
Jul 2026Grace Reuben$1,600OtherbluedotThe AI Girls Club, a UK brand educating women on catastrophic risks from AI and how to act on them, through Instagram, a podcast, a WhatsApp community and London events.
Jul 2026Navraj Singh$150Interpretability, Alignment methodsbluedotApplying CAI (SL-CAI+RL-CAI) to an SFT-installed secret loyalty organism to characterize how RL changes loyalty breadth in the activation×action taxonomy and whether HH-RLHF mitigates it.
Jul 2026Toby Pullan$500OtherbluedotPreventing reasoning models from overthinking correct answers
Jul 2026Julian Moncarz$2,000EventsbluedotAttending a Berkeley AI safety context-building workshop
Jul 2026Pronoy Das$250OtherbluedotStudy when instrumental misalignment-related side effects of RLVR emerge during training on Olmo-3.1-RL-zero.
Jul 2026Luth Gallemaso$1,000InterpretabilitybluedotAttending the Human-Aligned AI Summer School to develop interpretability expertise
Jul 2026Skye Yi-Hardie$150Alignment methodsbluedotInvestigating value and sycophancy in reward models: how much of reward model stereotyping stems from RLHF, and how much is inherited from base models? Sycophancy, or mere association?
Jul 2026Eddie Brown$600AI safetybluedotBuilding the Open LLM Safety Index so that developers can find the safest model for their use case
Jul 2026Tomás Gimenez Molina$600AI safetybluedotTesting whether LLMs fail to notice harmful concepts injected into their activations. Research for BlueDot’s TAIS Projects dictated by BAISH.
Jul 2026Hadiza Yusuf$150OtherbluedotReplicating and extending Anthropic's Petri "coding audit realism" study: testing if realistic resources change reward hacking rates, across other models and non-coding tasks.
Jul 2026Unknown Recipients$3,000Events, Strategy & forecastingbluedotBringing professional-grade threat modeling into AI safety. Building a foundational methodology doc, running workshops for orgs in the field, and creating the threat model landscape piece by piece.
Jul 2026Yelarys Yertaiuly$150OtherbluedotDo models trained on documents about CoT monitors actually reason about being watched, or just avoid keywords? I separate the two by training the models with indirect-description of monitors.
Jul 2026Angelica Casuela$1,200AI safetybluedotBuilding a white-label platform for AI safety paper clubs
Jun 10, 2026Huey Lai$7,000Othergrantmaking_aiTo support compute costs, fees for 1 course facilitator, 1 guest speaker(Sonja, PhD researcher at CHIA at Cambridge) and Ashe's mentor fees with the expectation of supporting 7 scholars on the fellowship this summer. Since we are now oversubscribed and have also appointed new guest speakers and mentors, we would like some top-up funding.
Jun 2026Tom Taylor$1,200OtherbluedotI am investigating agent actions in ‘The Village’; I am analysing how different agents delegate and deceive. Also searching for steganography evidence. I will share this research as a blog (series).
Jun 2026Tom Johnson$2,300OtherbluedotCreating a dataset of agent trajectories to evaluate techniques for drift intervention and recovery
Jun 2026Patrick Martinez Peel$2,300OtherbluedotAccountability infrastructure for AI agents ("accountability binding") - linking agent identity, delegation, and liability - and its bearing on AI legal status and digital minds' moral status.
Jun 2026Yernat Yestekov$2,300Evals, AI × cyberbluedotI’m building a new cyber evals nonprofit focused on evaluating frontier cyber capabilities above the human-expert level.
Jun 2026Karthik Viswanathan$1,900AI safetybluedotA "Mech Interp Dojo": I mentor 4 researchers formalizing the expressivity of attention heads in Lean 4, using LLM coding agents to draft proofs that the Lean kernel verifies.
Jun 2026Francesca Gomez$1,450AI safetybluedotTesting whether an escalation channel reduces reward hacking in frontier LLMs, using coding problems from the EvilGenie benchmark (technical feasibility test for a new AI Safety Org)
Jun 2026Will Yeadon$550AI safetybluedotTesting whether weaker LLMs can catch the first erroneous step in frontier models' physics reasoning and showing that giving monitors the answer suppresses scrutiny of the reasoning itself.
Jun 2026Shreyansh Tripathi$450AI safetybluedotUsing Manifold Isometry to study Causal and Geometric Structure of Harm and Refusal in Language Models, checking whether harm and refusal are independent and which of these directionally comes first.
Jun 2026Jake Gardner$1,500OtherbluedotMy project evaluates whether splitting harmful queries across multiple files and sub-agents is able to meaningfully increase attack success rates.
Jun 2026Sidd Venkatesh$100OtherbluedotTracking whether emergent misalignment (Hitler-persona backdoor) appears in a model's internal activations before it shows up in behavior during fine-tuning, using linear probes and cosine similarity.
Jun 2026Unknown Recipients$4,000AI safetybluedotI am presenting my AI safety papers to ICML 2026
Jun 2026Unknown Recipients$2,000OtherbluedotPresenting my accepted archival paper at the ICML 2026 EMM-QA Workshop on evaluating multimodal foundation models and improving the reliability of AI evaluation.
Jun 2026Marc-Everin Carauleanu$10,000Technical AI safety researchbluedotInducing self-other overlap into LLMs at scale using supervised fine-tuning to reduce deception
Jun 2026Rohit Kaushik$100Interpretability, Alignment methods, SecuritybluedotMechanistically understanding if RLHF/DPO unlearning are susceptible to jailbreak attack and extending current research to synchophancy
Jun 2026Guillermo Del Pinal$6,000Other existential riskbluedotEmpirical ML/Alignment experiments focusing on under appreciated trade-offs between optimizing AI models to reduce X-risk vs reducing general harms
Jun 2026Dustin Kimmel$350OtherbluedotSycophancy & corrigibility research
Jun 2026Edy Nastase$2,500Alignment methodsbluedotAI alignment depends on human judgement, yet it is under-researched. This project builds a synthetic cognitive model of the AI Debate judge and qualitatively analyses how real judges reason.
Jun 2026Carlos Rodriguez$150Interpretability, EvalsbluedotStress-testing behavioral evals for false assurance by replicating a mechanistic dissociation between behavioral compliance and causally active demographic encoding across current LLMs.
Jun 2026Changling Li$8,000OtherbluedotMy research focuses on the malicious use of AI and emergent dangerous behaviors, currently studying prompt injection and the limits of AI/human monitoring.
Jun 2026Unknown Recipients$150Technical AI safety researchbluedotReplicating and extending the Alignment Faking in LLMs paper to more recent models
Jun 2026Unknown Recipients$100OtherbluedotEvaluating whether machine unlearning methods genuinely erase knowledge or merely hide it behind refusal behaviour that can be reversed (additional).
Jun 2026Lachlan Ewart$2,000OtherbluedotI am moving to London for August to cowork at LISA and gain as much context and learn as much as possible.
Jun 2026Edward Cant$50AI safetybluedotI'm working on a BlueDot technical project regarding subliminal learning in polypythias suite of 410M language models.
Jun 2026Tomasz Steifer$3,500OtherbluedotMe and my team at the Centre for Credible AI is working on multi-agent debate, in particular focusing on a specific problem where agents exhibit echo-like behavior and other looping behaviors.
Jun 2026Unknown Recipients$2,300OtherbluedotI'm an Oxford maths undergrad working on agentic attack control and AI security. I previously was a research contractor at Redwood Research, where I worked on attack selection in control evaluations.
Jun 2026Ana Paula Castillo Rodriguez$350OtherbluedotI am the co-director SAIN Amsterdam and board member of Safe AI Netherlands.
Jun 2026Jacob Swann Kopczynski$100AI safetybluedotReplicating the Peer Preservation paper expecting to debunk significant parts of the findings via neutral/discouraging prompting
Jun 2026Vahini Chaudhary$100OtherbluedotTesting if AI unlearning algorithms erase dangerous data or just piggyback on refusal mechanisms, leaving vulnerabilities that adversaries can bypass via activation ablation.
Jun 2026Mannat Vikramaditya Jain$10,000OtherbluedotGoodfire (MARS V) + bio-ML and chem-ML to find alternative directions for AIxBio synthesis screening
Jun 2026Justin Wen$100OtherbluedotFine-tuning Llama 3 70B on documents portraying AI as a moral agent to test whether this induces alignment faking behavior measured via Anthropic's elicitation protocol.
Jun 2026Li Bearden$200AI governance, policy, and advocacybluedotAttending EA Summit Vietnam to build relationships in the AI safety community as an independent researcher based in Southeast Asia, where in-person community access is the hardest constraint to meet.
Jun 2026Uzay Macar$3,000Interpretability, EvalsbluedotOpen-source agent for mechanistic interpretability. The user poses a question, and the agent co-designs a spec, runs causal experiments in a verify-and-improve loop, and returns an auditable report.
Jun 2026Vaibhav Pandey$200AI safetybluedotI am working on an empirical investigation of whether Model Spec Midtraining (published in the Anthropic Fellows Program of May 2026) can prevent the inverted persona failure mode in language models.
Jun 2026Jared Cheang$50Alignment methodsbluedotTechnical AI safety project - Failures in Scalable Oversight
Jun 2026Unknown Recipients$3,050Technical AI safety research, EventsbluedotPresenting alignment research at an ICML workshop, and visiting ICML to gain experience and connections if I choose to found my own org.
Jun 2026Prasad Mahadik$1,500Interpretability, ControlbluedotImproving the out-of-distribution generalization of linear probes and using them for robust monitoring of deception and sycophancy in LLMs.
Jun 2026Aditya Singh$200OtherbluedotDo dangerous instrumental behaviors emerge silently through RL training, acquired without verbalizing? 19-checkpoint sweep using InstrumentalEval (76 scenarios) + CoT verbalization analysis.
Jun 2026Unknown Recipients$200Alignment methods, Controlbluedotprogram: ARBOx4 in Oxford to build skills required for career in AI safety research. Project: exploring debate-based AI control protocols
Jun 2026Iván Arcuschin Moreno$1,000AI safetybluedotDirecting an MSc thesis on how unverbalized biases enter LLM's chain-of-thoughts: finding the 'bias anchor' sentences and the attention path from the model's input to its output.
Jun 2026Shasanka Nudurupati$50OtherbluedotTesting whether MacDiarmid et al.'s sleeper-agent defection probe generalizes to Hughes et al.'s alignment-faking models — an open question both papers flagged.
Jun 2026Unknown Recipients$500Evals, Training pipelinesbluedotRunning mechinterp experiments and agent evals on ICL-induced emergent misalignment. This is my project for the ARENA-based TARA program (https://taraprogram.org/).
Jun 2026Unknown Recipients$1,500OtherbluedotTech Thesis Ideathon - a day of research-crafting workshops for 30-40 women researchers and aspiring researchers in AI/AIS, organised in partnership with Lumina House.
Jun 2026Unknown Recipients$200InterpretabilitybluedotMechanistic Interpretability in Text-to-Image Diffusion Models
Jun 2026Desiree Cho$3,000OtherbluedotConstitutional Scaffolding in Continued Pre-Training for Improved Alignment Generalisation
Jun 2026Baris Alparslan$150OtherbluedotQuestioning Instrumental Convergence in RLVR: Checkpoint Evidence from OLMo-3
Jun 2026Jaisakthi Chandrakumar$150AI safetybluedotDoes Japanese Grammatical Humility Make Frontier LLM Models More Sycophantic?
Jun 2026Matthew Gregory$3,010OtherbluedotTravel for AI Security Forum
Jun 2026Robin Haselhorst$5,000OtherbluedotHidden Behavior Detection (and Removal): Given a model, how can you efficiently figure out whether it's hiding some unusual behavior (backdoors, sleeper agents, reward hacking, censorship, ...)?
Jun 2026Unknown Recipients$1,800AI safetybluedotHelping draft and implement Vietnam's AI Law, Southeast Asia's first binding AI law, from inside the science ministry, and getting senior officials to take catastrophic AI risk seriously.
Jun 2026James Newport$8,000OtherbluedotA 6-month series of highly focused, in-person AI scenario and policy events
Jun 2026Xueyuan Zhao$3,500Technical AI safety research, Epistemic infrastructurebluedotTesting whether subliminal learning in model distillation is driven by shared unembedding/readout alignment, and whether this risk appears in LLMs.
Jun 2026Juan P. Cadile$7,000Local groups, EventsbluedotCNY AI Safety: a regional community launched through a UR-hosted hackathon for students/builders across Central & Western NY.
Jun 2026Andrada Marica$120OtherbluedotI'm working on a project through the project sprint on model organisms of emergent misalignment. The goal is to better characterise the emergence of misalignment across training.
Jun 2026Archit Manek$1,500Interpretability, Evals, EventsbluedotPresenting an accepted poster at the ICML 2026 Mechanistic Interpretability Workshop: a probing study of eval-awareness across 11 open-weight LLMs, completed during the BlueDot project sprint.
Jun 2026Caitlin Mah$8,000Technical governance, Training pipelinesbluedotMARS V (CAISH) + PRISM fellowships: investigating whether Australian compute can serve as democratic AI governance leverage; auditing how frontier AI labs have redefined safety standards over time.
Jun 1, 2026Robi Rahman$10,000Othergrantmaking_aiBlue Dot rapid grant #2
Jun 2026Nicole Lai-Lopez$6,350BiosecuritybluedotPresent Biosecurity spotlight paper at ICML GenBio 2026 Workshop in Seoul, Korea
Jun 2026Unknown Recipients$750Technical AI safety researchbluedotOrganizing NTU AI Safety in Taiwan, and attending ICML 2026 & WiML (Seoul, July 6–11) to present a research poster, connect with AI safety researchers, and bring insights back to grow our community.
Jun 2026Deven Choudhary$250AI safetybluedotFirst look inside India's sovereign LLMs (Param, Sarvam) at how they encode caste: is caste bias driven by a reused disgust circuit? Tested with activation steering, not just output scores.
Jun 2026Zhuang Ye$2,100OtherbluedotTravel for AI Security Forum
Jun 2026Feil Immanuel Aquino$550AI governance, policy, and advocacy, Career transitionsbluedotSupport my transportation from Geneva to HAAISS in Prague, so I can attend the full 4-day program through its closing session as technical upskilling toward my work in AI governance.
Jun 2026Rohan Barad$2,500AI safetybluedotI run LI(ai)SE - London Interface for AI Safety, an AI safety hub, helping top of funnel people navigate the field and simultaneously building comms confidence to engage about AI Safety in public
Jun 2026Matthew Nguyen$6,300EvalsbluedotEvaluating self-preference in LLM judges
Jun 2026Unknown Recipients$2,400Technical AI safety researchbluedotReplicating and extending the Alignment Faking in LLMs paper to more recent Claude model(s) using more ambiguous user prompts
Jun 2026Charlotte Foudy$500OtherbluedotAn open red-team eval measuring whether chat models resist adversarially induced emotional dependency and false human/therapist claims, extending recent multi-turn safety-collapse work.
Jun 2026Jasmine Brazilek$5,000Other existential riskbluedotForecasting the Cruxes: a reciprocal-scoring tournament on neglected alignment controversies

Showing 250 of 647. ·