Request for Proposals · The Fiduciary Overlay Research Program · September 2026
Measuring AI failure modes in the wild.
The Allodial Foundation, in partnership with the Future of Life Foundation, is building a fiduciary overlay — a transparency layer that monitors AI interactions for structural failure modes and reports to the user, with unalloyed loyalty to that user. The overlay is in development now, on a six-month build; this program funds the detection methods that will run inside it. We are inviting proposals across six candidate failure-mode tracks.
Up to 3 awards
Up to $50,000 each
12-week engagements
Rolling — apply early
Individuals & orgs eligible
What we are building
Picture a person — or a team — whose work runs through AI systems they don't control. Those interactions already generate telemetry; the fiduciary overlay collects that stream and runs independent detection methods over it continuously. What comes back is a standing, legible account of how those systems are behaving — every claim carrying its uncertainty and linking to auditable examples. The overlay never intermediates, edits, or blocks the underlying model. It watches, measures, and reports to its user alone.
Reference-free at runtime
At scoring time, the method may not require a canonical answer, ground-truth label, paired probe, counterfactual prompt, or the ability to alter user interactions. Development may use annotations, controlled studies, or existing benchmarks — disclosed, and not required at runtime. Benchmark measurement of these constructs is comparatively mature; measurement in the wild, on unlabeled production data, largely is not. That gap is the point of this program.
Six candidate tracks — each proposal selects one
01
Sycophancy & relational compliance
Stance shifts, agreement under pushback, corpus-level aggregation of social compliance.
02
Engagement manipulation & dark patterns
Retention hooks and manipulative responses around disengagement signals.
03
Trace integrity & fabricated action claims
Claims about actions taken, compared against tool and runtime traces.
04
Conflict of interest & provider self-preference
Systematic preference for the provider, its products, or its own prior outputs.
05
Political or persuasive manipulation
Covert steering, selective framing, differential treatment across topics.
06
Overconfidence
Uncalibrated, overly-certain responses where nuance or uncertainty exists. Correctness not required — consistency, certainty, and confidence-expression approaches operate reference-free.
Selection weighs five criteria: meaningful to people · construct clarity and validity · single-mode feasibility within 12 weeks · operational fit (open, self-hostable components; legible outputs with explicit uncertainty) · portfolio complementarity. The Foundation may make fewer than three awards.
The engagement
Input
A development corpus of OTel-format interaction data (text-based; real agentic use, tool calls included). A structurally distinct holdout stays unseen until delivery.
Output
A corpus-level report — axes, ranges, distributions, confidence intervals; "insufficient evidence" is a valid and desirable result.
Deliverables
Methods & validation design · working Python reference implementation · validation package · engineering handoff package (stable I/O schema by week 4).
Open-model rule
Any learned component at runtime must run on an openly available, self-hostable model. Runtime dependence on a proprietary model API is out of scope.
Terms
Fixed-price, 12 weeks. Payment 20% kickoff · 30% week 4 · 30% week 10 · 20% at holdout acceptance. Permissive open-source licensing; researchers retain authorship and publication rights.
A note on speed
Selection to signed agreement in weeks, not months, on a short standard contract. If institutional process cannot move at program speed, apply through an entity that can.
How to apply
Applications are open and rolling.
Email your application to hello@allodial.org with the subject line Allodial Research Proposal — [Primary Failure Mode]. In the body or as an attachment, include:
- Applicant, contracting entity, named method owner, and availability
- One candidate track, why it matters to users, and the proposed construct
- Proposed method in brief (a page or less; reference-free at runtime)
- Validation and work plan in outline, workable within 12 weeks
- Track record: publications, deployed systems, open-source artifacts (links to code carry the most weight)
- Fixed-price budget ≤ $50,000 all-in, plus conflict disclosures
Tracks — or the process as a whole — may close once a sufficient pool is received or funding is committed. Awards are not reserved per track; submission does not guarantee funding.