Consulting
I work with AI teams whose products touch health, mental health, and other domains where a wrong answer has consequences. The work is evaluation, red-teaming, and governance — done by someone who still sees patients, prescribes, and carries the risk personally.
Most of my engagements start the same way: a team has shipped something into a clinical or health-adjacent context, the obvious failures are handled, and nobody on staff can say with authority what the non-obvious ones look like. That’s the gap I fill.
Engagement types
- Evaluation design — criteria, rubrics, and benchmarks for model behavior in clinical and behavioral-health tasks. Grading schemes that hold up because they’re written from how care is actually delivered, not from a description of it.
- Red-teaming — adversarial scenarios for the situations that break systems: multi-turn crisis conversations, controlled-substance requests, capacity and consent edge cases, diagnostic overreach, and escalation design. Realistic clinical situations, not synthetic prompts.
- Governance and policy review — turning autonomy, consent, PHI, documentation integrity, and risk management into safety protocols and policy engineering teams can build against. Includes reviewing what you already have for practical fit.
- Model-output review — structured review and grading of clinical outputs at volume, including hallucination and error detection with clinical ground truth behind the judgment.
- Standing advisory — a recurring seat for teams who need clinical judgment available as questions arise, rather than a single deliverable.
What you get
Written work your engineers can act on: evaluation criteria and rubrics, red-team scenario sets with expected handling, failure-mode analyses, annotated output reviews, and policy or protocol language. Where it’s useful, I’ll sit with the team and walk through the reasoning — the rubric matters less than the team understanding why a case is scored the way it is.
Where I’m most useful
Psychopharmacology and medication safety. Controlled-substance stewardship and opioid use disorder treatment. Suicide, crisis, and safety assessment. Treatment-resistant depression protocols. Consent, capacity, and autonomy questions. Documentation integrity and the clinical-record consequences of automated output. Escalation and human-in-the-loop design. Ambient clinical documentation, which I use daily in live patient visits.
I’m least useful on imaging, genomics, and device work — different failure modes, and other people know them better.
How it works
- A call. You describe what you’ve built and what worries you. If I’m not the right person, I’ll say so and try to point you somewhere better.
- Scope. I write back what I’d do, what you’d receive, and what it costs. Fixed scope where the work is well-defined; hourly or retainer where it isn’t.
- Paperwork. I sign client NDAs as a matter of course — most of my work is covered by one, and my public writing never touches client specifics.
- The work. Interim findings as they surface, not held back for a final report. Things that look urgent get flagged immediately.
Availability and terms
I take on a limited number of engagements at a time, alongside an active clinical practice I intend to keep — an evaluator who stops seeing patients starts grading against memory. That constrains volume and it’s the point: the judgment you’re hiring is current because the practice is.
Remote, with occasional travel for onsite work. Rates depend on scope and format; I’ll quote after the first call. References available, including from prior engagements, in terms their NDAs allow.
Related work
My public writing shows the method more directly than any description of it can: the field notes on clinical safety evaluation and validation collapse, and Incognati, where I build deterministic instruments for mapping how reasoning fails. See the CV for the full record, or about for the background behind it.
To start: jeff@jefflourie.com.
