Consulting & speaking
Clear thinking for teams building with AI.
I help technical leaders, researchers, and organizations understand what language models can do, where evaluations can fail, and how to build more reliable AI systems.
Start a conversation01
Talks
Accessible, research-grounded talks for technical teams, leadership groups, conferences, and general audiences.
- When AI knows it is being evaluated An introduction to evaluation awareness—how models may recognize tests, alter their behavior, and weaken the evidence used to assess them.
- Backdoors and hidden behavior in language models A clear look at hidden triggers, deceptive model behavior, and the practical challenge of detecting failures before deployment.
- How to evaluate autonomous AI agents A practical framework for testing whether agents can complete multi-step tasks reliably, safely, and under realistic conditions.
- What today’s AI systems can—and cannot—do A research-grounded overview of current capabilities, limitations, and the implications for organizational decisions.
02
Training
Hands-on sessions that give teams shared concepts, practical methods, and concrete next steps.
- Designing useful and robust LLM evaluations Learn to define target behaviors, choose representative tasks and metrics, and stress-test evaluations against misleading results.
- Red-teaming models for hidden failure modes Apply structured probes and adversarial scenarios to uncover context-dependent, unsafe, or intentionally concealed model behavior.
- Building and benchmarking language-model agents Develop task environments, baselines, and measurements that reveal where agents succeed, fail, and need stronger safeguards.
- AI literacy for technical and non-technical teams Build a shared understanding of how modern AI works, where it remains unreliable, and how to use it responsibly.
03
Advisory
Focused support for consequential technical and strategic decisions involving language models and agents.
- Evaluation and reliability strategy Align evaluation methods with real deployment risks, product requirements, and the decisions your organization needs to make.
- Agent safety and deployment readiness Assess agent failure modes, operational controls, and evidence of reliability before systems reach users or sensitive environments.
- Research roadmaps and experiment design Prioritize high-value questions and design experiments that produce credible, decision-relevant evidence.
- Independent review of AI claims and results Pressure-test methods, evidence, and conclusions before important technical, investment, or deployment decisions.
Frequently asked questions
What to expect
Which service is the right fit?
Talks are designed to inform and provoke discussion, training builds practical skills, and advisory work supports a specific technical or strategic decision. If the right format is not obvious, describe your goal and I will recommend one.
Who are these services designed for?
I work with technical teams, leadership groups, researchers, event audiences, and organizations adopting language models or autonomous agents.
How technical can the content be?
The depth is adapted to the audience, from accessible AI literacy and executive briefings to detailed sessions on evaluation design, red-teaming, benchmarks, and agent safety.
Can a talk, workshop, or engagement be customized?
Yes. Topics, examples, exercises, and recommendations can be tailored to your organization’s goals, systems, and level of AI experience.
Are services available in English and French?
Yes. Talks, training, and advisory engagements can be delivered in English or French.
How do we get started?
Send a short email describing your audience, goal, preferred timing, and desired level of technical depth. I will suggest a suitable format and next steps.