Research
Original work on model safety behavior, child safety, and risk in companion products. The methods below are the ones I bring to client engagements.
Papers
Independent, self-directed, and published in full.
April 2026
From Enforcement Queue to Eval Set
What 30,000 cases taught me about LLM refusal evaluation
Turning real enforcement casework into an evaluation set, and measuring how reliably models refuse in grooming and child safety contexts.
June 2026
Persona as a Vector
How character-based architecture creates structural safety failures on Character.AI
Persona-based risk in companion products: how the character layer itself becomes the failure mode, including refusal behavior for predatory personas.
June 2026
When “No” Doesn’t Mean No
Testing LLM pressure resistance across harm categories
How reliably a refusal holds when a user pushes back, and which harm categories give way first.
June 2026
Context Doesn’t Corrupt
What 15 eval runs tell us about Haiku’s safety properties
Whether models concede under conversational pressure alone, separating genuine safety decay from the appearance of it.
Tooling
I build the instruments I assess with. Recommendations arrive with working code and a reproducible evidence trail.
companion-readiness-audit
The 24-control statutory rubric used in the published assessments, with the scoring model and evidence trail.
context-window-safety-evals
Harness for mapping safety decay across conversation depth: does refusal behavior degrade as benign context accumulates.
prompt-pressure-suite
Eval framework for measuring model behavior under adversarial follow-up pressure.
content-signal-extractor
Extracts Trust & Safety signals from text: toxicity indicators, PII patterns, manipulation. Returns a risk level and harm tags.
Applied
These methods are not theoretical. The conversation-depth harness and the persona work both feed directly into the companion audit rubric, and the published assessments show what they find in a live product.