An offline, deterministic red-teaming and evaluation harness for tool-using LLM agents — that measures and publishes its own detectors' error rate.
BUILDING SYSTEMS THAT TELL YOU WHEN THEY FAIL
I build LLM agent systems and the evaluation infrastructure that keeps them honest. At ClearAgent I built a LangGraph multi-agent service for agent compliance, and the harness that scores whether it actually works. Before that I shipped a production LLM pipeline that cut report delivery from days to hours across 200+ monthly reports. I care most about the part that usually gets skipped: measuring whether the thing you built does what you claimed, and publishing the answer when it doesn't.
- ClearAgentSan Francisco, CA (Remote)AI Engineer25'
- DataCorp Traffic Private LimitedRemote, UKAI/ML Engineering Intern25'
- HeadstarterSan Francisco, CASoftware Engineering Intern, Agents & Frontend24'
Selected Works
Stories.
View allJUL 2, 2026 · 2 MIN READ
Zero-Shot Beat Everything, and I Published It
Eleven conditions of chain-of-thought prompting and fine-tuning on table reasoning. The baseline won. Here is why that is worth reporting.
JUN 10, 2026 · 3 MIN READ
What "Compliant Agent Behavior" Actually Means
Turning EU AI Act obligations written for organizations into three dimensions a system can be scored against.