systems that tell you when they're wrong
I build LLM agent systems and the evaluation infrastructure that keeps them honest. At ClearAgent I built a LangGraph multi-agent service for agent compliance, and the harness that scores whether it actually works. Before that I shipped a production LLM pipeline that cut report delivery from days to hours across 200+ monthly reports. I care most about the part that usually gets skipped: measuring whether the thing you built does what you claimed, and publishing the answer when it doesn't.
I am finishing a Bachelor of Science in Data Science at University of California, Davis in Jun 2026, with a 3.60/4.0 GPA and 3x Dean's Honors List. Coursework closest to this work: Artificial Intelligence, Algorithms, Data Processing Pipelines, Databases, Big Data & High Performance Computing.
Experience
- ClearAgentSan Francisco, CA (Remote)AI Engineer25'
- DataCorp Traffic Private LimitedRemote, UKAI/ML Engineering Intern25'
- HeadstarterSan Francisco, CASoftware Engineering Intern, Agents & Frontend24'
Skills
- Languages
- Python, SQL, TypeScript, JavaScript, C++, R
- ML & AI
- PyTorch, Hugging Face Transformers, LangGraph, LangChain, RAG, Vector Embeddings, Cross-Encoder Reranking, Fine-Tuning, LLM-as-a-Judge, Multi-Agent Orchestration, Agentic Systems, Prompt Engineering, Model Evaluation, Model Context Protocol (MCP), LLM APIs (OpenAI, Anthropic), TensorFlow, scikit-learn
- Data & Experimentation
- Pandas, NumPy, ETL Pipelines, Experiment Design, Statistical Significance Testing, A/B Testing, Feature Engineering
- Infrastructure
- Docker, Kubernetes, AWS (EC2, S3, Lambda, RDS), PostgreSQL, ChromaDB, Kafka, Redis, FastAPI, REST APIs, GitHub Actions, CI/CD, Distributed Systems