Applied Scientist
Microsoft · GitHub (Vancouver, BC)
- Designed a multi-axis taxonomy and evaluated LLM classifiers for Copilot agent interactions, reaching 90%+ consistency on each intent axis and enabling offline-vs-online evaluation gap analysis.
- Adapted intent classification to multi-turn sessions using silver labels from a stronger-LLM panel calibrated with human feedback, then trained a GNN-based model that preserved accuracy at roughly 2% of the original classifier's cost and latency.
- Analyzed ~100k Copilot sessions with the Office of the CTO, identifying distribution gaps between offline evaluations and online usage that informed benchmark-instance selection for agent hillclimbing.
- Designed and built an outcome grader for task completion and user dissatisfaction, adding deeper LLM-based failure analysis for high-dissatisfaction sessions.