Applied Scientist at Microsoft · GitHub

Evaluating agentic AI at scale

Benchmark design, LLM-as-judge systems, and outcome grading for coding agents, backed by multimodal research and production ML.

Scroll

I work on the hard part of agentic AI: knowing whether it actually worked.

I'm drawn to the unglamorous half of this field. A demo that works once is easy; knowing whether a system is actually good, and staying honest when the numbers flatter you, is the part I find genuinely interesting. Most of my career has circled that question, from medical imaging research to the coding agents I work on now.

I like problems where the measurement is harder than the model. I'd rather ship something small I can defend than something impressive I can only demo, and I have a persistent weakness for making expensive things cheap. There's a specific satisfaction in getting the same answer for a fraction of the cost.

Most of what I build outside work starts as a tool I wanted myself. The agentic job finder came from being annoyed at tailoring resumes by hand; the market-analysis assistant came from wanting to understand my own trading decisions. I also like writing things down. A cluster tutorial I wrote for labmates at UBC is still one of my favourite things I've made, precisely because it was useful to people who weren't me.

I'm based in Vancouver, and I'm always happy to talk about agent evaluation, generative models, or why your benchmark might be lying to you.

  • 0 Years in ML
Portrait of Nima Kondori
  • Agent Evaluation
  • Benchmark Design
  • LLM-as-Judge
  • Outcome Grading
  • Coding Agents
  • Multimodal RAG
  • Diffusion Models
  • ControlNet
  • VLMs
  • Model Distillation
  • PyTorch
  • Transformers
  • MCP
  • Airflow
  • Kafka
  • Kubernetes
  • Azure
  • AWS
  • Python

Published work

    1. 2025

      HAPPI: Hyperbolic Hierarchical Part Prototypes for Image Recognition

      Hooman Vaseli · Victoria Wu · Nima Kondori · Nguyen Nhat Minh To · Andrea Fung · Ang Nan Gu · Purang Abolmaesumi

      ICCV Workshops, Beyond Euclidean

    2. 2024

      ControlEchoSynth: Boosting Ejection Fraction Estimation Models via Controlled Video Diffusion

      Nima Kondori · Hanwen Liang · Hooman Vaseli · Bingyu Xie · Christina Luong · Purang Abolmaesumi · Teresa Tsang · Renjie Liao

      CVPR Workshop

    3. 2023

      ProtoASNet: Dynamic Prototypes for Inherently Interpretable and Uncertainty-Aware Aortic Stenosis Classification in Echocardiography

      Hooman Vaseli · Ang Nan Gu · S. Neda Ahmadi Amiri · Michael Y. Tsang · Andrea Fung · Nima Kondori · Armin Saadat · Purang Abolmaesumi · Teresa S. M. Tsang

      MICCAI

All publications on Google Scholar ↗

The path so far

Experience

Apr 2026 - Present

Applied Scientist

Microsoft · GitHub (Vancouver, BC)

  • Designed a multi-axis taxonomy and evaluated LLM classifiers for Copilot agent interactions, reaching 90%+ consistency on each intent axis and enabling offline-vs-online evaluation gap analysis.
  • Adapted intent classification to multi-turn sessions using silver labels from a stronger-LLM panel calibrated with human feedback, then trained a GNN-based model that preserved accuracy at roughly 2% of the original classifier's cost and latency.
  • Analyzed ~100k Copilot sessions with the Office of the CTO, identifying distribution gaps between offline evaluations and online usage that informed benchmark-instance selection for agent hillclimbing.
  • Designed and built an outcome grader for task completion and user dissatisfaction, adding deeper LLM-based failure analysis for high-dissatisfaction sessions.
Apr 2025 - Apr 2026

Senior Machine Learning Scientist

Lily AI (Mountain View, CA), remote from Vancouver, BC

  • Built a multimodal RAG system combining product text, images, and VLM signals, improving attribute-extraction F1 by 25% over legacy methods while cutting token usage 10% per product.
  • Designed SmartQC, an automated quality-control system using 3 LLM judges and MCP-backed tools, reducing human QC cost by 30% through high-confidence attribute auto-approval.
  • Fine-tuned transformer-based vision models for image-driven attribute extraction, enabling scalable visual catalog understanding.
Dec 2022 - Jan 2024

Machine Learning Engineer

CanDry Technologies (Vancouver, BC)

  • Built a ViT-based classifier for production visual quality control, achieving 85% precision and sub-30ms Android inference with TFLite.
  • Developed a ControlNet synthetic-data pipeline generating 1,000 photorealistic minority-class samples, improving rare-state recall by 15%.
  • Built a PyTorch Lightning and W&B experimentation platform with 500+ automated hyperparameter trials, halving experiment-iteration time.
Oct 2020 - Sep 2022

Machine Learning Engineer

Scenebox (Vancouver, BC)

  • Built and maintained Python and JavaScript workflows for LiDAR, image, video, geolocation, and time-series data using Airflow, FFmpeg, Kafka, ROS bag processors, Redis, and Elasticsearch.
  • Rebalanced search and identity-management workloads by moving selected Elasticsearch-backed data into SQL storage, reducing query latency by 20%.
  • Designed, implemented, and tested a SageMaker Ground Truth integration for enterprise labeling workflows, accelerating annotation turnaround by 30% for 5+ clients.
  • Maintained CI/CD infrastructure and deployed Scenebox clusters into customer AWS VPCs.

Research Experience

Sep 2022 - Apr 2025

Graduate Research Assistant

University of British Columbia (Vancouver, BC)

  • Led the technical work of a 5-person research team developing ControlNet-style video diffusion models for controllable cardiac structure and motion in generated echocardiograms.
  • Applied consistency distillation to the video-generation pipeline, reducing latency 140× from ~7 minutes to ~3 seconds per video while preserving downstream utility.
  • Curated a 100k-video echocardiogram dataset and trained ejection-fraction estimation models reaching 87% accuracy in cardiac-function prediction.

Education

Sep 2022 - Apr 2025

Master of Applied Science, Electrical & Computer Engineering

University of British Columbia, GPA 93.1%

  • Thesis: Exploring Video Diffusion Models in Echocardiogram Generation.
Sep 2017 - May 2020

Bachelor of Applied Science, Electrical & Computer Engineering

University of British Columbia, GPA 89%

Service

Academic Service

  • Reviewer for IEEE Transactions on Medical Imaging (TMI).
  • Reviewer for the NeurIPS 2024 Adaptive Foundation Models Workshop.

Honors & Awards

  • 2024 IEEE TMI Distinguished Reviewer Certificate
  • 2023 Canada Graduate Scholarship (CGS-M)
  • 2019 NSERC Undergraduate Student Research Award
Download résumé (PDF)