RAG Evaluation Is Not a Score
Measure retrieval, generation, citations, no-answer behavior, latency, and cost as separate system layers.
AI Agent Evaluation Beyond Prompting
Build eval sets, trajectory checks, regression tests, and business-ready quality gates.
Enterprise Data Agents with Fabric and Foundry
Connect governed data, ontology, MCP, and orchestration in a practical enterprise architecture.
Testing and Evaluating Copilot Agents
Use schema contracts, golden sets, graders, human review, and release monitoring.
Causal Inference for Product Analytics
Connect experiments, observational evidence, ATE, CATE, uplift, guardrails, and decisions.
Uplift Modeling in Industry
Move from average treatment effects to targeting by incremental impact.
Welcome, I'm Yangming LiI build applied AI systems that teams can evaluate, trust, and ship.
Applied AI engineer and product builder focused on LLM systems, evaluation, RAG, MLOps, data products, and production AI workflows for healthcare, finance, and enterprise teams.
I help translate AI prototypes into reliable product systems: clear data contracts, reviewable outputs, evaluation loops, and decision tools that fit the workflow.
Explore Yangming Li's work
Use these crawlable links to jump into the main areas of the site: profile, projects, applied AI systems, evaluation, machine learning, data products, product thinking, writing, resume, and contact.
A faster read on trust, fit, and delivery
Built for teams that care about adoption, auditability, and production impact, not just model demos. The homepage now keeps the proof points visible and moves side quests into quieter corners.
Builder mindset
"I'm a tool builder. That's how I think of myself. I want to build really good tools that I know in my gut and my heart will be valuable. And then, whatever happens, is... you can't really predict exactly what will happen, but you can feel the direction that we're going. And that's about as close as you can get. Then you just stand back and get out of the way, and these things take on a life of their own."
More ways to explore
Beyond the main work and projects, I also keep study notes, essays, small experiments, investing notes, and certificates here. They give extra context on how I learn, think, and build.
If you are here for collaboration or hiring, start with About, Projects, Blog, Resume, or Contact. If you are curious, the other links are open too.
Building an AI agent?
Download a practical launch checklist for evaluating Copilot Studio, RAG, document AI, and enterprise AI agents before production.
Featured writing
Six practical starting points across AI evaluation, enterprise agents, and data products. Browse the complete archive on the blog index.
Notes
Working notes, study artifacts, and lower-priority references that support the main body of work.
-
Carnegie Mellon University Advanced NLP Course Notes
These are my study notes from CMU's Advanced Natural Language Processing course. The notes cover fundamental concepts and advanced topics in NLP.
-
MIT Data Structure and Algorithms Course Notes
These are my study notes from MIT's Data Structure and Algorithms course. The notes cover fundamental algorithms, data structures, and their practical implementations.
-
MIT Principles of Computer Systems (6.826) Course Notes
These are my study notes from MIT's Principles of Computer Systems course. The notes cover distributed systems, concurrency, fault tolerance, and system design principles.
-
MIT Computation Structures (6.004) Course Notes
These are my study notes from MIT's Computation Structures course. The notes cover digital systems design, Boolean logic, computer architecture, and assembly language programming.
Selected Work
Interactive product prototypes and decision tools: explore the workflow, test the assumptions, and export a concrete review artifact.
Applied AI for document and knowledge systems
Designing LLM-assisted systems for document transformation, retrieval, review, and operational handoff in enterprise environments.
Decision-support products for complex teams
Building analytics and product experiences that help healthcare, finance, and public-sector teams move from raw data to better operational decisions.
From experimentation to production delivery
Shaping the foundations for reproducible ML, faster iteration, and experiment-ready delivery with MLOps and engineering discipline.
Focus Room
A premium SwiftUI focus app prototype for deep work: a soft hold-to-enter threshold, layered ambient sound mixing, a subtle timer, and a fullscreen study room that slowly deepens as the session unfolds.
From assumptions to a reviewable decision
Four browser-based projects that connect quantitative modeling to an operational decision. Start with a scenario, change a lever, inspect the tradeoff, and export the current inputs and results as a Markdown brief.
Monte Carlo Risk Forecast Studio
Treat each run like a committee review: one path could be a quarterly portfolio outcome, another a launch program under delivery pressure. The upper chart shows how scenarios drift apart over time, and the histogram reveals where the ending cases really cluster.
- Problem
- A review team needs to know how often an uncertain plan clears its required hurdle.
- Implementation
- Simulate compounding paths, compare the central outcome with the 10–90% range, and resample to inspect tail sensitivity.
- Deliverable
- A risk review brief with assumptions, hurdle probability, finish distribution, and the current decision narrative.
Model & validation boundaries
Independent Gaussian step shocks on a normalized index. Random resampling changes results; this is a scenario model, not a calibrated market forecast or a task-level delivery schedule.
A balanced review setup where the hurdle still feels reachable, but the tail risk is visible enough to force a real decision.
A lead is checking whether the plan still clears its hurdle before committing more capital, timeline, or scope.
The center line may look calm, but a few bad shocks widen the tail quickly and change the story for stakeholders.
Lower the hurdle, extend the horizon, or reduce exposure and scope if the success odds drift too low.
How the forecast can unfold
Where the ending scenarios cluster
Experiment readiness, AI capacity, and segmentation
Work through a launch review, a capacity planning conversation, or a segmentation investigation. Change the assumptions, inspect the output, and export a decision brief for the next review.
A/B Test Power Simulator
Model the decision pressure behind a launch review: how much sample, how much noise, and how much real lift you need before a "winner" deserves trust.
- Problem
- A product team must decide whether traffic and expected lift support a credible experiment read.
- Implementation
- Compare baseline conversion, relative lift, sample allocation, variance, and daily traffic under a fixed-horizon power approximation.
- Deliverable
- An experiment readiness brief with estimated power, error tradeoffs, duration, and a review recommendation.
Model & validation boundaries
Illustrative two-group conversion model. Estimates assume fixed-horizon testing; repeated peeking, interference, and metric bias require separate checks.
Null vs. uplift distribution
How long you need to sit on the test
Short tests feel faster, but they usually buy speed by borrowing confidence from the future.
LLM Cost-Latency Simulator
Stress-test an AI system the way a platform lead would: traffic, context size, retries, and cache behavior all fight over the same latency and budget envelope.
- Problem
- A platform team needs a budget and latency envelope before scaling a copilot or tool-using agent.
- Implementation
- Compare model tiers, traffic, input/output tokens, caching, and retries to expose the largest cost and latency drivers.
- Deliverable
- A capacity planning brief with monthly spend, estimated P95 latency, failure risk, and cache savings.
Model & validation boundaries
Synthetic model profiles and heuristic latency/failure estimates. These are not current vendor prices or measured service benchmarks; calibrate against real traces before release.
Where the monthly spend really goes
How queueing and retries bend the p95
UMAP / HDBSCAN Manifold Simulator
Compress a synthetic high-dimensional population into a neighborhood map and watch density structure survive, split, or dissolve as overlap, local scale, and minimum cluster size shift.
- Problem
- An analyst needs to distinguish stable neighborhoods from an attractive but misleading cluster picture.
- Implementation
- Explore a seeded synthetic population using local-neighborhood projection and density threshold clustering, then stress-test overlap and cluster size.
- Deliverable
- A segmentation review brief with neighborhood retention, detected groups, noise, and sensitivity notes.
Model & validation boundaries
A UMAP-inspired projection and DBSCAN-style density approximation, not the UMAP or HDBSCAN libraries. Synthetic labels provide an internal check; real data needs stability and domain review.
How local neighborhoods fold into 2D
Where dense structure becomes noise
Investing
Notes on capital allocation, market structure, and the quieter parts of long-term decision-making.
-
风险投资其实是“防风险投资”
张师傅的退休实验室关于 VC 本质、Vintage、投人、一级与二级市场差异,以及投资智慧的一篇长文。
Essays & References
A quieter corner for essays, references, and ideas that inform how I build.
-
Why "Taste" Matters in Science — and in Technology
Exploring Nobel laureate Yang Zhenning's concept of 'taste' in research and how it applies to technology and product development. What separates the merely competent from the truly visionary in science and tech.
-
Interesting Resource: Calculating Empires
I recently discovered an fascinating interactive resource called "Calculating Empires: A Genealogy of Technology and Power Since 1500". This comprehensive visualization maps out the intricate relationships between technology, power, and human history over the past 500 years.
-
Knowledge Flow
An interactive platform for visualizing and exploring connected knowledge across various domains. Knowledge Flow helps discover relationships between concepts and ideas in a structured format.