Runtime Snapshot
yangming_profile.py live object
class YangmingLi: @staticmethod def builds() -> list: return ["AI system architecture", "LLM systems", "Statistical ML", "Data engineering", "Experiment infrastructure"] @staticmethod def serves() -> list: return ["Healthcare teams", "Finance teams", "Enterprise teams"] @staticmethod def outcomes() -> list: return ["Production AI systems", "Faster decision loops", "Reusable internal tooling"] # Instantiate Yangming Li builder = YangmingLi() print(f"Builds: {builder.builds()}") print(f"Serves: {builder.serves()}") print(f"Outcomes: {builder.outcomes()}")
>focus_areas = 5
>industries_covered = 4
>delivery_mode = "prototype to production"
Evidence Strip

A faster read on trust, fit, and delivery

Built for teams that care about adoption, auditability, and production impact, not just model demos. The homepage now keeps the proof points visible and moves side quests into quieter corners.

Industries
Healthcare operations Finance and risk Public sector delivery Research environments
Capability Areas
LLM systems Statistical ML Data engineering Decision-support products
Delivery Modes
Dashboards and scorecards APIs and microservices Copilots and RAG systems Monitoring-ready pipelines
Public Artifacts
Peer-reviewed publication Technical guides Selected portfolio work CFA and FRM credentials
Working Style

Builder mindset

Tool Builder Mindset
"I'm a tool builder. That's how I think of myself. I want to build really good tools that I know in my gut and my heart will be valuable. And then, whatever happens, is... you can't really predict exactly what will happen, but you can feel the direction that we're going. And that's about as close as you can get. Then you just stand back and get out of the way, and these things take on a life of their own."
Projects & Notes

More ways to explore

Beyond the main work and projects, I also keep study notes, essays, small experiments, investing notes, and certificates here. They give extra context on how I learn, think, and build.

If you are here for collaboration or hiring, start with About, Projects, Blog, Resume, or Contact. If you are curious, the other links are open too.

AI Agent Evaluation

Building an AI agent?

Download a practical launch checklist for evaluating Copilot Studio, RAG, document AI, and enterprise AI agents before production.

Six practical starting points across AI evaluation, enterprise agents, and data products. Browse the complete archive on the blog index.

Notes

Working notes, study artifacts, and lower-priority references that support the main body of work.

Selected Work

Interactive product prototypes and decision tools: explore the workflow, test the assumptions, and export a concrete review artifact.

Featured Interactive Experience

Focus Room

A premium SwiftUI focus app prototype for deep work: a soft hold-to-enter threshold, layered ambient sound mixing, a subtle timer, and a fullscreen study room that slowly deepens as the session unfolds.

Hold-to-enter ritual Ambient mixer Ghost UI Session evolution Local persistence
Good work. Take a breath.
Focus Timer
31:42
Focusing
Ambient Layers
Piano vinyl warmth
Rain window hush
Brown deep bed
Cafe soft distance
White clean edge
Applied engineering portfolio

From assumptions to a reviewable decision

Four browser-based projects that connect quantitative modeling to an operational decision. Start with a scenario, change a lever, inspect the tradeoff, and export the current inputs and results as a Markdown brief.

Risk Forecast Studio

Monte Carlo Risk Forecast Studio

Treat each run like a committee review: one path could be a quarterly portfolio outcome, another a launch program under delivery pressure. The upper chart shows how scenarios drift apart over time, and the histogram reveals where the ending cases really cluster.

10-90 risk band Sample cases Required hurdle
Capital & delivery risk · interactive prototype
Problem
A review team needs to know how often an uncertain plan clears its required hurdle.
Implementation
Simulate compounding paths, compare the central outcome with the 10–90% range, and resample to inspect tail sensitivity.
Deliverable
A risk review brief with assumptions, hurdle probability, finish distribution, and the current decision narrative.
Model & validation boundaries

Independent Gaussian step shocks on a normalized index. Random resampling changes results; this is a scenario model, not a calibrated market forecast or a task-level delivery schedule.

A balanced review setup where the hurdle still feels reachable, but the tail risk is visible enough to force a real decision.

Quarter-close risk review

A lead is checking whether the plan still clears its hurdle before committing more capital, timeline, or scope.

Variance can overpower the base case

The center line may look calm, but a few bad shocks widen the tail quickly and change the story for stakeholders.

Adjust before review day

Lower the hurdle, extend the horizon, or reduce exposure and scope if the success odds drift too low.

Scenario Paths

How the forecast can unfold

Preparing simulation...
Distribution

Where the ending scenarios cluster

Target not set
Applied decision projects

Experiment readiness, AI capacity, and segmentation

Work through a launch review, a capacity planning conversation, or a segmentation investigation. Change the assumptions, inspect the output, and export a decision brief for the next review.

Experiment Design Studio

A/B Test Power Simulator

Model the decision pressure behind a launch review: how much sample, how much noise, and how much real lift you need before a "winner" deserves trust.

Experiment launch review · interactive prototype
Problem
A product team must decide whether traffic and expected lift support a credible experiment read.
Implementation
Compare baseline conversion, relative lift, sample allocation, variance, and daily traffic under a fixed-horizon power approximation.
Deliverable
An experiment readiness brief with estimated power, error tradeoffs, duration, and a review recommendation.
Model & validation boundaries

Illustrative two-group conversion model. Estimates assume fixed-horizon testing; repeated peeking, interference, and metric bias require separate checks.

Evidence Map
Null vs. uplift distribution
Calibrating read...
Runtime Pressure
How long you need to sit on the test
Timing pending
1 day 1 week 2 weeks 1 month+

Short tests feel faster, but they usually buy speed by borrowing confidence from the future.

Applied AI Economics

LLM Cost-Latency Simulator

Stress-test an AI system the way a platform lead would: traffic, context size, retries, and cache behavior all fight over the same latency and budget envelope.

AI service capacity planning · interactive prototype
Problem
A platform team needs a budget and latency envelope before scaling a copilot or tool-using agent.
Implementation
Compare model tiers, traffic, input/output tokens, caching, and retries to expose the largest cost and latency drivers.
Deliverable
A capacity planning brief with monthly spend, estimated P95 latency, failure risk, and cache savings.
Model & validation boundaries

Synthetic model profiles and heuristic latency/failure estimates. These are not current vendor prices or measured service benchmarks; calibrate against real traces before release.

Cost Stack
Where the monthly spend really goes
Illustrative economics
Latency Budget
How queueing and retries bend the p95
Health pending
Representation Learning Studio

UMAP / HDBSCAN Manifold Simulator

Compress a synthetic high-dimensional population into a neighborhood map and watch density structure survive, split, or dissolve as overlap, local scale, and minimum cluster size shift.

Segmentation investigation · interactive prototype
Problem
An analyst needs to distinguish stable neighborhoods from an attractive but misleading cluster picture.
Implementation
Explore a seeded synthetic population using local-neighborhood projection and density threshold clustering, then stress-test overlap and cluster size.
Deliverable
A segmentation review brief with neighborhood retention, detected groups, noise, and sensitivity notes.
Model & validation boundaries

A UMAP-inspired projection and DBSCAN-style density approximation, not the UMAP or HDBSCAN libraries. Synthetic labels provide an internal check; real data needs stability and domain review.

Embedding Surface
How local neighborhoods fold into 2D
Projection pending
Density Frontier
Where dense structure becomes noise
Cluster scan pending

Investing

Notes on capital allocation, market structure, and the quieter parts of long-term decision-making.

Essays & References

A quieter corner for essays, references, and ideas that inform how I build.

  • Why "Taste" Matters in Science — and in Technology

    Exploring Nobel laureate Yang Zhenning's concept of 'taste' in research and how it applies to technology and product development. What separates the merely competent from the truly visionary in science and tech.

  • Interesting Resource: Calculating Empires

    I recently discovered an fascinating interactive resource called "Calculating Empires: A Genealogy of Technology and Power Since 1500". This comprehensive visualization maps out the intricate relationships between technology, power, and human history over the past 500 years.

  • Knowledge Flow

    An interactive platform for visualizing and exploring connected knowledge across various domains. Knowledge Flow helps discover relationships between concepts and ideas in a structured format.

Contact