Case studies
Breakout material for Why ML systems can fail in practice (Thu Aug 27).
Each group picks one case.
For each case: the system worked in some meaningful sense. Your job is to work out what failed anyway, and where in the lifecycle that failure lives.
- Case 1 · Reducing hospital readmissions (Camden, NJ)
- Case 2 · Screening job applicants
- Case 3 · Predicting sepsis in hospitalized patients (EPIC)
- Case 4 · Buying homes with an automated valuation
- Sources
Case 1 · Reducing hospital readmissions (Camden, NJ)
The goal. Reduce hospital readmissions among “super-utilizers” — patients with very high healthcare use.
The system. Identify the highest-risk, highest-cost patients from hospital data. For each one, a team of nurses, social workers, and community health workers visits after discharge to coordinate outpatient care and connect them to social services.
The reception. Nationally celebrated. Profiled in the New Yorker. Expanded to cities across the country. Before-and-after numbers looked strong.
The result. A randomized controlled trial of ~800 patients (NEJM, 2020) found no difference in 180-day readmissions between the program group and usual care — about 62% in both.
Your task
- The model identified high-risk patients accurately. So what failed?
- Why did the before-and-after numbers look good?
- Was the highest-risk patient the one whose readmission was most preventable?
- What would you change — in the model, the intervention, or the evaluation?
Case 2 · Screening job applicants
The goal. Automate the top of the hiring funnel: given a résumé, identify the strongest candidates for technical roles.
The system. From 2014, a model scored applicants one to five stars, “much like shoppers rate products.” It was trained on ten years of résumés submitted to the company. Gender was not an input.
What surfaced. By 2015 the team found the model was not rating candidates in a gender-neutral way. It penalized résumés containing the word “women’s” (as in “women’s chess club captain”), downgraded graduates of two all-women’s colleges, and favoured verbs more common in men’s résumés such as “executed” and “captured.”
The response. Engineers edited the model to be neutral to those specific terms. The team was disbanded around 2017. The tool was never rolled out broadly.
Your task
- Gender was never an input. How did the model learn it anyway?
- They removed the flagged terms. Why wasn’t that enough?
- What was this model actually trained to predict? Is that the same as “good candidate”?
- This was caught before reaching applicants at scale. What made that possible?
Case 3 · Predicting sepsis in hospitalized patients (EPIC)
The goal. Alert clinicians early when a hospitalized patient may be developing sepsis to prevent sepsis.
The system. A proprietary model built into a widely used electronic health record and deployed at hundreds of US hospitals. It rescored patients every 15 minutes and fired an interruptive alert above a vendor-recommended threshold. One of its inputs was whether a clinician had ordered antibiotics.
Vendor-reported performance. AUC 0.76–0.83.
Independent validation. 38,455 hospitalizations at one academic medical center; 7% developed sepsis.
| Metric | Result |
|---|---|
| AUC | 0.63 |
| Recall | 33% |
| Precision | 12% |
- Generated alerts on 18% of all hospitalized patients
- Of 2,552 septic patients, identified 183 (7%) whose sepsis was not already being treated in time
Your task
- Vendor 0.76–0.83, independent 0.63. How does that gap happen?
- What does an alert on 1 in 5 patients do to the people receiving alerts?
- What is the effect of using “clinician ordered antibiotics” as an input?
- It added value for 7% of septic patients beyond usual care. What does that say about the right baseline?
Case 4 · Buying homes with an automated valuation
The goal. Buy homes directly from sellers, make light repairs, resell at a margin (to make money)
The system. An established home-valuation model — mature, widely used, and reasonably accurate as an estimate — was used to decide which homes to buy and at what price. Purchases were made at scale.
What happened. Between April 2018 and November 2021, Zillow’s iBuying program purchased approximately 27,000 homes across 25 US cities. In Q3 2021 the company bought 9,680 homes and sold 3,032 Total program losses in 2021 reached $881 million. In November 2021 it shut the unit down, took a write-down of more than $500M, and cut about 25% of its workforce (~2,000 people).
Your task
- The valuation model was reasonably accurate. So what failed?
- What changes when an estimate stops informing a browsing customer and starts triggering a purchase?
- They were buying at volume in the markets they were forecasting. What does that do?
- Would a more accurate valuation model have prevented this?
Sources
Read these after the session
Case 1 — Finkelstein et al., Health Care Hotspotting — A Randomized, Controlled Trial, NEJM 2020 (free version) · Camden Coalition’s page on the trial
Case 2 — Reuters via CNBC · MIT Technology Review · ACLU on why it was predictable
Case 3 — Wong et al., External Validation of a Widely Implemented Proprietary Sepsis Prediction Model, JAMA Internal Medicine 2021 · 2024 replication, JAMIA Open · UT Austin case study