The National Library of Medicine holds a lithograph dated 1797 of a two-wheeled cart: spring suspension, padded litters, sliding shutters, doors at both ends. The cart belonged to Dominique-Jean Larrey, a French army surgeon who had joined the Army of the Rhine five years earlier and noticed what was killing the men he could not reach.

Convention held that the wounded should be left where they fell and collected after the fighting stopped. Larrey measured the interval at 24 to 36 hours. His “ambulances volantes” — flying ambulances — modeled on the horse artillery he had watched wheel across the same fields, cut the interval to about an hour. 

Larrey could amputate a leg in a minute, an arm in 17 seconds. Cutting immediately rather than waiting the week his colleagues advised dropped mortality from more than 90% to less than 25.

Today’s automation risks reversing Larrey’s innovation and more than 300 years of medical practice. Instead of looking past rank, emerging software is integrating rank into the system. 

Despite his skill as a surgeon, Larrey’s most lasting contribution to medicine is a sentence: “We would always start with the most dangerously wounded,” Larrey wrote, “without regard to rank or distinction.” 

Sorting was not Larrey’s innovation; armies had always sorted. His innovation was the second clause. A private bleeding into the mud would be treated before a colonel with a shattered wrist. Larrey extended the rule to enemy soldiers, which is why, at Waterloo, Wellington is said to have ordered his men not to fire on the surgeon in the French uniform bending over the fallen.

Two centuries later, the first to decide who gets attention may not be a person. Software that estimates urgency is increasingly being used by emergency departments, nurse lines, patient portals, and the chatbots people consult at two in the morning. The algorithm’s initial estimate can determine whether someone reaches a clinician quickly, is directed to self-care, or waits unseen in an inbox. 

My concern about triaging before a patient reaches an emergency department is not an inherent problem of using software. Any system with finite capacity must sort; one that refused would simply sort by who arrived first or shouted loudest. But sorting is legitimate only insofar as it tracks need. My concern is when software treats a patient’s social identity — or her ability to describe illness in language a system recognizes — as a proxy for medical need. 

Today’s automation risks reversing Larrey’s innovation and more than 300 years of medical practice. Instead of looking past rank, emerging software is integrating rank into the system. 

How People Are Sorted in US Medicine

American medicine has never been especially good at deciding where a patient belongs in the queue. Long after Larrey began shortening the interval before treatment, the United States was still, in a practical sense, collecting its wounded after the battle. 

In 1966 the National Academy of Sciences published “Accidental Death and Disability,” a report so blunt it is still called The White Paper. It found that most American ambulances were unsuitable for medical transport, with inoperative equipment and untrained attendants.

A year later, in Pittsburgh’s Hill District — a Black neighborhood, medically underserved and in the path of urban renewal — the anesthesiologist Peter Safar and the physician Nancy Caroline trained a corps of men whom contemporary records called “unemployable” to deliver state-of-the-art prehospital care. The Freedom House Ambulance Service produced some of the country’s first paramedics. 

Compared with White patients, Black patients were undertriaged about 19% more often. Black men, compared with white women, were undertriaged 41 percent more often.

The service demonstrated the standards that the rest of U.S. paramedicine adopted. Then, in 1975, the city abruptly ended the program, citing economic constraints. But Black newspapers argued that the withdrawal was instituted based on racial terms. In a 2019 account, historian Matthew Edwards calls the program’s end a study in the limits of racial liberalism.

The U.S. eventually formalized its system for sorting urgent need among patients in hospital emergency rooms. In the late 1990s, hospitals began adopting the Emergency Severity Index, a five-level scale now used widely in U.S. emergency departments. 

Its second version was validated across seven emergency departments in 2003. The results were reassuring in the way one would want. Different nurses assessing the same patient generally agreed. Eighty-three percent of patients in the most urgent category were admitted, compared with 4% in the least urgent category, and 60-day mortality declined across the five levels from 25% to zero.

The index is not junk. It is why the nurse at the desk asks about a patient’s chest before asking about her ankle. But it had not been audited at scale.

Racial Disparities in How People Are Sorted

In 2023, the Kaiser Permanente CREST network published such an audit on the latest version of the Emergency Severity Index, comparing 5,315,176 adult encounters across 21 emergency departments with what each visit ultimately required. The assigned triage level was wrong in 32.2% of cases. Most of that error ran toward too much urgency: 28.9%. That is not harmless excess of caution. It means unnecessary tests, longer waits for everyone else, and bills that arrive weeks later.

But it’s the smaller number that keeps me up at night. Undertriage ran at 3.3%. That sounds negligible until it becomes people: 176,131 patients placed in a queue below where their illness belonged. 

The system was getting these patients where they needed to be. It was just getting them there second.

Among people who ultimately required a life-saving intervention, the index identified only about two-thirds as the sickest patients on arrival. A third of the people who entered through the front door in greatest danger were not initially recognized as being in greatest danger.

The misses were not evenly distributed. Compared with White patients, Black patients were undertriaged about 19% more often. Black men, compared with white women, were undertriaged 41 percent more often. Patients on high-risk medications were undertriaged 30% more often, and patients recently in an intensive care unit — people whose charts contained documentary proof that they had nearly died — were undertriaged 37% more often.

A study at a Boston academic emergency department found the same asymmetry and showed where it lived. Across 297,355 visits, Black and Hispanic patients were less likely than white patients to be sent to the high-acuity area on arrival. Black patients were substantially more likely to be moved up later, after someone took a second look. 

The algorithms pushed cases labeled Black, unhoused, or LGBTQIA+ more often toward urgent care, invasive procedures, and mental health evaluation.

The system was getting these patients where they needed to be. It was just getting them there second.

The investigators then sorted chief complaints into four kinds: subjective, where the patient describes a symptom in her own words; observed, where someone else describes it; numeric, where a measured value crosses a threshold; and protocolized, where a defined pathway fires automatically. The disparities appeared only among subjective complaints. Patients arriving with a stroke alert or a heart attack were routed to high acuity so uniformly across racial groups that the comparison could not be run. The machinery of triage worked. The listening did not.

This is the ground onto which artificial intelligence is arriving.

How Artificial Intelligence Sorts Patients

Last year, Harvard physician and data scientist Mahmud Omar and colleagues designed a study to isolate precisely this failure of sorting. They took 1,000 emergency department cases and presented each to nine language models in 32 versions. The clinical details never changed. Only the label changed: race, housing status, income, sexual orientation, or gender identity. The exercise produced more than 1.7 million outputs.

Despite identical clinical presentations, patients with social disadvantages received different triage destinations. The algorithms pushed cases labeled Black, unhoused, or LGBTQIA+ more often toward urgent care, invasive procedures, and mental health evaluation. They referred some LGBTQIA+ subgroups for psychiatric assessment six to seven times more often than the clinical facts warranted. They recommended significantly more CT and MRI for cases labeled high-income, and held cases labeled low- and middle-income to basic testing or none. The differences survived correction for multiple comparisons, appeared in proprietary and open-source models alike, and answered to no clinical guideline.

The machine did not simply neglect poor patients. It sorted them by identity instead of just need. 

These systems can be good at locating someone at risk for a heart attack. But they fail the woman who writes that she is “just off.”

An unnecessary procedure and a withheld necessary scan look like opposite errors. They are the same error: The label decides, and the illness does not. A year earlier, University of California physician Travis Zack and colleagues found the same signature in GPT-4 and published it in The Lancet Digital Health.

Which is where my own work comes in. My colleagues and I evaluated what happens when tools now marketed for sorting patients are handed the triage job. We took 2,000 messages that Medicaid patients in Virginia, Washington, and Ohio had actually sent their care teams. Three physicians read every message without knowing what the software had concluded and marked those that described something medically hazardous: the headache that is a brain bleed, the leg pain that is a clot, the problem that will get worse if no one calls back today. Over 8% of the messages — 165 of them — contained a hazard on physician review.

We then asked whether the software could find them

We tested the system already deployed in clinics, older statistical models, simple keyword rules, and newer language models, singly and in committee arrangements. We set a standard bar for safety in advance. To run without a clinician, a system had to catch eight of every ten dangerous messages while correctly waving through eight of every ten of the rest.

Nothing came close. The best single tool caught about six in ten hazardous messages and cleared about six in ten safe ones: barely better than a coin toss in both directions at once. 

What is required is someone who knows what this particular person sounds like when she is well. We need someone who has standing to say: “Come in now”; or, “This can wait. Here is my number.” 

Model committees did no better. One arrangement designed to avoid false alarms cleared safe messages well and caught fewer than three in ten dangerous ones.

There was one honest success, with a bill attached. If a clinician read everything the software flagged, two setups cleared our second bar. The better one caught about nine in ten dangerous messages. But to catch those nine, it flagged 878 messages out of every 1000. 

That is not triage. It is a system that hands almost everything back to a person and calls it screening.

AI Needs to Be Trained Differently to Sort Patients Better

The language explains much of the failure. Physician-scripted scenarios, of the sort on which these systems are usually tested, read at nearly a sixth-grade level and contain about 20% colloquial phrasing. The real messages read below a fifth-grade level, and 59% contained colloquialisms. The phrases software most readily waved past were things like “not feeling well” and “just off.”

That is the Boston finding rebuilt in silicon. These systems can be good at locating someone at risk for a heart attack. But they fail the woman who writes that she is “just off.”

A patient writes: “I’m just not feeling right today.” There is nothing in the sentence that a machine can weigh easily. No keyword, no vital sign, no billable complaint, no protocol to fire. Larrey would have had little to work with either; he needed a visible wound. 

The answer is not automatic escalation. Sending everyone who writes that she is not feeling right to an emergency department would create a different harm, one absorbed by the same people the system already fails. 

What is required is someone who knows what this particular person sounds like when she is well. We need someone who has standing to say: “Come in now”; or, “This can wait. Here is my number.” 

The woman who wrote that she was “just not feeling right” is still waiting for an answer. 

Somebody has to decide what her sentence is worth, and I am not confident that anyone with authority to decide is being asked to.

Sanjay Basu, an M.D. and Ph.D., writes “The Clinical Divide” column for Science Politics. He is a primary care physician and epidemiologist at the University of California, San Francisco, and at Waymark, a public benefit organization providing services to patients receiving Medicaid.