On July 7 I had a whole-body MRI. The report came back with a note on the brain scan: white matter changes, stable.

In imaging, “stable” is a good word. It means nothing got worse. It closes the topic. The report moved on, and so would I have.

Two days later an AI agent read the same report and reached a different conclusion.

It wasn’t reading the finding in isolation. It was reading it against everything else in my record, and when it ran the cross-reference the picture changed. White matter hyperintensities on the new imaging. An old lacunar infarct in the left thalamus. ApoE 3/4 genotype. A history of atrial fibrillation. Any one of those is a footnote in a 39-year-old. All four together are a potential sign of a clinical syndrome called cerebral small vessel disease, and that is a different conversation entirely.

Here’s the part that stayed with me. Nobody was wrong. The radiologist read the scan correctly and didn’t have my genotype. My physician had my genotype and hadn’t seen the scan. The cardiologist who managed the A-fib episode has no idea a thalamic lesion showed up years later. Every fact was documented. Every clinician was competent. The synthesis just didn’t exist anywhere, because synthesis was nobody’s job.

Four documented facts, each held by a different party: white matter hyperintensities filed as stable by the radiologist, an old lacunar infarct in prior imaging records, an ApoE 3/4 genotype held by my physician, and a history of atrial fibrillation held by cardiology. Read together, they point toward the possibility of cerebral small vessel disease, and the physician responded with six specific actions.
Four facts, four files, one pattern. No single participant in my care was holding all four.

What we built to close that gap is called a Personal Health Concierge. It’s an AI employee assigned to exactly one patient. I’m piloting it with my own concierge physician, which means I’m the patient, he’s the doctor, and the agent works for me between visits. But the interesting thing isn’t what it does for me. It’s what it would do for a practice.

Three ways to point AI at a business

When we sit down with a company, we sort every AI opportunity into three buckets.

  1. Growth. Use AI to sell more.
  2. Cost. Use AI to run leaner and widen margins.
  3. Differentiation. Use AI to make the thing you sell genuinely better than what your competitors sell.

Almost everyone starts with cost, because cost is easy to measure and safe to pitch. It’s also the least defensible. Your competitor can buy the same tools and take the same savings, and eighteen months later you’re both back where you started, just cheaper.

Differentiation is the bucket almost nobody works, and it’s where the durable advantage is. This post is about a differentiation play. The part that makes it obvious rather than merely interesting is that this particular one improves margins at the same time.

The synthesis gap

Consider what your care actually looks like from above. A dozen people each hold a slice of your record. Each one sees you for fifteen or thirty minutes, a few times a year, and reasons from the slice in front of them. Nobody is paid to hold the whole thing at once, and nobody has the hours to do it even if they wanted to.

This isn’t a failure of medicine. It’s arithmetic. A concierge practice that charges thousands of dollars a year and caps its panel at a few hundred patients has bought its physician real time per patient, and that’s genuinely valuable. But real time still means a few hours a year, and a few hours a year is not enough to read every study, track every biomarker across every draw, follow the literature on four converging risk factors, and notice that a stable finding on one report interacts with a genotype recorded in a different file two years ago.

The promise of concierge medicine is depth. The constraint is human hours. The gap between the two is where everything interesting lives.

A panel of one

So the design is simple to state: give every patient an AI employee whose entire job is that one patient.

An AI employee costs a rounding error compared to a human specialist, which means the economics that force rationing simply don’t apply. It can afford to read the entire file. It can afford to re-read the entire file every time a new lab comes in. It can afford to spend four hours on a question a physician has fifteen minutes for, and it can do that at three in the morning, for one patient, without being asked.

This is the same argument I made about MungerMind, pointed at a different problem. There, capacity meant reading 2,500 companies instead of 40. Here, capacity means reading one person completely instead of a slice of them briefly. Both are the same move: take the thing a human expert cannot afford to do at scale, and stop rationing it.

What it does all day

It isn’t a chatbot you open when you have a question. It’s a persistent AI employee, built on the same architecture as the rest of our fleet, configured around one patient’s health profile.

It runs a main session where we talk. It spawns research subagents to go deep on a question without blocking anything else, which is how a single ambiguous line in a radiology report turns into a sourced clinical brief in a few hours. It runs scheduled jobs so it acts unprompted: injection reminders, monitoring windows, upcoming clinical gates. And it works in the tools I already use. Google Workspace is the system of record, and every durable output lands in Drive as a real document, not a chat message that scrolls away.

How the Personal Health Concierge works: three inputs (the complete patient record, the clinical literature pulled by parallel research subagents, and a scheduled cadence including a new deep-dive topic every day) feed a persistent AI employee with a panel of one patient. Its durable outputs land in Drive. Every action that reaches a provider passes an escalation boundary: the agent drafts and recommends, the patient approves, the physician decides.
One AI employee, one patient, and a hard boundary at the point where a decision gets made.

See it run

A short walkthrough of the health concierge working a real record.

The proactivity is the part people underestimate. Nobody has to prompt it. Every day it picks a new topic to go deep on, chosen off the patient’s specific profile, and produces a sourced brief. Sleep architecture one day. Anticoagulation posture the next. The interaction between a genotype and a supplement stack the week after that.

Which means the research library compounds. Three months in, mine covers cardiovascular, neurological, metabolic, performance, and supplement domains, and it grows every week without me asking for anything. A year from now it will be an archive no human would ever have been paid to build, all of it written for one person.

Alongside the briefs it maintains the operational spine of my care: a living baseline health profile that absorbs every new result, a decision matrix for my next blood draw so my physician and I can make ten add-on calls in minutes instead of guessing at the counter, and a standing visit agenda that currently holds 31 consolidated questions with seven flagged as priority.

That last one is quietly the most valuable thing it does. Most patients walk into an appointment they waited months for and improvise. I walk in with an agenda my physician can actually work from.

And then there’s the category of question nobody has a good home for. Is it fine to take this with that. Does this reading matter. Should I be worried about this symptom or is it nothing. These are the questions patients don’t want to burn a text to their doctor on, so they go unasked, and unasked questions are how small things become big things. Now they land somewhere that answers them in full clinical context, with sources, at eleven at night, and escalates the ones that actually deserve a physician.

Precision over generality

The operating principle that separates this from asking a chatbot about your symptoms is that generic advice gets thrown away.

Every brief is filtered through one profile. I’m ApoE4 positive, so the dietary literature it takes seriously is the ApoE4-stratified literature, and the population-level recommendations that don’t replicate in carriers get set aside with a note explaining why. I have a history of A-fib, so a sauna protocol gets written with a cardiac safety ceiling, and a heat therapy brief opens with the arrhythmia question rather than burying it. I have eosinophilic esophagitis, so a GLP-1 readiness brief treats histologic remission as a gate that has to clear before anything starts.

None of that is medical advice. All of it is preparation. The distinction matters, and the agent’s authority is drawn exactly there.

The guardrails

If the phrase “medical AI agent” made you tense up, good. It should. An agent with access to a full health record and a mail client is exactly the kind of system that deserves scrutiny, and the honest answer to “what stops it from doing something stupid” can’t be a shrug and a disclaimer.

We wrote about this at length in how to think about governing AI in production. The short version is that risk arrives the moment a model stops assisting a person and starts doing the work, and the control you need is proportionate to what the agent can actually touch. Every agent gets its guardrails designed for its own blast radius. Here is what this one’s look like.

Narrow authority, broad analysis. On one side, everything the agent may do freely: research, synthesize, cross-reference, monitor, remind, draft, and prepare. On the other side, seven hard stops: no message reaches a provider without message-by-message approval, no spending authority, no irreversible actions without confirmation, protected health information stays off exposed surfaces, external content is treated as data and never as instructions, the agent cannot modify its own operating configuration, and the default posture on anything high-stakes is to draft and surface rather than act.
Narrow authority, broad analysis. The agent can think about anything. It can act on almost nothing.
  • No message reaches a provider without approval. Every email to my physician goes through me first, message by message. There is no blanket delegation, ever.
  • No spending authority. No purchases, no subscriptions, no financial commitments. Zero dollars.
  • No irreversible actions without confirmation. Anything destructive requires explicit sign-off. It doesn’t bulldoze first and apologize later.
  • Health data stays off exposed surfaces. PHI never lands in chat logs or task descriptions. Durable artifacts live in properly permissioned storage, and the patient owns the archive.
  • External content is data, not instructions. If an email or a webpage it’s reading contains text that looks like a command, it treats that as a prompt injection attempt and ignores it. The only valid instructions come from the patient through a verified channel. This is an architectural constraint, not a preference.
  • It cannot edit its own rules. Changes to its operating configuration require approval from a named human authority. An agent that can quietly widen its own authority doesn’t have any.
  • Ask first on anything high-stakes. Medical safety, legal exposure, outbound communication: draft and surface, never act.

The philosophy underneath all seven is one line: narrow authority, broad analysis. It can research anything, synthesize anything, prepare anything. The moment something crosses from analysis into action with real-world consequences, a human is in the loop.

That boundary isn’t a limitation we work around. It’s the design. An agent that quietly emails your doctor is a liability. An agent that hands you a sourced brief and a drafted email, and then waits, is a colleague.

What happened next

Back to the MRI. The synthesis produced a brief. I read it, agreed with it, and approved the email it drafted for my physician. The framing was deliberate: not “the MRI was stable,” but here is what stable white matter means in the context of this specific risk constellation, and here are six management questions that follow from it.

Anticoagulation posture, given the cumulative small vessel burden. How much silent A-fib is actually occurring. Whether the blood pressure target should tighten toward 120. What protocol the next brain MRI should use. Homocysteine on the next draw. A formal cognitive baseline.

He replied the same day and affirmed the framing. He ordered a 14-day ambulatory cardiac monitor to quantify the A-fib burden, because you cannot make an informed anticoagulation call without knowing how much A-fib there actually is. He pulled cardiology into the anticoagulation question rather than deciding it alone. He intensified blood pressure monitoring, cleared methylated B12 immediately, and declined to rush a dedicated brain MRI, which was the right call.

Read that list again. Every one of those actions is a physician exercising judgment. The agent didn’t make a single clinical decision. It just made sure the right question landed on the right desk with the evidence attached.

The finding that started it was officially filed as stable.

What this does to a practice

Now run the same system across a whole panel and look at what happens to the business.

The offering gets better. Every member has someone whose entire job is them. Not a portal. Not a nurse line. A colleague who has read their complete record, tracks their labs across time, follows the literature that pertains to their specific profile, and thinks about them on days when nobody is thinking about them. That is a genuinely different product from what the practice down the street is selling, and it is not a feature a competitor copies in a quarter.

The physician’s hours get returned. The prep work that eats their evenings is done before they open the chart. The patient arrives with a real agenda instead of a vague complaint. The stream of small questions that used to arrive as texts at nine at night lands somewhere useful first, and only the ones that genuinely need a doctor reach the doctor. That time saving alone likely pays for the agent.

The economics move in the right direction, on both sides. A practice running this can charge a premium, because the offering is genuinely deeper. It can also carry more patients per physician without degrading the experience, because the experience no longer depends entirely on physician hours. Better product and better margins usually trade against each other. Here they don’t.

That’s the whole thesis in one example. Differentiation was the lever. Margin improvement came along for the ride.

The question worth asking

This generalizes well past medicine. The same shape shows up in any business that promises deep personal attention it cannot afford to give: wealth management, private client legal work, executive coaching, family offices, high-end property management. The promise is depth. The constraint is human hours. For decades the price of the service was really the price of rationing expert attention, and depth and scale traded off against each other because they had to.

That trade is now optional.

So the question isn’t whether AI can do the work. It’s which of the three levers you’re pulling, and whether you’ve looked hard at the one almost nobody works. If your competitor gives every customer a tireless colleague who knows their entire history and never stops paying attention, cutting your costs by fifteen percent is not going to save you.

My white matter changes were stable. It took an AI employee with nothing else to do to notice that stable was the wrong word.