Pharmaceutical Trials Optimized for Approval

The Hook

The trial worked beautifully. 847 patients with moderate-to-severe disease, randomized, double-blinded, placebo-controlled. The drug outperformed placebo with p < 0.001. The FDA approved it. The press release said “breakthrough.”

The prescribing begins. Doctors give it to their patients — patients who are older than the trial subjects, sicker than the trial subjects, taking more medications than the trial subjects, and living lives more complicated than a clinical trial protocol permits. The drug works less well. Significantly less well. The side effects are more frequent. The benefit is smaller. The doctors see the gap between what the trial promised and what the prescription delivers.

The doctors cannot reconcile the gap. The trial evidence says it works. Their clinical experience says it works less well than the evidence suggests. The evidence outranks the experience. The evidence is the trial. The trial was designed to produce the evidence.

The system is grading its own homework.


The Conventional Frame

Clinical trials are the gold standard of medical evidence — and correctly so. The randomized controlled trial controls for bias, isolates the treatment effect, and produces evidence that is more reliable than any alternative. The system works.

The system also has a specific, structural optimization pressure. A Phase III trial costs $50-100 million. The company must recoup the investment through FDA approval. The rational response to a $50 million bet: maximize the probability of winning. Every design choice that increases approval probability is a good business decision. The problem is not that companies are dishonest. The problem is that the system’s incentive structure (approval = revenue) produces designs optimized for approval rather than for the population that will actually receive the drug.

Design choices that maximize approval probability: narrow inclusion criteria (select patients most likely to respond — younger, healthier, fewer comorbidities), placebo comparison rather than active-comparator (easier to beat placebo than to beat the best existing treatment), surrogate endpoints (lab values that correlate with clinical outcomes but can be measured faster), short duration (the minimum required for statistical significance), and enrichment designs (pre-screening for likely responders before randomization).

Each choice is legitimate. Each is approved by the FDA. Each is standard practice. And each introduces a specific, predictable gap between the trial population and the real-world population.


The Reframe

The system is self-sealing because the EVIDENCE BASE — the only instrument medicine has for evaluating treatments — is produced by the system in a way that systematically inflates the apparent efficacy.

The trial selects for the best responders the trial produces positive results the results become the evidence the evidence is used to prescribe to a broader population the broader population responds less well the gap between evidence and reality widens but the EVIDENCE (which is the trial results) does not change the evidence continues to say the drug works more prescription more gap

The seal: a clinician who says “this doesn’t work as well as the trial suggested” is told “the evidence says it works.” The clinician’s real-world observation — which IS the signal about the efficacy-effectiveness gap — is classified as “anecdotal” and overruled by the trial evidence. The hierarchy (trial > observation) is defensible in general. It becomes a seal when the trials are systematically optimized to produce the result the hierarchy trusts.

The system cannot self-correct because the correction signal (real-world underperformance) is LOWER in the evidence hierarchy than the confirmation signal (trial results). The instrument that would detect the problem is ranked below the instrument that produces it.


The Scores

Factor Score Justification
F1: Mortality & Irreversibility 7 Patients receiving drugs that are less effective than the evidence suggests experience treatable conditions undertreated
F2: Scale 9 Every drug approved through the current trial system; every patient prescribed based on trial evidence
F3: Compression Depth 5 The compression is diffuse — slightly less benefit per patient across millions of patients
F4: Time Sensitivity 7 The evidence base accumulates with each new optimized trial; the gap widens with each cycle
F5: Voice Deficit 5 Clinicians can speak but are outranked by the evidence their observations contradict
F6: Proximity Gap 7 Trial design optimization consultants and real-world evidence researchers are in separate rooms
F7: Temporal Displacement 5 The gap is present immediately upon prescription but invisible without real-world comparison data
F8: Normalization 7 “Evidence-based medicine” normalizes the hierarchy that seals the circle
F9: Hallway Dependency 7 Breaking the seal requires trial design + real-world evidence + regulatory reform in conversation
F10: Knowledge Readiness 7 Real-world evidence methodologies are mature; the integration with regulatory standards is the gap
F11: Entry Cost 5 Regulatory change is slow; real-world evidence studies can be conducted independently
F12: Cascade Potential 8 The self-sealing evidence-base problem applies to every domain that uses controlled studies as its primary evidence standard

Hiddenness Score: 44.6 Actionability Score: 48


The Collision Partners

Real-world evidence (RWE) researchers measure what happens after approval. Their data consistently shows smaller effects than trials. They are treated as supplementary evidence rather than as a correction to the trial evidence. The specific transferable knowledge: the gap between trial efficacy and real-world effectiveness is MEASURABLE and SYSTEMATIC — it follows predictable patterns based on how narrow the trial’s inclusion criteria were, whether the comparator was placebo or active treatment, and how short the trial duration was. These patterns could be used to ADJUST trial results — producing an “expected real-world effectiveness” estimate alongside the trial efficacy number.

Insurance actuaries bear the cost of the gap. They pay for drugs based on trial-predicted efficacy. They experience real-world effectiveness. They have the financial data that quantifies the gap — the difference between the cost-effectiveness the trial predicted and the cost-effectiveness the real world delivered. The specific transferable knowledge: the actuarial data IS the correction signal the system lacks. If insurers published the gap between expected and actual cost-effectiveness for each drug, the evidence base would have a second instrument — one that is independent of the trial system and that measures what actually happens.


Where to Start

If you are a regulatory authority: require that every drug approval include an EXPECTED REAL-WORLD EFFECTIVENESS estimate alongside the trial efficacy number. The estimate is calculable from known predictors of the gap:

The two numbers — trial efficacy and expected real-world effectiveness — published side by side would make the gap visible for the first time. The gap is currently invisible because only one number is reported.

If you are a clinician: when a drug performs differently in your practice than the trial suggested, report it — not to the pharmaceutical company but to a real-world evidence registry. The FDA’s Sentinel System, IQVIA, and Flatiron Health collect post-market outcomes data. Your clinical observation — “this drug works less well in my patients than the trial predicted” — becomes actionable data when aggregated across thousands of clinicians. The hierarchy that ranks your observation below the trial evidence is the seal the circle operates through. Your data, accumulated, is the instrument that breaks it.


The Circle

Tier 3, self-sealing — The evidence base is produced by optimization that inflates it.

Trials optimized for approval efficacy overpredicts real-world performance underperformance attributed to patients trial design maintained metrics confirm system works

The circle begins with a $50-100 million bet. A pharmaceutical company has invested years and vast resources in developing a drug. The Phase III trial is the final test. The rational business decision is to maximize the probability of approval. Every design choice that increases approval probability is a good investment: narrow inclusion criteria that select the patients most likely to respond, placebo comparators that are easier to beat than the best existing treatment, surrogate endpoints that can be measured faster, and short trial durations that capture initial efficacy without waiting for effectiveness to decay. Each choice is legitimate. Each is FDA-approved. Each widens the gap between what the trial measures and what the real world delivers.

The drug is approved. Doctors prescribe it to their patients — patients who are older, sicker, taking more medications, and living more complicated lives than the trial’s carefully selected participants. The drug works less well. The effect is smaller. The side effects are more frequent. The gap between trial efficacy and real-world effectiveness is consistent and documented across drug classes. But the correction signal — the clinical observation that the drug underperforms its evidence — is ranked below the evidence itself in the hierarchy of medical proof. The trial is a randomized controlled study. The clinical observation is “anecdotal.” When the two conflict, the trial wins.

The system seals itself through this hierarchy. When real-world underperformance is observed, it is attributed to the patients (they are non-compliant, they have comorbidities, they take concomitant medications) rather than to the trial design (which selected against exactly these real-world conditions). The trial design is maintained for the next drug. The next trial produces the same inflation. The evidence base accumulates — each entry slightly inflated by the optimization, each real-world shortfall attributed to patient factors rather than evidentiary factors. The system grades its own homework and gives itself high marks while the patients in the real world receive treatments that perform below their published evidence.

What breaks it is publishing expected real-world effectiveness alongside trial efficacy, calculated from the known predictors of the gap: how narrow were the inclusion criteria, what percentage of real-world patients would have been excluded, was the comparator placebo or active treatment, was the duration sufficient to capture real-world patterns? These adjustments are calculable from existing data. Two numbers published side by side — what the trial showed and what the real world can expect — would make the gap visible for the first time, and visibility is what the seal prevents.