Published
8 September 2026GAIF-FA-001Global AI ForumEvery Tuesday
FAILURE ATLAS · CASE GAIF-FA-001 · 1 OF 50 The model was not the problem. A cancer centre spent USD 62.1 million on a decision support system that never reached a decision. The audit took no position on whether the technology worked. It did not need to. SUBJECT Oncology Expert Advisor, built on IBM Watson COMMISSIONING BODY MD Anderson Cancer Center PERIOD June 2012 to February 2017 CLASSIFICATION Mode 1. No workflow CANCELLED 11 minute read · Evidence grade A $62.1 M Spent against a contract written for $2.4 million 12 Extensions to a six-month contract 52 Months from signature to expiry 0 Patients ever treated on its recommendation GAIF-FA-001 THE ATLAS SIX MODES THE RECORD WHAT HAPPENED FACTORS IBM'S POSITION APPLICABILITY EVIDENCE THE ATLAS What this is, and what it is for The Failure Atlas is a catalogue of fifty enterprise AI programmes that did not work, each reduced to a single structural cause and a single test a buyer can apply before signing. Most published AI case studies are written by people with something to sell. Success stories are marketing with a customer logo on them, and post-mortems are usually written by the losing vendor's competitor. Neither is much use to somebody holding a budget. This Atlas is written from the buyer's side and takes no money from anyone. It is deliberately narrow. Each case answers one question: what structural condition made this fail, and how would a buyer have detected it before spending? It is not a technology review. In most of these cases the technology worked. That is the finding, not a caveat. Method Only programmes with a public, documented record: audits, filings, regulatory findings, court papers, or first-party disclosure. Every figure traced to a named primary source. Where sources conflict, the conflict is shown rather than resolved. The subject organisation's own position is stated in its own terms, in every case, whether or not it is convincing. One primary failure mode per case. Contributing modes are listed separately and never blended into the headline. No case is included because a vendor is unpopular, and none is excluded because a vendor is a partner. No vendor paid to appear in this Atlas, and none paid to be left out of it. There is no sponsored tier, no review copy, and no right of reply beyond publishing the subject's stated position, which is done in every case. CLASSIFICATION The six modes Fifty programmes, six ways of failing. The modes describe structure rather than technology, which is why they survive changes in the technology. A case is assigned one primary mode and any number of contributing modes. MODE 1 · PRIMARY IN THIS CASE No workflow The system functioned and was never wired to the moment a human decides. Output arrives beside the decision rather than inside it. GAIF-FA-001 MODE 2 · CONTRIBUTING No data path The system cannot reach the data it needs from inside the institution, so humans carry the data to it by hand. GAIF-FA-001 MODE 3 No owner No named person is accountable for the decision the system was built to change, so no one can authorise the change in practice. MODE 4 · CONTRIBUTING Wrong bottleneck Visibility, prestige or ambition selected the use case rather than a measured constraint. The system solves something that was not costing anything. GAIF-FA-001 MODE 5 No baseline Nobody measured what the humans did before, so improvement can never be demonstrated and the programme is defended on anecdote. MODE 6 · CONTRIBUTING No stop condition No agreed circumstance under which the programme would be halted, so it is extended instead of judged. GAIF-FA-001 Modes are structural, not technical. A programme can fail on more than one and almost all do. The primary mode is the one whose removal would have changed the outcome. FINDING The model was not the problem. It was never connected to the moment an oncologist actually decides. A recommendation that arrives beside the decision rather than inside it is not a recommendation. It is a document. Sixty-two million dollars bought a document. CASE GAIF-FA-001 The record SUBJECT Oncology Expert Advisor (OEA) A Watson-powered clinical decision support tool, built with IBM for MD Anderson Cancer Center. Distinct from IBM's Watson for Oncology product, whose clinical content was developed with Memorial Sloan Kettering. The two are routinely confused. COMMISSIONING BODY University of Texas MD Anderson Cancer Center Houston, Texas. Part of the University of Texas System. PRINCIPAL VENDOR IBM With PricewaterhouseCoopers engaged on project support. SECTOR Healthcare. Clinical decision support, oncology. JURISDICTION United States PERIOD June 2012 to February 2017 Contract signed June 2012. IBM support withdrawn September 2016. Final contract extension expired 31 October 2016. Audit published February 2017. STATED OBJECTIVE To let community oncologists deliver care at the standard of MD Anderson specialists Described in the audit, paraphrasing the project's originator, as elevating the standard of cancer care worldwide. SYSTEM DEPLOYED Decision support over curated case data Version 1.0 covered a single condition, lower-risk myelodysplastic syndrome. Scope later grew to five further leukemias and lung cancer. No other cancer type was ever added. Disposition QUANTUM USD 62.1 million IBM approximately $39.2M. PwC approximately $23M. Original contract: $2.4M for six months, extended twelve times. OUTCOME Cancelled. Never used on a patient. The contract was allowed to expire. No clinical production deployment occurred at any scale. PRIMARY MODE Mode 1. No workflow. CONTRIBUTING MODES Mode 2, no data path. Mode 4, wrong bottleneck. Mode 6, no stop condition. EVIDENCE GRADE A. Independent audit by the parent institution, plus first-party statements from both sides. SECTION ONE What actually happened In June 2012 MD Anderson contracted IBM to build the Oncology Expert Advisor: a system that would read a patient's record, cross-reference the literature, propose treatment options and match patients to trials. The first version was scoped at six months and $2.4 million, covering one leukemia subtype. The contract was extended twelve times over the following four years. Scope grew to five further leukemias and lung cancer. IBM was paid approximately $39.2 million and PricewaterhouseCoopers approximately $23 million. The final figure was $62.1 million. EXHIBIT A The contract that became a programme Original scope against final outlay, USD millions Contracted $2.4M 6 months, one leukemia subtype Spent $62.1M 52 months, 12 extensions, never used on a patient 26x the contracted value. The scope grew to five further leukemias and lung cancer. No other cancer type was ever added. Source: University of Texas System Administration special review of procurement procedures, 2017. IBM $39.2M, PwC approx $23M. During the same period MD Anderson replaced its record system, moving from the legacy ClinicStation platform to Epic. OEA had been built against ClinicStation. It was never integrated with Epic, and it was never going to be without a rebuild that nobody funded. The practical consequence is the whole case. To use OEA, a clinician had to leave the record system where the patient's data lived, open a separate interface, and enter the data again by hand. Then read the output, and go back to the record system to act. EXHIBIT B The gap the money never crossed The clinical path, and where the system was actually connected Patient record Epic EHR Assessment Oncologist Decision Oncologist Order Epic EHR Oncology Expert Advisor separate interface, manual data entry never wired in Source: UT System audit, 2017; IEEE Spectrum, 2019. OEA was built against the legacy ClinicStation record system and was never integrated with Epic. In September 2016 IBM withdrew support, noting internally that the system was not ready for human investigational or clinical use. The last contract extension expired that October. In February 2017 the University of Texas System published a special review of the procurement. No patient was ever treated on the basis of an OEA recommendation. SECTION TWO Contributing factors Four conditions were present. Only the first is the primary mode; the others made it survivable for four years. 1 No defined clinical workflow At no point was there a specification of which clinician, at which point in which pathway, would consult the system, and what they would do differently as a result. The output had no destination. This is Mode 1, and removing it would have changed the outcome. 2 No data path into the institution The migration to Epic was a known, scheduled, institution-wide programme. A decision support tool built against the outgoing system, with no funded integration path to the incoming one, was obsolete on a published timetable. Manual re-entry is not an integration strategy; it is the absence of one. 3 Prestige selected the use case The objective was to elevate the worldwide standard of cancer care. That is an ambition, not a bottleneck. Nothing in the record identifies a measured constraint in the oncology pathway that the system was built to relieve. Mode 4. 4 Governance did not stop it The audit found procurement standards were not followed, contracts structured to sit below oversight thresholds, and gift funds committed before receipt. Twelve extensions to a six-month contract is not a series of decisions to continue. It is the absence of a decision to stop. Mode 6. The audit also noted a governance conflict: the project's originator was a senior faculty member married to the institution's then president. The Forum records this as a control weakness, not as a motive. SECTION THREE The subject's position IBM's stated position is that the system worked. A spokesperson said the recommendations were accurate, agreeing with expert opinion around 90 per cent of the time, and that the research and development project was a success which could likely have been deployed had MD Anderson chosen to take it forward. The Forum's view is that this is substantially correct, and that it is the most important sentence in the case. The audit took no position on the scientific basis or the functional capability of the technology. It described the difficulty of assimilating the system into the hospital. The independent reviewer, given four years of evidence, found the failure was not in the model. Two qualifications belong on the record. A 2018 peer-reviewed paper on OEA's information extraction reported accuracy of 90 to 96 per cent on clear concepts such as diagnosis, but 63 to 65 per cent on time-dependent information such as therapy timelines, which is the harder half of an oncology record. And IBM's internal September 2016 assessment, that the system was not ready for clinical use, sits uneasily beside its public defence. Neither qualification changes the finding. A system with a 90 per cent concordance rate and no route to a clinician produces exactly the same clinical benefit as a system with a 40 per cent concordance rate and no route to a clinician. Both produce none. SECTION FOUR Applicability This is the part a buyer can use. It is not advice about oncology or about IBM. It is a test that applies to any AI programme in any function, and it is designed to be failed early and cheaply rather than late and expensively. THE TUESDAY TEST Ask these three before signing. Answer them in writing. 1 Name the moment a human currently decides. Not the process. The moment. Which screen, which meeting, which minute. If the answer is a department or a workflow rather than a decision point, there is no destination for the output. 2 Name the person who owns that moment. A named individual who can change how that decision is made without asking anyone else. If the owner is a committee, the system will be evaluated rather than used. 3 Say what they will do differently on the Tuesday after go-live. In one sentence, in their words, describing an action they take or stop taking. Not "have better information". An action. If you cannot name all three, you are not buying a system. You are commissioning a document, and you are building this case again with a better model. The test is deliberately hostile to the pilot. Most enterprise AI pilots pass a technical evaluation and fail question three, because question three is the only one that requires somebody to change what they do. That is also why the test is cheap: it costs one meeting, and it can be run before any money moves. Where this case would have stopped Checkpoint What was known at the time What the test would have surfaced Contract signature, 2012 Objective stated as elevating worldwide cancer care No decision moment named. Fails question 1 before any money moves. Epic migration decision Institution-wide record system replacement scheduled The named decision moment moves to a system the tool cannot reach. Fails question 3. Extension 3 or 4 Still no clinical production use No stop condition exists. The absence is itself the finding. EVIDENCE Sources and grading Evidence grade A: an independent audit by the parent institution, supported by first-party statements from both the commissioning body and the vendor, and by peer-reviewed publication. PRIMARY University of Texas System Administration. Special review of procurement procedures related to the UTMDACC Oncology Expert Advisor project, published February 2017. Source of the quantum, the contract history, the twelve extensions and the procurement findings. FIRST PARTY IBM. Public statements on OEA accuracy and deployability, February 2017. Internal assessment of readiness for clinical use, September 2016, as reported in the audit. FIRST PARTY MD Anderson Cancer Center. Public statements, 2017. PEER REVIEWED The Oncologist , 2018. Reported OEA information extraction accuracy of 90 to 96 per cent on clear concepts and 63 to 65 per cent on time-dependent information. REPORTED The Cancer Letter and Houston Chronicle , February 2017, first reporting on the audit. Forbes , February 2017, carrying IBM's response. IEEE Spectrum , 2019, on the integration failure. What this case is not It is not an assessment of Watson for Oncology, the separate IBM product whose clinical content was developed with Memorial Sloan Kettering and which was sold internationally. Reporting on that product, including the 2018 investigation into unsafe recommendations, concerns a different system and is filed separately in this Atlas. Conflating the two is the most common error in secondary coverage of this case, and the Forum has treated it as a distinct matter. It is also not a verdict on clinical AI. Systems of this type are in production use today. The distinguishing feature is not the model. Corrections Sources differ on the start date, variously given as 2012 or 2013, and on the end date, given as 2016 or 2017. This Atlas dates the case from contract signature in June 2012 to publication of the audit in February 2017, and records the cancellation as October 2016. Corrections are published as versioned amendments on the artifact itself and never applied silently. Source. University of Texas System Administration special review, February 2017; IBM and MD Anderson public statements, 2017; The Oncologist, 2018; contemporaneous reporting as listed above. Research current as of 8 September 2026. Case 1 of 50. No vendor paid to appear in this Atlas, and none paid to be left out of it. The Global AI Forum holds no commercial relationship with IBM or MD Anderson Cancer Center. Failure Atlas · GAIF-FA-001 · September 2026 · v1.0 Case 1 of 50 No vendor paid to appear in this Atlas, and none paid to be left out of it.
Cite this asGlobal AI Forum, IBM Watson for Oncology at MD Anderson, 8 September 2026. GAIF-FA-001.
