How a neural network is trained: a guide for hospital leaders
What a model is, what data it learns from, how it is measured, why performance drops in another hospital, and which questions to ask a vendor before signing.
A hospital director who receives an offer for an artificial intelligence system does not need to know how to write code. They do need to understand enough about how the model was trained and verified to ask the right questions and to recognise incomplete answers. This article explains the process without mathematics, with figures from published studies and with the obligations of the European framework.
What a model is and what it learns from
A neural network is a very large mathematical function with millions or billions of internal numbers called parameters or weights. At the start these numbers are random and the model knows nothing. Training is the process by which the parameters are adjusted, step by step, until the model gives correct answers on the examples it sees. The model does not memorise medical rules; it learns statistical correlations between images or laboratory values and the labels it is given.
Before anything else, the data is split into three sets with different roles:
- The training set: the material the model learns from; this is where the parameters are adjusted.
- The validation set: used by the team to compare variants and to decide when to stop training.
- The test set: put aside at the very beginning and used once, at the end.
If someone looks at the results on the test set and then changes the model, that set becomes, in effect, a second validation set, and the reported performance no longer says anything about new patients.
Labels: who assigns them and how much they agree
To learn, a model needs labels: every image or episode of care has to carry the correct answer ("referable retinopathy", "pneumonia present", "sepsis at hour X"). Clinicians assign the labels, and the quality of the model cannot exceed their quality.
A well-documented example is Gulshan et al. (JAMA, 2016) on diabetic retinopathy: 128,175 fundus photographs were labelled by 54 ophthalmologists and final-year residents, each image receiving between 3 and 7 independent grades, with the majority decision as the final label.
The reason several grades per case are needed is the limited agreement between specialists. Elmore et al. (JAMA, 2015) asked 115 pathologists in the United States to interpret breast biopsies independently, 6,900 interpretations in total. Concordance with the consensus reference diagnosis was 75.3%: 96% for invasive carcinoma, 84% for ductal carcinoma in situ (DCIS), 87% for benign lesions without atypia, and only 48% for atypia.
A model cannot be better than the labels it learned from. Ask who labelled, how many people were involved and how disagreements were resolved.
The training loop and overfitting
Training is a loop repeated millions of times. In plain words, each step has four moments:
- Forward pass: the model receives an example and produces a prediction, usually a probability.
- Loss: a number measuring how far the prediction is from the correct label.
- Backpropagation: the calculation that determines, for each parameter, in which direction it should be adjusted so that the loss decreases.
- Update: each parameter moves a little in the indicated direction. How far it moves is set by the learning rate: too large, and the model jumps past the solution; too small, and training takes a very long time.
One complete pass through the training set is called an epoch; a model is usually trained for tens of epochs, and after each one the team measures the loss on the validation set.
This is where the phenomenon that matters most to a buyer appears: overfitting. At some point the model starts memorising the particularities of the training examples instead of learning the general pattern. The sign is simple: the training loss keeps falling, but the validation loss stops improving and then rises. The chart below is an illustration with example data, not the result of a study.
A serious vendor can show this chart for their model; if they do not have it, they do not know where they stopped or why.
External validation: why performance drops in another hospital
A model that passes its internal test well has not yet proven anything about your hospital. Performance drops almost systematically on data from other institutions, with other equipment, other populations and other protocols.
Zech et al. (PLOS Medicine, 2018) trained a pneumonia detection model on 158,323 chest radiographs from three institutions. The model trained on Mount Sinai Hospital data reached an area under the curve (AUC) of 0.802 on the internal test, but 0.717 on data from the National Institutes of Health and 0.756 on data from Indiana. The explanation was in the data: pneumonia prevalence was 34.2% at Mount Sinai against 1.2% and 1.0% at the other two. A model trained only to guess which hospital a radiograph came from succeeded in 99.95% of cases. The network had learned to recognise the hospital, not the disease.
In digital pathology, Campanella et al. (Nature Medicine, 2019) trained a model on 44,732 slides from 15,187 patients, with an AUC of 0.991 for prostate cancer on the internal test. On more than 12,000 consultation slides from other institutions, the AUC fell by about 6 points; simply changing the scanner cost about 3 points.
The best-known commercial example is the sepsis prediction model built into a widely used hospital information system. Wong et al. (JAMA Internal Medicine, 2021) evaluated it on 27,697 patients and 38,455 hospitalizations at Michigan Medicine: AUC 0.63, against 0.76 to 0.83 in the developer's documentation. At the threshold in use, sensitivity was 33%, specificity 83% and positive predictive value 12%. The system generated alerts for 18% of hospitalizations and missed 67% of patients with sepsis.
The problem is structural. Wu et al. (Nature Medicine, 2021) analysed the 130 AI medical devices cleared by the FDA between 2015 and 2020: 126 had been evaluated retrospectively only, 93 published nothing on multi-site evaluation, and of the 41 that did, 4 had been evaluated at a single site and 8 at two.
The metrics that matter clinically
Vendors usually report "accuracy", a figure that depends on prevalence and says little about the clinical decision. The table below shows the metrics that should be requested.
| Metric | What it answers | Why it matters to the hospital |
|---|---|---|
| Sensitivity | Of the patients with the disease, how many are detected? | Missed cases (false negatives) |
| Specificity | Of the patients without the disease, how many are correctly ruled out? | False alarms and alert fatigue |
| AUC (area under the ROC curve) | How well the model separates sick from healthy, across all thresholds | Comparing models; says nothing about a specific threshold |
| Positive predictive value | When the model alerts, how often is it right? | Depends strongly on local prevalence |
| Calibration | When the model says "30% risk", does it happen in 30% of cases? | Decisions based on probability, not just on ranking |
The same model has several sensitivity-specificity pairs, depending on the chosen threshold. In Gulshan et al. (JAMA, 2016), the algorithm was reported at two operating points on the EyePACS-1 set: 90.3% sensitivity with 98.1% specificity, or 97.5% sensitivity with 93.4% specificity, for the same AUC of 0.991. The threshold is a clinical and organisational decision, not a technical one.
Calibration matters just as much. A model can have a good AUC and still systematically overestimate risk (unnecessary interventions) or underestimate it (missed cases). Wong et al. reported both poor discrimination and poor calibration for the sepsis model. Ask for the calibration curve, not only the AUC.
How much data and how much compute
There is no universal number, but published studies give the orders of magnitude. The models that made history in image diagnosis were trained on hundreds of thousands of labelled examples: 128,175 retinal photographs in Gulshan et al. (2016), 158,323 radiographs in Zech et al. (2018). In pathology, where one slide contains billions of pixels, Campanella et al. (2019) used 44,732 whole slides.
Two techniques reduce the need for labels. Transfer learning starts from a model trained on another task and fine-tunes it on medical data. Foundation models are pre-trained on large volumes of unlabelled data and then adapted with few labels. RETFound, published by Zhou et al. in Nature (2023), was pre-trained on 1.6 million retinal images without any label (904,170 fundus photographs and 736,442 OCT scans). Once adapted, it outperformed the comparison models for heart failure prediction with only 10% of the labels, and for diabetic retinopathy it reached comparable results with 45-50% of the data.
Compute cost spans very different orders of magnitude. According to Zhou et al., pre-training RETFound took about 14 days on 8 NVIDIA A100 graphics cards, while adapting it to a new task required about 70 minutes per 1,000 images on a single T4 card, a common piece of cloud equipment. At the other end, Epoch AI (Cottier et al., 2024) estimates the amortized hardware and energy cost of training GPT-4 at about 40 million USD and Gemini Ultra at about 30 million USD, with a growth of 2.4 times per year since 2016.
Data governance and the European framework
Training data comes from real patients, so its governance is a legal matter before it is a technical one. Three regulations matter.
GDPR. Regulation (EU) 2016/679 defines, in Article 4, the controller as the entity which "determines the purposes and means" of processing (point 7) and the processor as the entity which processes data "on behalf of the controller" (point 8). When a hospital makes data available for training, the hospital remains the controller and decides the purpose; the vendor acts, as a rule, as a processor under a written contract. Pseudonymisation (point 5) leaves the data attributable to a person "with the use of additional information", so it remains personal data. The recommended practice is de-identification as complete as possible before the data leaves the hospital, with the re-identification key kept exclusively at the hospital.
EHDS. Regulation (EU) 2025/327 on the European Health Data Space, adopted on 11 February 2025, regulates the secondary use of health data. Article 53(1)(e)(ii) lists explicitly, among the permitted purposes, "training, testing and evaluation of algorithms, including in medical devices, in vitro diagnostic medical devices, AI systems and digital health applications". Access goes through health data access bodies, and Article 54 provides that users process the data only on the basis of a data permit. Under Article 105, the regulation applies from 26 March 2027, and Chapter IV on secondary use from 26 March 2029.
AI Act. Regulation (EU) 2024/1689 entered into force on 1 August 2024 and applies in general from 2 August 2026. AI systems that are medical devices, or safety components of them, and undergo a conformity assessment by a notified body are high-risk (Article 6(1) and Annex I). For them, Article 10 requires training, validation and testing data sets that are relevant, sufficiently representative and, to the best extent possible, free of errors and complete, with examination of possible biases that may affect health and safety, and Article 72 requires a post-market monitoring system that "actively and systematically" collects and analyses performance data throughout the system's lifetime. The timetable has been amended: Regulation (EU) 2026/1744 (the Digital Omnibus on AI), adopted on 8 July 2026, sets the application of the obligations for Annex III systems from 2 December 2027 and for Annex I systems, which include medical devices, from 2 August 2028.
The monitoring required by Article 72 has a direct technical justification. A hospital's population, equipment and protocols change, and a model trained on last year's data can gradually lose performance, a phenomenon called drift.
What this means for a hospital in Romania or Moldova
For a hospital evaluating an AI project, whether buying it or developing it with a partner, the practical steps are these:
- Define the task and the threshold together with clinicians: which decision the model supports, which error you prefer (false alarms or missed cases) and who responds to the alert.
- Take inventory of your data before the quotation: how many cases you have, on which equipment they were produced, which labels already exist in the information system.
- Plan labelling as a clinical project: at least two independent graders and a mechanism for resolving disagreements, with time allocated.
- Request a local validation before the production contract: the model runs on a retrospective sample from your hospital, labelled by your clinicians.
- Sign the data processing agreement before a single record leaves the hospital: hospital as controller, vendor as processor, de-identified data, key kept at the hospital, purpose limited to training and validation.
- Set up monitoring from the start: who measures sensitivity, specificity and alert rate monthly, and at which threshold the system is paused for re-evaluation.
Questions to ask a vendor about training and validation:
- From how many institutions does the data come, and how many cases are in each set (training, validation, test)?
- Was the test set used once, at the end? Who had access to it?
- Who assigned the labels, how many graders per case, and how were disagreements resolved?
- Do you have the training and validation loss chart for the delivered model?
- At how many external hospitals was the model validated, and by how much did the AUC drop compared with the internal test?
- What are the sensitivity, specificity and positive predictive value at the proposed threshold, and what does the calibration curve look like?
- On which equipment and protocols was it trained? What happens if we change the scanner?
- How do you monitor drift after installation, and what documentation do you provide under Article 72 of the AI Act?
- Will our data be used to train models delivered to other clients? Under what conditions?
Consdinamic builds software and artificial intelligence to order, with its deepest specialisation in healthcare, and works with Aiforia (Finland) on de-identified data sets for digital pathology. Its own products run daily in a private medical network in Romania (Gral Medical, 29 locations), and local validation with a data agreement before any training is the practice applied in its projects.
Conclusion
Training a neural network is not a black box for hospital leadership: three data sets with separate roles, labels assigned by several clinicians, a training chart showing where the model stopped, an external validation with the performance drop stated honestly, metrics at a threshold chosen together with physicians, and monitoring after installation. The studies cited show that the performance drop between hospitals is the rule, not the exception, and GDPR, EHDS and the AI Act (with the 2 August 2028 deadline for medical devices) turn these good practices into obligations. A hospital that asks the questions in this guide before signing protects its patients, its budget and its clinicians' time.
- Zech et al., Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs, PLOS Medicine, 2018 — 158,323 radiographs from 3 institutions; AUC 0.802 internal, 0.717 and 0.756 external; prevalence 34.2% vs 1.2% and 1.0%; hospital identified in 99.95% of images
- Wong et al., External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients, JAMA Internal Medicine, 2021 — 27,697 patients, 38,455 hospitalizations; AUC 0.63 vs 0.76-0.83 in the developer's documentation; sensitivity 33%, specificity 83%, PPV 12%; alerts on 18% of hospitalizations
- Wu et al., How medical AI devices are evaluated, Nature Medicine, 2021 — Stanford HAI policy brief, 2022 — 130 FDA devices 2015-2020; 126 evaluated retrospectively only; 93 with no information on multi-site evaluation; 4 evaluated at one site; 59 without sample size; median 300
- Elmore et al., Diagnostic Concordance Among Pathologists Interpreting Breast Biopsy Specimens, JAMA, 2015 — 115 pathologists, 6,900 interpretations; 75.3% concordance with the consensus reference; 96% invasive carcinoma, 84% DCIS, 48% atypia, 87% benign
- Gulshan et al., Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy, JAMA, 2016 — 128,175 images labelled by 54 ophthalmologists, 3-7 grades per image; AUC 0.991; two operating points: 90.3% / 98.1% and 97.5% / 93.4%
- Campanella et al., Clinical-grade computational pathology using weakly supervised deep learning on whole slide images, Nature Medicine, 2019 — 44,732 whole-slide images from 15,187 patients; AUC 0.991 prostate; drop of about 6 AUC points on external slides and 3 points on another scanner
- Zhou et al., A foundation model for generalizable disease detection from retinal images (RETFound), Nature, 2023 — 1.6 million unlabelled images; 8 A100 GPUs for about 14 days; fine-tuning about 70 min per 1,000 images on one T4 GPU; comparable results with 10-50% of labels
- Cottier et al., The rising costs of training frontier AI models, Epoch AI / arXiv, 2024 — amortized hardware and energy cost: GPT-4 about 40 million USD, Gemini Ultra about 30 million USD; growth of 2.4x per year since 2016
- Regulation (EU) 2024/1689 on artificial intelligence (AI Act), EUR-Lex — Article 10 (training, validation, testing data sets), Article 72 (post-market monitoring), Article 113 (in force 1 August 2024, general application from 2 August 2026)
- Regulation (EU) 2026/1744 (Digital Omnibus on AI), EUR-Lex — adopted 8 July 2026; obligations for Annex III high-risk systems from 2 December 2027 and for Annex I systems (medical devices) from 2 August 2028
- Regulation (EU) 2025/327 on the European Health Data Space (EHDS), EUR-Lex — adopted 11 February 2025; Article 53(1)(e)(ii): training, testing and evaluation of algorithms; Article 54: only under a data permit; Chapter IV from 26 March 2029
- Regulation (EU) 2016/679 (GDPR), EUR-Lex — Article 4: definitions of controller (point 7), processor (point 8), pseudonymisation (point 5) and data concerning health (point 15)
Tell us what you need. We come back with a prototype, not with slides.