In this register-based study, using data from Närhälsan's electronic health record system in the Västra Götaland Region (VGR) and linkage with data from Statistics Sweden (SCB), the research questions will be addressed through the development and validation of AI-based models. At a later stage of the process, the ability of the AI models to predict foot ulcers will be compared with that of statistical models.
From Asynja Whisp, Närhälsan's electronic health record system in VGR, data will be retrieved for all adult patients (18 years or older) with diagnoses (according to ICD-10) who either have a diabetes diagnosis (E10-E14) or have been prescribed any diabetes medication after the age of 18, covering the period from 2014 to 30 June 2025.
Based on, among other variables, diagnostic codes (ICD-10), procedure codes (KVÅ), visit types, visit frequency, ECG parameters, and free-text/clinical notes, predictors will be identified, such as neuropathy, impaired circulation, previous ulcers, antibiotic treatment, foot deformities, and skin status. The data will be validated and, if necessary, supplemented with additional parameters.
Methods to address the research questions
Machine learning-based models will be trained to predict the risk of developing foot ulcers. Cross-validation will be used to identify optimal hyperparameters for each model. In the first phase, the models' ability to discriminate between patients with diabetic foot ulcers and patients without foot ulcers will be evaluated. In the second phase, the models' ability to prospectively predict ulcer development will be assessed. Redundant variables will be excluded, and the models will be retrained in an iterative process to increase robustness.
The models will be combined with conformal prediction to integrate uncertainty estimation into the predictions and to identify patients for whom the model is unsuitable for prediction. Finally, the most predictive variables will be identified using Shapley values (SHAP).
Statistical models
Using electronic health record data from the Asynja Whisp care information system in VGR primary care, together with SCB data and scientific and empirical evidence, variables and categories that constitute potential risk factors for foot ulcers will be identified. A case-control design will be applied, in which the control group consists of people with diabetes who have not developed foot ulcers, compared with patients who have developed foot ulcers.
In the development of statistical prediction models, the workflow involves analysing populations, i.e. all patients with diabetes without foot ulcers compared with all patients with diabetes who have foot ulcers. This allows investigation of potential associations between the occurrence of foot ulcers in patients with diabetes and other factors. In collaboration with the medical profession, causal relationships underlying the occurrence of foot ulcers will be identified. A model will be developed that describes chains of causation leading to the occurrence of foot ulcers in patients with diabetes, and information will be provided on the degree of certainty of the model.
Based on the results of the models (both AI-generated and statistical), strengths and weaknesses of each approach will be compared. Validation of the developed models will be performed on an independent dataset to ensure that the results are generalisable and robust over time.
The validation strategy ensures that the model performs well on new patients and not only on the dataset from which it was developed. Outcome measures for validation include sensitivity (how well the model identifies those who truly have a high risk of foot ulcers), specificity (how well the model avoids false alarms), and positive predictive value (PPV). Furthermore, the model will be interpreted to ensure transparency and clinical interpretability. Development, testing, and validation will be conducted in collaboration with patient representatives, clinicians, and researchers.

