Deliverables
D2.1 Manuscript on MDR, IVDR and AI Act
Artificial intelligence is reshaping healthcare by introducing adaptive, data-driven technologies into medical devices, diagnostics, and clinical workflows. Because AI systems rely on continuous learning and the large-scale processing of sensitive health data, their integration into medical technologies has prompted the EU to build a layered regulatory framework involving the Artificial Intelligence Act (AI Act), the General Data Protection Regulation (GDPR), and the Medical Device and In Vitro Diagnostic Regulations (MDR/IVDR). This Deliverable, produced within WP2 of the REALM project, offers an inte-grated analysis of how these regimes jointly apply to AI in the medical-devices sector.
The study has three main objectives. First, it maps the obligations imposed on key actors—manufac-turers, providers, deployers, operators, controllers, and processors—highlighting where terminology differs but responsibilities converge across the AI Act, MDR/IVDR, and GDPR. Second, it examines how horizontal (AI Act, GDPR) and sector specific (MDR/IVDR) regulations interact in practice, identifying areas of complementarity as well as risks of fragmentation. Guidance from bodies such as the MDCG and the forthcoming AIB is considered essential for ensuring coherent interpretation. Third, the anal-ysis assesses the practical impact of this regulatory environment on innovation, market access, and healthcare delivery, with particular attention to the obligations of healthcare institutions deploying AI systems.
Methodologically, the study draws on binding EU legislation and relevant soft-law instruments, apply-ing a comparative and teleological interpretation to clarify overlaps and divergences. Because all four regulations have extraterritorial reach, their obligations extend to non-EU actors whose AI systems affect EU patients. A functional comparative approach further shows how legal requirements translate into conformity assessment, quality-management processes, and post-market surveillance duties.
D4.1 Federated Cloud-Based Data Repository
Deliverable D4.1 documents the development and deployment of the federated, cloud-based data repository infrastructure designed under Task 4.1. REALM data repository is a cornerstone of the REALM platform, enabling secure, privacy-preserving, and interoperable access to anonymized, multimodal real-world data (RWD) and synthetic data to support the evaluation of AI-based medical device software (MDSW). The architecture follows a federated approach that accommodates direct access to data. This access mode allows national nodes to host and manage data locally while enabling AI model evaluation within their secure environments.
The repository is powered by a robust infrastructure stack integrating Docker, Kubernetes, Kafka, PostgreSQL, and OHDSI tooling. These technologies collectively support data ingestion, transformation (extract, transform, load (ETL) operations), cohort generation, and AI evaluation through the RIANA dashboard, and semantic alignment using the OMOP Common Data Model. Extensions were developed to support genomics, radiology, and oncology use cases, and synthetic data generators were implemented for scenarios where real-world data is limited or incomplete. Deployment requirements and configurations are documented to facilitate the setup of the national REALM data nodes.
A central component of this infrastructure is the standardized data catalogue (T4.3) based on the OMOP Common Data Model, extended to support genomics (G-CDM), radiology (R-CDM), and oncology domains. A comprehensive ETL pipeline and mapping methodology, combining OHDSI tools and custom scaffolding, was developed to enable the transformation of heterogeneous data into OMOP Common Data Model compliant datasets, with real-world and synthetic datasets used to validate the pipelines.
The work is further supported by extensive engagement with data providers, resulting in a cooperation model that emphasizes shared value over direct compensation. This includes technical support, benchmarking insights, and opportunities for joint dissemination.
Looking forward, this deliverable opens several promising directions for future exploration and uptake. These include operationalizing the indirect access model in alignment with national and EU-level regulatory frameworks, such as those defined under the European Health Data Space (EHDS). REALM’s national nodes could potentially serve as complementary infrastructures under the governance of Health Data Access Bodies (HDABs). Further opportunities lie in expanding the network of participating data providers, improving data and model interoperability with EHDS standards, and strengthening automated tooling for data mapping and quality assurance.
D4.2 Data Management Plan V1.0
This data management plan (DMP) outlines the strategies, processes, and resources employed to ensure effective management of data throughout the lifecycle of REALM project. The key objectives of the DMP are to promote data integrity, facilitate data sharing and reuse, and ensure compliance with relevant regulations in Europe and funding agency requirements from Horizon Europe. To provide a comprehensive data management overview, this DMP took the inputs from five demonstrators who will provide real-world data and use cases in the project and considered the feedback from members working in WP1 (Michel Dumontier), WP2 (Birgit Wouters), WP3 (Ilias Siniosoglou), and WP5 (Bart Elen).
The goal of REALM project is to create an inclusive platform for evaluation and certification of software in healthcare where both the developers and the regulatory bodies have access to a standardized set of technology stack and data. The research activities include building REALM architecture, establishing real-world and synthetic data repository, integrating REALM tools and testing and monitoring REALM tools. In REALM, data plays a crucial role in achieving the project objectives, robust data management practices have been implemented to maximize the value and impact of the generated research outputs.
This DMP describes the types and formats of data, the size of the data, origin and provenance of the data that will be re-used and generated in the project. Then, the DMP elaborates how the data and other research outcomes in REALM project will be made Findable, Accessible, Interoperable, and Reusable. We classify data into 1) real-world data (that are individually identifiable) from five demonstrators in the project, 2) anonymize data (that are not individually identifiable) which are requested from biobanks and data repositories (such as UK Biobank, FINNGEN), and 3) synthetic data (that are not individually identifiable) which are generated based on the real-world data. Based on the different nature of the data, the project proposes different strategies to make these data accessible and reusable for other parties under different level of restrictions.
Data storage and security are given utmost importance in the project. The data is stored in private clouds or trusted data repositories which provides secure and reliable storage solutions. Robust data security measures, including encryption, access controls, and regular backups, will be implemented to protect the integrity and confidentiality of the data. Long-term preservation of the data is crucial for future reference and potential reuse. The project uses persistent identifiers for all datasets and research outcomes and archives and stores data using the trusted services such as university or national research infrastructure to ensure the longevity and accessibility of the data beyond the project duration. The DMP adheres to relevant compliance and ethical considerations. This involves EU GDPR, and national data protection regulations in state members, informed consent procedures, or anonymization/de-identification methods to protect the privacy and confidentiality of individuals or entities involved in the research.
In conclusion, this first version of DMP for REALM project provides a comprehensive overview of managing data in the project. The strategies and practices outlined in the DMP ensure data integrity, facilitate data sharing and reuse, and adhere to relevant compliance and ethical considerations. The achievements attained through the implementation of the DMP highlight the project's commitment to robust data management practices and its contribution to advancing scientific knowledge.
After this first version of DMP (V1.0), we will keep updating the DMP every six months throughout the project lifecycle. The new data, methods, software, and other research outputs that are created from all the work packages in REALM must be added in the updated version of DMP (V2.0). Any major changes, deletion, or extension from the first version of DMP will be highlighted in the updated DMP (V2.0).
D4.3 Data Management Plan V2.0
This updated Data Management Plan (DMP V2.0) presents the updated status and evolution of data management practices in the REALM project, reflecting developments and lessons learned since the initial DMP (D4.2). The core objectives remain the same: to ensure data integrity, enable secure and ethical data sharing, and promote compliance with European regulations and Horizon Europe requirements. This version integrates updated inputs from all demonstrators, along with progress and refinements from WPs 2, 3, 5 and 6.
REALM aims to build a trustworthy, inclusive platform for evaluation and certification of AI-enabled healthcare software by combining real-world, synthetic, and anonymized data with a standardized technology stack. Over the past 30 months, significant progress has been made in establishing the REALM architecture, integrating tools, and building the data repositories that support testing and validation workflows. Data management has played a central role in ensuring that these activities remain aligned with FAIR principles and regulatory expectations.
This DMP V2.0 documents the types and volumes of data collected or generated, including updates to data provenance, metadata standards, and data quality practices. We have classified and managed three categories of data:
- personal real-world data from demonstrators under strict access controls.
- anonymized datasets from biobanks and trusted repositories; and
- synthetic data generated to support algorithm development and benchmarking.
This DMP V2.0 captures the REALM project's commitment to responsible and transparent data management. It ensures that data assets remain FAIR, secure, and privacy-preserving, supporting the immediate project goals and broader scientific and regulatory communities in healthcare AI.
D4.4 Standard Data Quality Framework (SDQF)
This deliverable demonstrates the work carried out for the Standard Data Quality Framework (SDQF) regarding the REALM platform. The first step was to collect and find out what technical teams and users need. Surveys, interviews, and expert feedback were used for this purpose. Requirements collected include functional requirements (e.g., data and model uploads), security requirements (e.g., user privacy and encryption), and ethical requirements (e.g., fairness and transparency). These will allow the platform to operate effectively within real healthcare environment and be compliant with EU legislation.
Moreover, key stakeholders and players, including hospitals, research institutions, health industries and developers, regulators (including Health Technology Assessment (HTA) agencies) and patient groups, were identified. Understanding their roles and needs ensured that the platform will be effective and ethical with real end-users.
Additionally, a Standard Data Quality Framework (SDQF) was developed by adapting quantitative and qualitative dimensions developed for assessing the quality of medical data from the QUANTUM consortium3, an EU-funded project on defining health data quality labels in the context of European Health Data Space (EHDS). This framework enables the assessment of how broad, accurate and useful medical datasets used for evaluation of AI-based Medical Device Software (MDSW) are, which is critical in validating AI models for healthcare applications. Overall, this deliverable outlines the foundation for designing a secure, fair and effective AI test platform by combining strong legal, technical and ethical requirements with definite user and system requirements.
D4.5 Authorized access agreements with European Medical Data Pools and Digital Twins
Europe’s diverse medical data sources offer significant potential for research, innovation, and decision-making, yet accessing them raises ethical, legal and privacy concerns that need to be establish through robust data access agreements. This deliverable builds upon the effort made on the Data Management Plan (DMP), outlines the data providers, initial strategies, processes, and resources employed to identify relevant data pools and implement access agreements with European medical data pools and digital twin community.
This first version of the data access agreement template provides a comprehensive overview of the workflow for getting signed access agreements with data providers.
D4.6 Multimodal Synthetic Data Generator
Deliverable D4.6 “Multimodal Synthetic Data Generator” reports on the outcomes of Task 4.4, carried out between months 5 and 35 of the REALM project. The task was designed to achieve the main objectives: to generate synthetic multimodal medical data and complete patient profiles, to improve accessibility of clinical data with privacy guaranteed.
The work presented in this deliverable has demonstrated that these objectives have been fulfilled. Synthetic data has been successfully generated across all modalities required for the REALM use cases, including time-series, CT, chest X-ray, and electronic health records. Privacy-preserving mechanisms such as differential privacy have been integrated into the synthetic data generation framework, ensuring that the data can be shared and reused without compromising confidentiality.
A distinctive achievement of Task 4.4 is the development of a comprehensive evaluation framework that assesses the quality of synthetic data across multiple dimensions: fidelity, utility, diversity, fairness, and privacy.
A central purpose of the synthetic data framework in Task 4.4 is to enable more accurate subgroup evaluation. By conditionally generating samples for under-represented demographic and clinical strata, the framework alleviates the small-sample problem that typically undermines fairness assessment. The resulting evaluation suites yield narrower confidence intervals, greater statistical power, and more stable estimates of performance across intersections (e.g., age × sex × ethnicity) while preserving privacy through differential privacy and strict non-memorization safeguards. In practice, synthetic data supplements hold-out real data, providing calibrated, modality-specific test sets that allow equitable and evidence-based validation of models across time-series, CT, chest X-ray, and EHR modalities Beyond methodological advances, Task 4.4 has delivered tangible outputs in the form of open-source code, evaluation benchmarks, and synthetic datasets integrated into the REALM data repository in full compliance with FAIR principles. These outputs provide a lasting contribution to the consortium and to the wider research community.
The results of Task 4.4 are closely linked to other REALM activities, feeding into the statistical and deep learning approaches of Task 5.1, supporting the data fusion and integration work of Task 5.2, underpinning the RIANA dashboard in Task 5.3, and operating within the federated data repository established in Task 4.3. Integration with the REALM architecture and the AI orchestrator (Tasks 3.1 and 3.2) ensures that synthetic data generation is embedded within the overall system design.
The next steps will focus on the complete integration of synthetic data generation into the REALM production environment. This will occur within Work Package 6, where all REALM components will be brought together into the REALM evaluation platform. The platform will provide the means to conduct rigorous, privacy-preserving evaluations of the five REALM use cases, ensuring that the innovations developed in Task 4.4 contribute directly to clinical translation and regulatory impact.
In conclusion, this deliverable demonstrates that multimodal synthetic data generation is feasible, reliable, and impactful. By combining advanced generative methods, formal privacy guarantees, and subgroup evaluation for fairness, Task 4.4 has fulfilled its objectives and positioned synthetic data as a cornerstone of the REALM framework for trustworthy healthcare AI.