FAIR software pipelines for FAIR data management

Coursework

This course presents the journey of research data from raw instrument output to reusable and FAIR-compliant datasets grounded in the data lifecycle. It explores the transition from rigid and legacy pipelines to modular and scalable metadata-driven workflows built for modern automation and deployment. The discussion is shaped by two practical workflow examples: a SEM image-classification pipeline that progresses from raw microscope output to a live classification service, and an Oxford Nanopore sequencing pipeline integrated with the ORFEO HPC infrastructure. Together, these cases trace the full arc from data mess to reproducible and shareable datasets. The course is intended for anyone working with experimental data workflows and provide a hands-on understanding of FAIR-by-design principles and modern data flow.

Key Topics

  • The data lifecycle, from generation to interpretation
  • FAIR principles and the FAIRification workflow
  • FAIR-by-design concept
  • Why data pipelines break: inconsistent metadata, format chaos, and scientist resistance to curation
  • Traditional pipelines vs. modern data flows: rigidity and technical debt vs. modularity and metadata-driven automation
  • Case studies: SEM image annotation, dataset publication, and the TriDAS classification service. along with nanopore basecalling, the live-basecalling bottleneck, and HPC infrastructure design
  • Building a modular pipeline on HPC: orchestration tools (Jenkins, Nextflow/nf-core) and design principles for robustness and usability

Learning Outcomes

  • By the end of the course, participants will be able to:
  • Explain the data lifecycle and how FAIR principles apply at each stage
  • Apply the FAIR-by-design approach
  • Recognize the causes of data mess and propose practical fixes
  • Compare traditional and modern pipeline architectures and their trade-offs
  • Walk through a real pipeline end-to-end, from raw data to published dataset
  • Reason about the infrastructure (storage, network, compute) needed for high-throughput, near-real-time processing
  • Sketch a modular HPC-native pipeline design with clear data handoffs and error handling
  • Evaluate orchestration tools like Nextflow/nf-core or Jenkins for a given research setting 

Lecturer(s)

placeholder person image
Lecturer
Area Science Park
Ahmed Khalil