BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055402Z
LOCATION:Online
DTSTART;TZID=America/Chicago:20211115T103000
DTEND;TZID=America/Chicago:20211115T105500
UID:submissions.supercomputing.org_SC21_sess342_ws_worksp101@linklings.com
SUMMARY:A Recommender System for Scientific Datasets and Analysis Pipeline
 s
DESCRIPTION:Workshop\n\nA Recommender System for Scientific Datasets and A
 nalysis Pipelines\n\nMazaheri, Kiar, Glatard\n\nScientific datasets and an
 alysis pipelines are increasingly being shared publicly in the interest of
  open science. However, mechanisms are lacking to reliably identify which 
 pipelines and datasets can appropriately be used together. Given the incre
 asing number of high-quality public datasets and pipelines, this lack of c
 lear compatibility threatens the findability and reusability of these reso
 urces. We investigate the feasibility of a collaborative filtering system 
 to recommend pipelines and datasets based on provenance records from previ
 ous executions. We evaluate our system using datasets and pipelines extrac
 ted from the Canadian Open Neuroscience Platform, a national initiative fo
 r open neuroscience. The recommendations provided by our system (AUC=0.83)
  are significantly better than chance and outperform recommendations made 
 by\ndomain experts using their previous knowledge as well as pipeline and 
 dataset descriptions (AUC=0.63). In particular, domain experts often negle
 ct low-level technical aspects of a pipeline-dataset interaction, such as 
 the level of pre-processing, which are captured by a provenance-based syst
 em. We conclude that provenance-based pipeline and dataset recommenders ar
 e feasible and beneficial to the sharing and usage of open-science resourc
 es. Future work will focus on the collection of more comprehensive provena
 nce traces, and on deploying the system in production.\n\nTag: Online Only
 , Cloud and Distributed Computing, Scientific Computing, Workflows\n\nRegi
 stration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
