BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20211207T055411Z
LOCATION:223
DTSTART;TZID=America/Chicago:20211114T141500
DTEND;TZID=America/Chicago:20211114T144500
UID:submissions.supercomputing.org_SC21_sess434_ws_cafcw125@linklings.com
SUMMARY:Probing Decision Boundaries in Cancer Data Using Noise Injection a
 nd Counterfactual Analysis
DESCRIPTION:Workshop\n\nProbing Decision Boundaries in Cancer Data Using N
 oise Injection and Counterfactual Analysis\n\nJain, Shah, Mohd-Yusof, Wozn
 iak, Brettin...\n\nAdvanced analyses and computations based on gene expres
 sions are prone to errors as they depend on experimental design, chemical 
 operations/measurements, and data analysis. The assembly and aggregation o
 f such data for creating deep neural network models may further influence 
 the accuracy of these analyses. For example, the CANDLE [1] NT3[2] Benchma
 rk uses a table of laboratory-obtained data mapping RNA expression data to
  a normal or tumor designation and is used to make predictions about given
  expression samples. In this work, we use the NT3 Benchmark to study the e
 ffects of injecting bad data at different rates to study the impacts on th
 e resulting predictions. Our data manipulations include flipping classific
 ation labels (label noise) and introducing noise in gene expressions (feat
 ure noise). We present results for the performance of both the base NT3 Be
 nchmark and NT3 with the addition of the abstention class in the presence 
 of various types of injected noise. For higher noise levels, the ability o
 f the base network to correctly predict the normal/tumor classification (a
 s measured by the validation accuracy) degrades significantly. Use of the 
 abstaining classifier allows the model to learn when the labels have becom
 e unreliable and abstain from providing a prediction in that case, while r
 etaining accuracy. \n\nCounterfactual examples are an example-based inter
 pretability technique used by the explainable AI community.The technique a
 ims to mirror human counterfactual reasoning by finding a minimal subset o
 f changes to an input example so that a machine learning model classifies 
 the input into a different class. We demonstrate the use of counterfactual
  examples to identify the normal directions to the decision boundary "from
  normal to tumor" and perform further analysis to identify specific overex
 pressed genes, or “perturbation vectors” (the difference between the gener
 ated example and the original input). Perturbation vectors were separated 
 by class and clustered into groups. From the clustered perturbation vector
 s, identify those features which are important for classification. Noise w
 as injected only on the genes corresponding to the counterfactual perturba
 tion vector while keeping the label the same. We found that for a trained 
 NT3 model without abstention, this does in fact lead to steeper degradatio
 n of accuracy compared to with incremental noise injection on a randomly c
 hosen set of indices.\n\nThe top gene symbol is PLOD2, which is considere
 d to be the highway of cancer cell-migration as per a 2017 article[3]. Oth
 er genes identified in the counterfactual analysis include LRTM1, RGS5, TP
 53I13, MAN1B1, TRRAP and TP53I13 which have all been found overexpressed a
 nd linked to studies of urothelial, lung, renal, bladder, ovarian and bone
  cancer respectively. We believe that other genes found here might serve a
 s a good starting point and even lead to new discoveries in the area cance
 r research.\n\nThe major contributions of this work include 1) a study mo
 del performance on incremental noise injection to input data and, 2) use o
 f abstention classifiers to combat noisy data in the NT3 dataset, 3) a tec
 hnique to highlight the decision boundary of the NT3 model and identify ke
 y genes for cancer research with counterfactual analysis.\n\nTag: Applicat
 ions, Computational Science, Education and Training and Outreach, HPC Comm
 unity Collaboration, HPC Training and Education, Machine Learning and Arti
 ficial Intelligence, Performance, Workforce\n\nRegistration Category: Work
 shop Reg Pass
END:VEVENT
END:VCALENDAR
