Responsible Data Science Lab at Purdue
We study problems at the intersection of data management and machine learning to build trustworthy and responsible decision-making systems. Our aim is to develop systems that enable explainability, fairness, and accountability of data-driven decision-making systems. We are particularly interested in:
- Explaining and debugging fairness violations in machine learning models and data science pipelines:
- How can we determine sources of unexpected errors and bias in machine learning model outcomes?
- How can we decompose unexpected or discriminatory behavior of data science pipelines in terms of the different pipeline stages?
- Can we effectively generate post hoc explanations for the outcomes of machine learning models?
- Data integration and data quality:
- How can we leverage expert feedback to improve data cleaning techniques for machine learning?
- Can we use the final outcomes in data science pipelines to inform intermediate pipeline choices?
- How can we intertwine pipeline stages with downstream analytics to improve upon the end goals?
We are always looking for motivated Ph.D. students to collaborate with. If you are interested in data management and/or responsible data analytics, feel free to contact us with your CV/resume and a couple of sentences describing your research interests, and consider applying to Purdue SACC!
Sponsors We are thankful for the generous funding award and gift from our sponsors: NSF, Google, and CASMI.
news
| Sep 4, 2026 | The Responsible Data Science Lab had a strong presence at VLDB this week. Jahid presented his paper on PipeLens in the main conference while Ambarish presented his paper on Splice in the QDB workshop. |
|---|---|
| Sep 4, 2026 | Spring and summer graduations! Ike (co-chaired with B. C. Min) defended his Ph.D. thesis this summer. Harshita and Ananya defended their M.S. theses in spring and summer, respectively. Congratulations, Ike, Harshita, and Ananya! |
| Jul 8, 2026 | Our paper on Scalable Goal-Oriented Source Selection got accepted to the 15th International Workshop on Quality in Databases at the 52nd International Conference on Very Large Data Bases (VLDB). Congratulations, Ambarish! |
| Jul 3, 2026 | Our paper on Identifying Interventions for Resolving Malfunctioning Data Science Pipelines got accepted to the 52nd International Conference on Very Large Data Bases (VLDB). Congratulations, Jahid and Tejendra! |
| May 1, 2026 | Ph.D. student Jahid Hasan successfully passed his preliminary examination. Congratulations, Jahid! |
selected publications
- Explanations for Machine Learning Pipelines under Data DriftIn Workshop on Human-In-the-Loop Data Analytics (HILDA) at the ACM International Conference on Management of Data (SIGMOD), 2025