Repository logo
Log In(current)
  1. Home
  2. Colleges & Schools
  3. Tickle College of Engineering
  4. Engineering Publications and Other Works
  5. Min H. Kao Department of Electrical Engineering and Computer Science
  6. Electrical Engineering and Computer Science Publications and Other Works
  7. The Role of Data Filtering in Open Source Software Ranking and Selection
Details

The Role of Data Filtering in Open Source Software Ranking and Selection

Date Issued
April 1, 2024
Author(s)
Malviya-Thakur, Addi  
Mockus, Audris  
DOI
https://doi.org/10.1145/3643664.3648210
Permanent URI
https://trace.tennessee.edu/handle/20.500.14382/13215
Abstract

Faced with more than 100M open source projects, a more manageable small subset is needed for most empirical investigations. More than half of the research papers in leading venues investigated filtering projects by some measure of popularity with explicit or implicit arguments that unpopular projects are not of interest, may not even represent “real” software projects, or that less popular projects are not worthy of study. However, such filtering may have enormous effects on the results of the studies if and precisely because the sought-out response or prediction is in any way related to the filtering criteria.


This paper exemplifies the impact of this common practice on research outcomes, specifically how filtering of software projects on GitHub based on inherent characteristics affects the assessment of their popularity. Using a dataset of over 100,000 repositories, we used multiple regression to model the number of stars –a commonly used proxy for popularity– based on factors such as the number of commits, the duration of the project, the number of authors and the number of core developers. Our control model included the entire dataset, while a second filtered model considered only projects with ten or more authors. The results indicated that while certain characteristics of the repository consistently predict popularity, the filtering process significantly alters the relationships between these characteristics and the response. We found that the number of commits exhibited a positive correlation with popularity in the control sample but showed a negative correlation in the filtered sample. These findings highlight the potential biases introduced by data filtering and emphasize the need for careful sample selection in empirical research of mining software repositories. We recommend that empirical work should either analyze complete datasets such as World of Code, or employ stratified random sampling from a complete dataset to ensure that filtering is not biasing the results.

Subjects

Empirical software en...

Missing data problem

Filtering

Sampling

Mining software repos...

Disciplines
Computer Sciences
Databases and Information Systems
Data Science
Software Engineering
Recommended Citation
Addi Malviya-Thakur and Audris Mockus. 2024. The Role of Data Filtering in Open Source Software Ranking and Selection. In International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE ’24 ), April 16, 2024, Lisbon, Portugal. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3643664.3648210
Embargo Date
May 5, 2025
File(s)
Thumbnail Image
Name

3643664.3648210.pdf

Size

518.97 KB

Format

Adobe PDF

Checksum (MD5)

eff373d2cdc0b280d7960687271dc149


University Libraries

1015 Volunteer Boulevard
Knoxville, TN 37996
865-974-4351

Map & Directions
Donate to the Libraries
  • About
  • John C. Hodges Society
  • Speaking Volumes magazine
  • Outreach
  • Directory
  • Employment
  • Policies
  • Library Intranet
University of Tennessee power T logo

The University of Tennessee, Knoxville
Knoxville, Tennessee 37996
865-974-1000

Events
A-Z
Apply
Privacy
Map
Directory
Give to UT
Accessibility

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science