Abstract
This topic investigates the quality and correctness of nine automatically assigned labels for merge-conflict resolutions. The labels were reconstructed by replaying historical Git merges and mapping the tentative conflict representation to the final developer-committed file. Such mappings are reproducible and scalable, but they can be affected by formatting changes, reordered code, structural edits, ambiguous correspondence between conflict chunks, and compound resolutions. The project will develop methods to detect potentially incorrect, ambiguous, or unstable labels and will validate a sample of these cases through systematic inspection. The work combines supervised, unsupervised, and weakly supervised methods for label-quality assessment.
Supervision
- Reza, Darooei
Motivation
The nine labels are central to all subsequent classification experiments. If some labels are incorrect or ambiguous, the model may learn artifacts of the labeling procedure rather than actual developer resolution behavior. This problem is particularly important because the final Git repository records the resolved file but does not explicitly preserve the exact correspondence between every original conflict chunk and its final resolution.
The current mapping procedure matches context sections and conflict alternatives to the final file. It can assign labels such as selecting one side, combining alternatives, deleting content, applying another compound pattern, or performing a non-canonical edit. However, the supplied study explicitly acknowledges that deterministic text matching may misclassify cases involving formatting changes, reordered code, or semantically equivalent edits. It also notes that compound labels may omit the order in which alternatives were combined and that structurally difficult cases are excluded conservatively.
Goal
The goal is to assess the reliability of the nine-label dataset and develop a method for identifying suspicious labels. The student should:
Document the precise definition and extraction procedure for all nine labels.
Quantify label distributions by project, programming language, file category, time period, and merge.
Identify possible sources of label noise
Requirements
Good Python programming skills.
Basic knowledge of supervised and unsupervised machine learning.
Familiarity with classification metrics, confusion matrices, model confidence, and cross-validation.
Interest in data quality, empirical software engineering, and reproducible research.
Basic knowledge of Git
Pointers
P. Elias, H. D. S. Campos, E. Ogasawara, and L. G. P. Murta, “Towards Accurate Recommendations of Merge Conflicts Resolution Strategies,” Information and Software Technology, vol. 164, Art. no. 107332, 2023, doi: 10.1016/j.infsof.2023.107332.
A. Boll, Y. van Dok, M. Ohrndorf, A. Schultheiß, and T. Kehrer, “Towards Semi-Automated Merge Conflict Resolution: Is It Easier Than We Expected?” in Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering, 2024, pp. 282–292, doi: 10.1145/3661167.3661197.