mbilal.works
Available · September 2026
Talk
← Back to Portfolio
Practitioner Engineering Publication · arXiv cs.SEOPEN ACCESS
Self-Initiated Study · May 2026

Given, When, Then, Again: Mining Subscenario Refactoring Candidates in Behaviour-Driven Test Suites with ML Classifiers and LLM-Judge Baselines

Co-authored arXiv paper mining BDD subscenario refactoring candidates across 5.3M slices using ML classifiers benchmarked against LLM judges. XGBoost classifier (F₁ = 0.891) beat a tuned rule baseline and two open-weight LLMs (GPT and Ling) at p < 10⁻⁴.

01. Context & Motivation

Research Context & Problem Framing

No prior work automated which recurring step subsequences in BDD test suites are worth extracting as refactoring candidates, or which of the three standard patterns (Background, reusable-scenario, shared-step) applies.

02. Methodology & System Design

Methodology & Implementation

Mined every contiguous L-step window (L ∈ [2,18]) across a 339-repo corpus, keyed by paraphrase-robust cluster IDs from the foundational BDD paper. Applied SBERT/UMAP/HDBSCAN for paraphrase clustering. Trained XGBoost extraction-worthy classifier under 5-fold CV; compared against rule baseline and two open-weight LLM judges on a human-labelled 200-slice pool.

03. Empirical Findings & Evidence

Findings & Key Results

XGBoost F₁ = 0.891 (95% CI [0.852, 0.927]) beat LLM judges at F₁ ≤ 0.728 (McNemar p < 10⁻⁴). 75% / 59.5% / 11.7% of scenarios carry within-file / within-repo / cross-org refactoring candidates respectively.

04. Topics & Methodology Stack
PythonXGBoostSBERTUMAPHDBSCANStatistical Testing
mbilal.works · personal hub at bilalm.me · Impressum · PrivacyBuilt to a DESIGN.md spec · Bricolage Grotesque + JetBrains Mono · MMXXVI