01. Context & MotivationResearch Context & Problem Framing
LLMs are increasingly used to produce test oracles, but no clear account existed of where these oracles draw their authority. Prior secondary studies organised by oracle form or LLM technique, not by the source of the verdict's trust basis.
02. Methodology & System DesignMethodology & Implementation
Conducted a systematic literature review under PRISMA 2020 guidelines. From 2,436 records, LLM pre-filtering followed by independent dual human screening (Cohen's κ = 0.79) and full-text assessment yielded 54 included studies. Analysed along three axes: source of authority, oracle form, and adjudication mechanism.
03. Empirical Findings & EvidenceFindings & Key Results
Specification-derived authority covers ~52% of studies (28/54); remaining 26 reach verdicts without any specification. 'LLM-as-a-judge' names a mechanism, not a basis for trust. Taxonomy's sparse and empty regions constitute an explicit research agenda.
04. Topics & Methodology StackSystematic ReviewPRISMA 2020LLM EvaluationTest Oracle AnalysisCohen's κ