mbilal.works
Available · September 2026
Talk
← Back to Portfolio
Research Data Infrastructure · TRR 266LIVE SYSTEM
TUM × LMU × Bocconi × IESE · 2024

STOXX Europe 600 Climate Disclosure Pipeline

Data ingestion infrastructure built for a TRR 266 working paper, extracting machine-readable corporate annual report sections for NLP analysis.

01. Problem & Context

The Operating Challenge

Academic researchers studying climate disclosure compliance needed financial statements and audit reports from six years of STOXX Europe 600 annual reports (2018–2023) converted into structured NLP datasets. Manual coding across thousands of unstructured PDF pages was impossible.

02. Architecture & Solution

Technical Implementation

Developed an automated data pipeline using PyMuPDF and LLM-backed section identification. Built an explicit page-level provenance tracking layer so that any extracted sentence in the final NLP dataset could be traced back directly to its exact page in the source PDF.

03. Measurable Impact

Results & Verification

Extracted 6.3M+ structured paragraphs across STOXX Europe 600 corporate filings, serving as the empirical foundation for SSRN 4763140.

04. Technology Stack
PythonLLM PipelinePyMuPDFPandasDuckDBNLP Classifier
mbilal.works · personal hub at bilalm.me · Impressum · PrivacyBuilt to a DESIGN.md spec · Bricolage Grotesque + JetBrains Mono · MMXXVI