mbilal.works
Available · September 2026
Talk
← Back to Homepage13 Systems • 5 Engineering Publications

Systems engineered & shipped.

Technical portfolio spanning production applications, BI middleware, custom AI pipelines, and practitioner-led engineering publications on arXiv & SSRN.

Production Systems, Apps & Infrastructure

Production SaaS products, client BI architectures, autonomous agent pipelines, and research data engineering infrastructure.

Advisory Engagement & Implementation

Unified Operations Middleware & BI Architecture

Architected custom Node.js middleware bridging Toggl, Absence.io, and internal CRM tool APIs into a centralized Power BI analytics engine for executive leadership.

Power BINode.jsExpressREST APIsTogglAbsence.ioRender
49% → 65%
Efficiency lift in 30 days (€50,000/yr annualised value)
Enterprise Platform & Analytics Architecture

Enterprise Salesforce & Tableau Analytics Architecture

Architected NVision's enterprise Salesforce Sales Cloud, Data Cloud, and Tableau BI analytics infrastructure from scratch across Europe & North America, establishing a single source of truth for master data with 1-click access to relevant commercial insights, custom automation flows, PandaDoc quote integrations, Apex/Lightning components, and automated ETL pipelines.

Salesforce Sales CloudData CloudTableauPandaDocEinstein Activity CaptureApexLightning ComponentsFlowsETL PipelineApps ScriptExcel
€84M+
Sales opportunity pipeline supported across Europe & North America
Production SaaS · Personal InfrastructureLIVE

Rental Intelligence & Custom Tuned AI Agents

Self-hosted Munich rental intelligence system with Custom Tuned AI Agents monitoring the rental market every 3–7 minutes, scoring ads via multi-stage LLM chains, and pushing instant Web/Mobile/Telegram app alerts.

Python 3.14FastAPIHTMXAlpine.jsSQLite (WAL)LiteLLMCloudflare Workers AI & Agents
1,812
Automated test cases verifying agent execution cadences & multi-provider chain
Applied AI System & EngineLIVE

Multi-Model LLM CV Tailoring Pipeline

Multi-pipeline agentic system that ingests job descriptions, extracts structured experience evidence, scores ATS keyword coverage, and generates tailored DOCX/LATEX CV drafts with real-time SSE progress streaming.

Python 3.14FastAPIsentence-transformersCustom AI AgentsPydanticpython-docx
28+
LLM models orchestrated across multiple providers with 5 pipeline architectures
AI Infrastructure & ToolingLIVE

Automated Job Search & Match Engine

Self-improving custom tuned AI Agents monitoring 60+ platforms including job boards and enterprise ATS platforms, scoring role match via MPNet bi-encoders and Thompson-sampling feedback loops.

Python 3.14Streamlitall-mpnet-base-v2Cross-EncoderSQLitePytest
60+ Platforms
Across 311 tracked enterprise ATS APIs with bi-encoder reranking
Data Engineering & Triage Infrastructure

Terabyte-Scale Google Drive to Notion Database Migrator

Google Drive → Notion byte-integrity migrator designed for unattended cloud deployment on lightweight VM infrastructure, executing 6-pass deterministic triage, property-based testing, and chunked uploads.

Python 3.14SQLiteNotion APIGoogle Drive APIHypothesismypy --strict
355,286
Files & folders processed across Terabyte-scale Drive inventory
iOS Widget Suite · Open SourceLIVE

LifeGrid — Live iOS Widget Suite

Eight independent Scriptable widgets that turn abstract 'time remaining' into a daily visual signal. Each one hand-rolls its own date math and gradient rendering with zero dependencies, and every widget below runs as a live port in the browser.

JavaScriptScriptable (iOS)DrawContext APINative iOS Widget Rendering
1:1 port
Every widget below runs the exact math and gradients from the original iOS build

Practitioner Engineering Publications

Practitioner-led empirical studies and open publications on arXiv and SSRN, synthesizing production learnings across software testing, BDD optimization, LLM verification, and AI energy economics.

Practitioner Engineering Publication · IEEE Access · SLROPEN ACCESS

LLM-Based Test Oracles: Source-of-Authority Taxonomy—A Systematic Literature Review

Co-authored systematic literature review (PRISMA 2020) of 54 studies on LLM-based test oracles, published in IEEE Access, introducing a source-of-authority taxonomy. Organises the field by where oracle verdicts derive their authority, a dimension missing from prior secondary studies.

Systematic ReviewPRISMA 2020LLM EvaluationTest Oracle AnalysisCohen's κ
54 Studies
Screened from 2,436 records under PRISMA 2020 with Cohen's κ = 0.79
Practitioner Engineering Publication · arXiv cs.SEOPEN ACCESS

All Green, Still Broken: Real-Flow Verification Lessons

arXiv paper reporting on a production LLM-integrated, multi-market rental-search assistant whose 1,553-test automated suite passed continuously yet shipped user-facing defects. Introduces a four-seam defect classification framework based on analysis of 252 bug-fix commits.

PythonLLM IntegrationMulti-Market TestingBrowser AutomationDefect Analysis
252 Bug-Fix Commits
Classified by seam: 44% escaped through 4 boundaries unit tests cannot observe
Practitioner Engineering Publication · SSRNOPEN ACCESS

The Governed Jevons Paradox of Artificial Intelligence: A Brief, Evidence-Anchored Synthesis of Mechanisms and Moderators in the AI–Energy–Economy Nexus

Evidence-anchored synthesis paper reconciling AI's dual energy narrative: while efficiency gains reliably cut energy and carbon intensity per task, cheaper compute triggers Jevons rebound effects that expand absolute energy demand.

SSRN PaperJevons ParadoxAI Energy NexusMacroeconomic ModelingPolicy & Grid Governance
4 Mechanisms
Reconciling AI energy intensity reductions against absolute compute demand rebound across 6 moderators
Practitioner Engineering Publication · arXiv cs.SEOPEN ACCESS

Given, When, Then, Again: Mining Subscenario Refactoring Candidates in Behaviour-Driven Test Suites with ML Classifiers and LLM-Judge Baselines

Co-authored arXiv paper mining BDD subscenario refactoring candidates across 5.3M slices using ML classifiers benchmarked against LLM judges. XGBoost classifier (F₁ = 0.891) beat a tuned rule baseline and two open-weight LLMs (GPT and Ling) at p < 10⁻⁴.

PythonXGBoostSBERTUMAPHDBSCANStatistical Testing
F₁ = 0.891
XGBoost classifier beating two open-weight LLM judges at p < 10⁻⁴
Practitioner Engineering Publication · arXiv cs.SE / cs.NIOPEN ACCESS

Software Testing at the Network Layer: Automated HTTP API Quality Assessment and Security Analysis of Production Web Applications

Co-authored arXiv paper presenting an automated HTTP API quality assessment framework. Playwright-based browser instrumentation captured 108 HAR files across 18 production websites, applying 8 heuristic anti-pattern detectors to produce a composite quality score per site.

Python 3.10+PlaywrightNode.js 18+HAR AnalysisSecurity Heuristics
108 HAR Files
Across 18 production sites, 11 categories, 175 MB of captured traffic
mbilal.works · personal hub at bilalm.me · Impressum · PrivacyBuilt to a DESIGN.md spec · Bricolage Grotesque + JetBrains Mono · MMXXVI