Architecture

Problems solved, not tools listed.

Architecture problems, systems and decisions across twelve years: from master data and validation frameworks to external data at Uber scale and the foundations behind AI.

A reusable data foundation for risk, products, analytics and AI, not one-off pipelines. Ingestion, medallion warehouse layers, domain marts and operational distribution from a single architectural spine.

AWSRedshiftGlueS3KinesisdbtTerraform

Large-scale company risk-data architecture that ingests, enriches, normalizes and distributes company-level intelligence across countries. Built as a reusable risk pool for underwriting, products, analytics and intelligent systems, and today the data foundation beneath Cowbell's OMNI AI agents.

Risk DataEnrichmentMulti-CountryData Marts

Cowbell · 2024 - Present

GlueFlux

Read case study

Metadata-driven processing framework. Pipelines defined in YAML so teams onboard Spark, Python and dbt workloads without rebuilding orchestration, deployment and operational patterns every time.

YAMLGlueECSSparkdbtPlatform

Redshift as a reusable enterprise warehouse: Bronze, Silver, Gold and domain marts with dbt models, separating ingestion, normalization, business modeling and consumption.

RedshiftdbtMedallionData Marts

Moving curated data to regional applications, search and APIs while respecting residency, PII, latency and operational independence. Distribution as a platform capability, not a copy job.

Multi-RegionDMSPostgreSQLGovernance

Trusted data and platform foundations that enable intelligent product experiences: governed flows, grounding and operational access for agents and AI-driven insurance workflows.

Agentic AIRAGGroundingLLMs

Ext: Uber, via Nineleaps · 2018 - 2019

External Data Platform

Read case study

Owned the pipelines that brought external data into Uber's ecosystem: ingestion, validation and normalisation at terabyte scale across billions of records, on Spark and Kafka.

SparkKafkaPythonExternal DataData Quality

Data and model-computation pipelines behind Blue Yonder's supply-chain AI: preparing enterprise data, running ML models for demand forecasting and delivering predictions at scale on Azure.

AzurePySparkSynapseMachine LearningCI/CD

ATMS and UDL: Python and AWS frameworks that use compiler-style models to generate every feature combination a processor must pass before release, so coverage is designed, not guessed.

PythonAWSTest AutomationCompiler DesignCombinatorics

American Megatrends · 2014 - 2017

Master Data & Release Platform

Read case study

Centralised scattered data into a master data system with regional replication that respected PII and data laws, and built AMISVN, a Python version-control layer on Subversion with release automation.

PythonMDMReplicationSubversionAutomation