Ext: Uber, via Nineleaps · 2018 - 2019
External Data Platform
Owned the pipelines that brought external data into Uber's ecosystem: ingestion, validation and normalisation at terabyte scale across billions of records, on Spark and Kafka.
The problem
Uber relied on data from outside its own systems. Every external source arrived with its own format, cadence and quality, and each consumer had been integrating it differently.
Key decisions
- Treat every external source as a contract, with schema and freshness checks at the boundary.
- One canonical model per entity, so consumers stop re-mapping the same feed.
- Stream where freshness matters, batch where volume does.
- Make data quality visible to consumers, not only to the pipeline owner.
Outcome
External data became a dependable, shared asset inside Uber's ecosystem instead of a collection of one-off integrations, running at terabyte scale across billions of records.
Architecture flow
- External and partner sources
- Ingestion: batch and streaming
- Validation and schema checks
- Normalisation into canonical models
- Curated external datasets
- Consumers across Uber's platforms
More architecture work
Book a conversationEnterprise Data Foundation
A reusable data foundation for risk, products, analytics and AI, not one-off pipelines. Ingestion, medallion warehouse layers, domain marts and operational distribution from a single architectural spine.
Read case studyGlobal Risk Data Platform
Large-scale company risk-data architecture that ingests, enriches, normalizes and distributes company-level intelligence across countries. Built as a reusable risk pool for underwriting, products, analytics and intelligent systems, and today the data foundation beneath Cowbell's OMNI AI agents.
Read case studyGlueFlux
Metadata-driven processing framework. Pipelines defined in YAML so teams onboard Spark, Python and dbt workloads without rebuilding orchestration, deployment and operational patterns every time.
Read case study