Ext: Uber, via Nineleaps · 2018 - 2019

External Data Platform

Owned the pipelines that brought external data into Uber's ecosystem: ingestion, validation and normalisation at terabyte scale across billions of records, on Spark and Kafka.

The problem

Uber relied on data from outside its own systems. Every external source arrived with its own format, cadence and quality, and each consumer had been integrating it differently.

Key decisions

  • Treat every external source as a contract, with schema and freshness checks at the boundary.
  • One canonical model per entity, so consumers stop re-mapping the same feed.
  • Stream where freshness matters, batch where volume does.
  • Make data quality visible to consumers, not only to the pipeline owner.

Outcome

External data became a dependable, shared asset inside Uber's ecosystem instead of a collection of one-off integrations, running at terabyte scale across billions of records.

SparkKafkaPythonExternal DataData Quality

Architecture flow

  1. External and partner sources
  2. Ingestion: batch and streaming
  3. Validation and schema checks
  4. Normalisation into canonical models
  5. Curated external datasets
  6. Consumers across Uber's platforms