Cowbell · 2024 - Present
GlueFlux
Metadata-driven processing framework. Pipelines defined in YAML so teams onboard Spark, Python and dbt workloads without rebuilding orchestration, deployment and operational patterns every time.
The problem
Pipelines were expensive and repetitive to build and run. Each one re-implemented orchestration, logging, retries and deployment.
Key decisions
- Configuration over code for onboarding new workloads.
- One router for Spark, Python and dbt job types.
- Environment-aware execution so the same config runs across stages.
- Shared operational patterns instead of per-pipeline conventions.
Outcome
Started as a fix for one ingestion problem and became the standard framework for data workloads, cutting processing cost for that workload category by more than 95%.
Architecture flow
- YAML configuration
- Framework parser
- Job type router
- Spark and Python on AWS Glue
- dbt on ECS
- Standard logging and operations
More architecture work
Book a conversationEnterprise Data Foundation
A reusable data foundation for risk, products, analytics and AI, not one-off pipelines. Ingestion, medallion warehouse layers, domain marts and operational distribution from a single architectural spine.
Read case studyGlobal Risk Data Platform
Large-scale company risk-data architecture that ingests, enriches, normalizes and distributes company-level intelligence across countries. Built as a reusable risk pool for underwriting, products, analytics and intelligent systems, and today the data foundation beneath Cowbell's OMNI AI agents.
Read case studyMedallion Data Warehouse
Redshift as a reusable enterprise warehouse: Bronze, Silver, Gold and domain marts with dbt models, separating ingestion, normalization, business modeling and consumption.
Read case study