Cowbell · 2024 - Present

GlueFlux

Metadata-driven processing framework. Pipelines defined in YAML so teams onboard Spark, Python and dbt workloads without rebuilding orchestration, deployment and operational patterns every time.

The problem

Pipelines were expensive and repetitive to build and run. Each one re-implemented orchestration, logging, retries and deployment.

Key decisions

  • Configuration over code for onboarding new workloads.
  • One router for Spark, Python and dbt job types.
  • Environment-aware execution so the same config runs across stages.
  • Shared operational patterns instead of per-pipeline conventions.

Outcome

Started as a fix for one ingestion problem and became the standard framework for data workloads, cutting processing cost for that workload category by more than 95%.

YAMLGlueECSSparkdbtPlatform

Architecture flow

  1. YAML configuration
  2. Framework parser
  3. Job type router
  4. Spark and Python on AWS Glue
  5. dbt on ECS
  6. Standard logging and operations