Nineleaps · 2023
Sensitive-Field Rules
API Gateway and a document-store stream land in S3 through Kinesis and Firehose. A rules API and a field library write the curated prefix. S3 Batch replays history. Redshift is what analytics query.
The problem
The same fields arrived from APIs and from application tables. Each job decided which of those fields could pass through, so the landing files and the history did not match.
System design
01
API ingress
API Gateway sits in front of internal and external callers. Events go to Kinesis. Firehose writes them, as files, under an S3 landing prefix. Glue holds the catalog for that prefix. KMS covers the stream and the bucket.
02
Table changes
Application tables emit a change stream. That stream uses the same Kinesis and Firehose path and lands beside the API files. One landing prefix, two producers.
03
Rules API and field library
A small API stores the rule for each JSON field: leave it, encrypt it, or decrypt it. The Python library that copies an object reads those rules at runtime. The field list is not compiled into the job. Adding a field is an API call.
04
Curated prefix
The copy from the landing prefix to the curated prefix is where the library runs. API files and table changes go through the same function, so both producers get the same treatment.
05
History replay
An S3 Batch operation lists older landing objects, from both producers, and runs them through that function once. History is not left on the previous treatment.
06
Warehouse and the upstream copy
The curated prefix loads to Redshift, which is what analytics query. A separate path, Lake Formation in front of Athena, was proved for governed query and for deleting a subject's rows. Upstream databases replicate in with DMS, and alarms sit on those replication instances.
Key decisions
- Two producers, one landing prefix. API events and table changes do not get two pipelines.
- Field protection is data in a rules API, read by one library, not a branch in each job.
- KMS on the stream and the bucket. The field library is a second, selective layer on top of that.
- Replay older objects with S3 Batch through the same library. Do not write a second backfill job.
- Redshift serves analytics. The landing prefix is not a query surface.
Outcome
A new field was a rule. Landing files, table changes and history went through one library, and analytics queried the curated prefix in Redshift.
Architecture flow
- API Gateway
- Kinesis, then Firehose
- S3 landing prefix, Glue catalog
- Document-store stream, same path
- Rules API and field library
- S3 curated prefix
- S3 Batch replay
- Redshift
- DMS alarms
More work
Book a conversationEnterprise Data Foundation
A reusable data foundation for risk, products, analytics and AI, not one-off pipelines. Ingestion, medallion warehouse layers, domain marts and operational distribution from a single architectural spine.
Read case studyGlobal Risk Data Platform
Large-scale company risk-data architecture that ingests, enriches, normalizes and distributes company-level intelligence across countries. Built as a reusable risk pool for underwriting, products, analytics and intelligent systems, and today the data foundation beneath Cowbell's OMNI AI agents.
Read case studyGlueFlux
Metadata-driven processing framework. Pipelines defined in YAML so teams onboard Spark, Python and dbt workloads without rebuilding orchestration, deployment and operational patterns every time.
Read case study