Nineleaps · 2023

Sensitive-Field Rules

API Gateway and a document-store stream land in S3 through Kinesis and Firehose. A rules API and a field library write the curated prefix. S3 Batch replays history. Redshift is what analytics query.

The problem

The same fields arrived from APIs and from application tables. Each job decided which of those fields could pass through, so the landing files and the history did not match.

System design

  1. 01

    API ingress

    API Gateway sits in front of internal and external callers. Events go to Kinesis. Firehose writes them, as files, under an S3 landing prefix. Glue holds the catalog for that prefix. KMS covers the stream and the bucket.

  2. 02

    Table changes

    Application tables emit a change stream. That stream uses the same Kinesis and Firehose path and lands beside the API files. One landing prefix, two producers.

  3. 03

    Rules API and field library

    A small API stores the rule for each JSON field: leave it, encrypt it, or decrypt it. The Python library that copies an object reads those rules at runtime. The field list is not compiled into the job. Adding a field is an API call.

  4. 04

    Curated prefix

    The copy from the landing prefix to the curated prefix is where the library runs. API files and table changes go through the same function, so both producers get the same treatment.

  5. 05

    History replay

    An S3 Batch operation lists older landing objects, from both producers, and runs them through that function once. History is not left on the previous treatment.

  6. 06

    Warehouse and the upstream copy

    The curated prefix loads to Redshift, which is what analytics query. A separate path, Lake Formation in front of Athena, was proved for governed query and for deleting a subject's rows. Upstream databases replicate in with DMS, and alarms sit on those replication instances.

Key decisions

  • Two producers, one landing prefix. API events and table changes do not get two pipelines.
  • Field protection is data in a rules API, read by one library, not a branch in each job.
  • KMS on the stream and the bucket. The field library is a second, selective layer on top of that.
  • Replay older objects with S3 Batch through the same library. Do not write a second backfill job.
  • Redshift serves analytics. The landing prefix is not a query surface.

Outcome

A new field was a rule. Landing files, table changes and history went through one library, and analytics queried the curated prefix in Redshift.

KinesisFirehoseS3KMSRedshift

Architecture flow

  1. API Gateway
  2. Kinesis, then Firehose
  3. S3 landing prefix, Glue catalog
  4. Document-store stream, same path
  5. Rules API and field library
  6. S3 curated prefix
  7. S3 Batch replay
  8. Redshift
  9. DMS alarms