binbash

AWS Data Platforms

Data Modernization & Analytics

Lakehouses, pipelines and warehouses on AWS — governed, portable, and ready for the agents that read them.

WHAT IT IS

Your data, in one place your engineers and your agents can both use

Most teams do not have a data problem. They have six data problems in six systems, and no single answer to “what is actually true?”

What we do

Seven capabilities, one lakehouse

  • For Ximple, S3 is zoned raw, processed and curated — partitioned and compressed as Parquet with Snappy, with a Glue Data Catalog tying the zones together. InstaCrops needed the same shape of foundation under its own zone names — Raw-Bronze, Silver and Gold — which binbash designed as an advisory engagement: the client’s own engineers built it.

  • For Ximple, ingestion runs on Glue Crawlers and Glue-plus-DMS pipelines, with DynamoDB Streams feeding Firehose for incremental, change-data-capture loads. Celes’s ingestion crosses clouds entirely: a DataSync agent moves data from GCP into AWS.

  • For Ximple, Redshift Spectrum queries S3 directly through an external Glue-catalog database, with materialized views on top. Vistapath runs on Redshift Serverless. Celes needed to stay portable across clouds, so its warehouse is StarRocks on EKS instead — open source, no lock-in.

  • Vistapath’s dbt models run on Redshift, feeding QuickSight dashboards that are themselves templated as infrastructure-as-code — so a change ships as a pull request, not a ticket to a data team.

  • Vistapath’s computer-vision pipeline starts once PHI has been stripped: SageMaker endpoints de-duplicate images by feature vector and segment them, with pgvector on RDS Postgres holding those vectors. SQS feeds an EventBridge Pipe into Step Functions that routes each image to instalabel.ai for annotation, and the callbacks land as intermediate state in DynamoDB. A Glue ETL job then turns the finished annotations into tabular data for the warehouse to serve.

  • Celes’s cross-provider lake — Hudi on S3, StarRocks on EKS — is watched end to end: Prometheus and Grafana for metrics, Amazon OpenSearch and FluentBit for centralized logs. KEDA scales workloads to zero between jobs, and OpenBao holds the secrets none of it runs without.

  • Vistapath strips PHI in a production-account Lambda before any of it reaches the data-science account. Akua keeps PII and non-PII in separate S3 buckets, isolated across separate accounts, with least-privilege IAM throughout — including the Glue Data Catalog and Athena tables the cleaned results land in.

How we do it

Built the way we build everything

As-code, always

Every lakehouse ships from the Leverage reference architecture (le-tf-infra-aws) — reviewable, versioned, repeatable, built in OpenTofu first, with Terraform fully supported too.

Portable by default

Open formats on S3 — Parquet, Hudi — with StarRocks as the warehouse. No proprietary lock-in, so you can leave whenever you want.

Regulated-industry data

HCLS and fintech platforms carry PHI, PII and KYB data that has to stay isolated, encrypted and least-privilege from day one — not bolted on after the fact.

AWS-funded

AWS funding programs backed four of these five engagements, one of them in full. The fifth — Vistapath’s computer-vision pipeline — ran through AWS Partner Network, its design reviewed and approved by an AWS Solutions Architect.

Leverage, all the way down

Leverage is the reference architecture that lands your AWS foundation — and the same one that builds and runs the data platform on top of it, from account baseline to lakehouse, as code.

One lake, every downstream reader

Sources land in S3, get curated once, and are served to Athena, Redshift Spectrum and the agents that read it.

Reference architecture diagram: source systems flow through Glue, DMS and Kinesis Firehose into an S3 raw zone, get transformed by Glue ETL into a processed zone, and are cataloged once by AWS Glue Data Catalog — which fans out to Athena, to Redshift Spectrum through dbt into QuickSight, and to Bedrock, because the lake is what the agents read.

Delivered, not proposed

5 data engagements

Ximple, Celes, Vistapath, InstaCrops and Akua — lakehouse foundations, ingestion pipelines, a data warehouse, a computer-vision pipeline and a governed catalog.

5 countries

Mexico, Colombia, the United States, Chile and Uruguay — one client platform delivered in each.

5 architecture types

Data warehouse, data lake, data lakehouse, data pipelines and computer-vision data pipelines — the shapes we build, each one delivered for a named client.

Ready to modernize your data platform?

Let’s architect a lakehouse that scales with you.