Open to Lead & Senior Data Engineering roles

Data platforms that hold up at 200M records a day.

I'm Pronnoy Dutta, a Lead Data Engineer with 6+ years building cloud-native lakehouse platforms on AWS and Databricks. I own the architecture, lead the team, and ship pipelines that finish on time, cost less, and don't page anyone at 3 AM.

  • Databricks Data Engineer Professional
  • AWS Security — Specialty
  • AWS Solutions Architect — Associate
pharma_commercial_dailyHealthySLA 99.9%
delta lake · medallionsourcesrx_claimsparquetcrm_eventsjsonsales_opscsv · orcIngestionframeworkGlue · SparkBronzerawSilvercleanedGoldmodeled✓ schema✓ nulls✓ RI✓ countsdata quality gatesserveRedshiftwarehouseDashboardsBI · 30+Products10+ users
06:00:04ingest 10+ sources landed · ~200M rows
  • 6+YearsBuilding production data platforms
  • 200M+Records / dayProcessed on AWS & Spark
  • 99.9%SLAOn a 50M-patient platform
  • $40KSaved / yearIn cloud compute
  • 30%FasterDaily runtime, 9h → 6h
  • 60%Fewer repeatsOf P1/P2 incident failures

01About

Architecture, delivery, and the team behind both.

I lead data engineering for a pharma commercial analytics platform at Axtria that serves 50M patients. I own the whole system: the architecture on AWS, a team of four engineers, the roadmap for 8+ data products, and the client conversation when something matters.

Before that I moved legacy Hive workloads onto Spark and designed the SCD Type-2 models that keep 3+ years of history auditable. At Infosys I built a retail warehouse that became the executive team's single source of truth.

Role
Project Lead, Axtria
Team
4 data engineers
Platform
50M patients · 99.9% SLA
Based in
Gurugram, India · open to remote

How I work

  1. Platforms, not pipelines

    Build the framework once. A new source should be a config change, not a project.

  2. Observable by default

    If a pipeline can fail silently, it will. Quality gates and monitoring ship with the code.

  3. Model for the audit

    History is a feature. Every number on a dashboard should be traceable to its source.

  4. Cost is a metric

    Every shuffle has a bill. Spark tuning is cloud spend, so I treat it like one.

  5. Teams scale, heroes don't

    Design reviews, code reviews and clear ownership over late-night heroics.

02Experience

Three promotions in three years.

Analyst to Project Lead at Axtria, after starting out in retail analytics at Infosys.

  1. 2022Analyst
  2. 2023Associate DE
  3. 2024Senior DE
  4. 2025Project Lead

Axtria

Pharma commercial analytics · Gurugram, India

Mar 2022 — Present
  1. Project LeadCurrent

    Apr 2025 — Present

    Own the architecture and delivery of a 50M-patient pharma data platform, and lead the team that builds it.

    • 99.9%SLA
    • 4Engineers led
    • 8+Data products
    • −60%Repeat failures
    • Own end-to-end architecture of cloud-native pipelines on AWS (S3, Glue, EMR, Redshift), sustaining a 99.9% SLA.
    • Lead 4 engineers through sprint planning, design reviews and code reviews; own the technical roadmap for 8+ data products and act as the client's primary technical contact.
    • Consolidated fragmented ETL workflows into a reusable ingestion framework, cutting new-source onboarding from weeks to days across 10+ downstream consumers.
    • Drove RCA on 20+ P1/P2 incidents and cut repeat failures by 60% with Grafana and Prometheus monitoring.
    • AWS
    • S3
    • Glue
    • EMR
    • Redshift
    • PySpark
    • Grafana
    • Prometheus
  2. Senior Data Engineer

    May 2024 — Apr 2025
    • 9h → 6hDaily runtime
    • $40KSaved / year
    • 15+Pipelines migrated
    • 200MRecords / day
    • Migrated 15+ legacy Hive/SQL pipelines to distributed PySpark on EMR. Daily runtime dropped from 9 to 6 hours and compute costs fell by $40K a year, from fixing data skew, partitioning and oversized shuffles at 200M records/day.
    • Designed SCD Type-2 dimensional models that give 3+ years of patient-level historical auditability.
    • Standardized multi-format ingestion (JSON, CSV, Parquet, ORC) through the Glue Data Catalog with schema-on-read, plus lineage and schema governance across 10+ upstream systems.
    • PySpark
    • EMR
    • Glue
    • Hive
    • S3
    • Parquet
    • Python
  3. Associate Data Engineer

    Mar 2023 — Apr 2024
    • 12Zero-defect releases
    • −80%Downstream defects
    • 5Analytics use cases
    • Built production ETL/ELT pipelines with PySpark, AWS Glue and Control-M, with zero defects across 12 consecutive releases.
    • Added data quality gates (null, referential integrity, row count) that cut downstream defects by 80%.
    • Modeled star-schema tables in Redshift for 5 analytics use cases.
    • PySpark
    • Glue
    • Control-M
    • Redshift
    • SQL
    • Python
  4. Analyst

    Mar 2022 — Mar 2023
    • 350hSaved / year
    • 30+Dashboards
    • Built a Python framework that automates Tableau refreshes across 30+ dashboards, removing 350 hours of manual work a year.
    • Maintained PySpark and Control-M ingestion workflows with data quality validation for pharma commercial datasets.
    • Python
    • Tableau
    • PySpark
    • Control-M

Infosys

Retail analytics · Remote

Nov 2020 — Mar 2022
  1. Systems Engineer

    Nov 2020 — Mar 2022
    • 3d → 2hReporting cycle
    • 500+Stores
    • 12+KPIs
    • −40%Report time
    • Built monthly sales analytics pipelines on S3, Python and PySpark for a retail client with 500+ stores across 3 regions, cutting reporting from 3 days to under 2 hours.
    • Designed a star-schema warehouse (8 fact and dimension tables) and KPI logic for 12+ metrics, which became the single source of truth for executive dashboards.
    • Optimized SQL on 100M+ row transaction tables, making report generation 40% faster.
    • S3
    • Python
    • PySpark
    • SQL

03Selected work

Problems worth solving.

Four engagements: what was broken, what I changed, and what it was worth.

measured

Performance · Cost

Cutting a 9-hour batch window to 6

Problem
15+ legacy Hive/SQL pipelines processing 200M records a day were eating the batch window and the cloud budget.
Approach
Rebuilt them as distributed PySpark on EMR. Profiled every stage, fixed skewed keys and bad partitioning, and removed oversized shuffles.
  • 30%Faster
  • $40K/yrCompute saved
  • PySpark
  • EMR
  • Hive
  • S3
oneframeworksourcesconsumersillustrative

Platform · Architecture

A reusable ingestion framework

Problem
Every new data source meant a new hand-built ETL workflow. Onboarding took weeks and maintenance kept growing.
Approach
Consolidated the fragmented workflows into one standardized framework with shared patterns for landing, validation and schema handling.
  • Weeks → daysSource onboarding
  • 10+Consumers served
  • Glue
  • PySpark
  • S3
  • Redshift
Illustrative SCD Type-2 dimension: each territory change adds a new versioned row
hcp_idterritoryvalid_fromvalid_tocurrent
10482NORTH-012022-01-012023-06-30false
10482NORTH-042023-07-012024-12-31false
10482CENTRAL-022025-01-019999-12-31true
illustrative

Data modeling · Governance

3+ years of patient history, fully auditable

Problem
Analysts needed to answer “what did we know, and when?”, but overwrites destroyed history.
Approach
Designed SCD Type-2 dimensional models with lineage tracking and schema governance across 10+ upstream source systems.
  • 3+ yrsPoint-in-time history
  • 10+Governed sources
  • PySpark
  • Redshift
  • Glue
  • SQL
illustrative

Reliability · Leadership

Fewer 3 AM pages

Problem
The same production failures kept coming back on a platform with a 99.9% SLA commitment.
Approach
Led root-cause analysis on 20+ P1/P2 incidents, then turned each finding into Grafana/Prometheus monitoring and alerts.
  • −60%Repeat failures
  • 99.9%SLA sustained
  • Grafana
  • Prometheus
  • AWS

04Stack

The tools I build with.

Highlighted tiles are the ones I use every day.

Lakehouse & Big Data

  • Databricks(daily driver)
  • Apache Spark(daily driver)
  • Delta Lake(daily driver)
  • Medallion(daily driver)
  • Hadoop
  • Hive
  • YARN
  • Parquet

AWS Cloud

  • S3(daily driver)
  • Glue(daily driver)
  • EMR(daily driver)
  • Redshift(daily driver)
  • Lambda
  • IAM
  • DynamoDB
  • Terraform

Languages

  • Python(daily driver)
  • SQL(daily driver)
  • PySpark(daily driver)
  • FastAPI
  • Java
  • C#
  • Bash
  • Linux

Orchestration & DevOps

  • Airflow(daily driver)
  • Control-M
  • GitHub Actions
  • Git
  • Docker
  • Kubernetes
  • Grafana
  • Prometheus

Modeling & Governance

  • Dimensional modeling(daily driver)
  • SCD Type-2(daily driver)
  • Star schema(daily driver)
  • Data quality(daily driver)
  • Data lineage
  • Schema governance
  • CDC
  • Batch & stream

05Credentials

Certified where it counts.

Amazon Web Services

Security — Specialty

Amazon Web Services

Solutions Architect — Associate

Cisco

CCNA

Bharati Vidyapeeth College of Engineering, Pune · 2016 — 2020

B.Tech, Computer Science

Official member, Security team

AWS Community Builders

06Contact

Building a data platform, or fixing one?

I'm open to Lead and Senior Data Engineering roles, remote or in Gurugram, Bangalore, Hyderabad or Pune. Email is the fastest way to reach me.

Gurugram, India