Axtria
Project LeadCurrent
Apr 2025 — PresentOwn the architecture and delivery of a 50M-patient pharma data platform, and lead the team that builds it.
- 99.9%SLA
- 4Engineers led
- 8+Data products
- −60%Repeat failures
- Own end-to-end architecture of cloud-native pipelines on AWS (S3, Glue, EMR, Redshift), sustaining a 99.9% SLA.
- Lead 4 engineers through sprint planning, design reviews and code reviews; own the technical roadmap for 8+ data products and act as the client's primary technical contact.
- Consolidated fragmented ETL workflows into a reusable ingestion framework, cutting new-source onboarding from weeks to days across 10+ downstream consumers.
- Drove RCA on 20+ P1/P2 incidents and cut repeat failures by 60% with Grafana and Prometheus monitoring.
- AWS
- S3
- Glue
- EMR
- Redshift
- PySpark
- Grafana
- Prometheus
Senior Data Engineer
May 2024 — Apr 2025- 9h → 6hDaily runtime
- $40KSaved / year
- 15+Pipelines migrated
- 200MRecords / day
- Migrated 15+ legacy Hive/SQL pipelines to distributed PySpark on EMR. Daily runtime dropped from 9 to 6 hours and compute costs fell by $40K a year, from fixing data skew, partitioning and oversized shuffles at 200M records/day.
- Designed SCD Type-2 dimensional models that give 3+ years of patient-level historical auditability.
- Standardized multi-format ingestion (JSON, CSV, Parquet, ORC) through the Glue Data Catalog with schema-on-read, plus lineage and schema governance across 10+ upstream systems.
- PySpark
- EMR
- Glue
- Hive
- S3
- Parquet
- Python
Associate Data Engineer
Mar 2023 — Apr 2024- 12Zero-defect releases
- −80%Downstream defects
- 5Analytics use cases
- Built production ETL/ELT pipelines with PySpark, AWS Glue and Control-M, with zero defects across 12 consecutive releases.
- Added data quality gates (null, referential integrity, row count) that cut downstream defects by 80%.
- Modeled star-schema tables in Redshift for 5 analytics use cases.
- PySpark
- Glue
- Control-M
- Redshift
- SQL
- Python
Analyst
Mar 2022 — Mar 2023- 350hSaved / year
- 30+Dashboards
- Built a Python framework that automates Tableau refreshes across 30+ dashboards, removing 350 hours of manual work a year.
- Maintained PySpark and Control-M ingestion workflows with data quality validation for pharma commercial datasets.
- Python
- Tableau
- PySpark
- Control-M