Srisurya Chunchu
Data Engineer • Atlanta, GA • s******************@gmail.com • 404****115 • drivetube.ai/•••••
Professional Summary
Data Engineer with 5+ years of experience designing cloud-native ETL/ELT and streaming pipelines across banking, healthcare, and financial services. Built production-grade medallion-architecture pipelines on Azure Databricks/ADF and AWS Glue/EMR processing multi-TB daily, delivered sub-90s fraud-detection latency, engineered ML feature datasets, and enforced SOX/PCI/PHI compliance for regulated domains.
Technical Skills
Programming Language: Python,Scala,Shell Scripting,SQL
Databases: Azure Synapse,Amazon Redshift,Google BigQuery,Azure Synapse Analytics
Cloud Platforms: Microsoft Azure,ADLS Gen2,Amazon Web Services,S3,AWS Glue,Kinesis,Google Cloud Platform
DevOps & Infrastructure: Continuous Integration,Continuous Deployment,Docker,Kubernetes,Terraform,Jenkins,Azure DevOps
API & Integrations: RESTful APIs,Git
Messaging & Monitoring: Azure Event Hubs,Apache Kafka
Data Engineering & Processing: Amazon EMR,Apache Spark,Apache Hadoop,Apache Hive,Azure Databricks,Delta Lake,Parquet,Avro,ORC,PySpark,Apache Airflow,Medallion Architecture
Data Warehousing: Data Modeling,Star Schema,Lakehouse
Data Analysis & Visualization: Power BI,Tableau,Looker
Machine Learning & AI: Machine Learning
MLOps: Feature Engineering,MLOps,Azure Machine Learning,Amazon SageMaker
Data Integration & ETL: Azure Data Factory,Google Cloud Dataflow,Change Data Capture
Work Experience
Bank of America
Atlanta, GA
Azure Data Engineer
May 2025 – Present
Banking/financial services — built data engineering pipelines supporting transaction processing, fraud detection, payments and finance analytics.
Tech Stack: Azure Data Factory, Databricks, Delta Lake, ADLS Gen2, Azure Synapse Analytics, Azure Event Hubs, PySpark, Azure Monitor, Power BI
- Built Azure Data Factory and Databricks pipelines using Bronze/Silver/Gold medallion architecture to ingest and transform transaction and account data, processing 5+ TB daily to support fraud, payments and finance reporting.
- Developed PySpark streaming jobs on Azure Event Hubs to process real-time card authorization events, reducing fraud-detection alert latency from ~4 minutes to under 90 seconds.
- Designed curated Gold-layer data marts in Synapse and Delta Lake for GL, payments, and Customer 360 workloads, improving downstream BI query performance by 40%.
- Optimized Databricks Spark jobs on 500M+ row transaction tables via partitioning, join reordering, compaction and small-file management—cutting job runtimes by 45% and lowering compute costs.
- Implemented column-level encryption, dynamic data masking and role-based access controls across Synapse and ADLS Gen2 to achieve SOX and PCI-DSS compliance requirements.
- Built an Azure Monitor-backed observability layer and PySpark health checks across 20+ production workflows to track pipeline health and data freshness, reducing incident response time by 30%.
Wellstar Health System
Atlanta, GA
AWS Data Engineer
Dec 2024 – Apr 2025
Healthcare — built data pipelines and a data warehouse to consolidate Epic EHR and claims data for population health and operational analytics.
Tech Stack: AWS Glue, Apache Airflow, Amazon S3, Amazon Redshift, EMR, PySpark, Kafka, Kinesis Firehose, Power BI
- Built end-to-end ETL pipelines on AWS Glue and Airflow to ingest Epic EHR clinical data from multiple hospitals into a centralized S3 data lake, consolidating patient records for analytics.
- Designed Redshift schemas, tuned WLM and revised distribution keys to improve complex analytics query performance by 40%.
- Developed PySpark jobs on EMR to process structured claims and encounter data, producing encounter-level datasets for revenue cycle and managed care analytics.
- Implemented Kafka producers and Kinesis Firehose to stream admission/discharge events, enabling bed-management dashboards with sub-5-minute data freshness.
- Built a PHI-aware data quality framework with validation and referential-integrity checks that identified a patient-record duplication issue causing 25% of downstream matching failures.
- Partnered with clinical informatics and analytics stakeholders to operationalize Power BI dashboards for care-gap reporting and operational KPIs used by site leadership.
Fidelity Information Services (FIS)
Bengaluru, India
Data Engineer
Jun 2022 – Dec 2023
Financial services / payments — developed high-throughput pipelines for ACH, wire and card settlement feeds and near-real-time payment reporting.
Tech Stack: PySpark, Apache Airflow, Google BigQuery, Google Pub, Sub, Google Dataflow, Looker, Tableau
- Built PySpark and Airflow pipelines to ingest ACH, wire and card settlement feeds into Google BigQuery, supporting peak volumes of 50M+ transactions per day.
- Developed a real-time ingestion layer using Pub/Sub and Dataflow, replacing a T+1 batch reconciliation flow with near-real-time payment reporting.
- Redesigned BigQuery schemas with date-based partitioning and clustering to reduce average query costs for payments analytics by 50%.
- Integrated data from multiple core banking and risk systems into a unified transaction model to reconcile discrepancies between operations and finance reporting.
- Delivered Looker and Tableau dashboards tracking chargebacks, authorization declines and interchange revenue for product and commercial teams.
- Maintained SLA adherence for end-of-day settlement pipelines by implementing reprocessing routines and coordinating schema changes with upstream feed providers.
Johnson & Johnson (via Cognizant)
Hyderabad, India
Data Engineer
Jan 2022 – May 2022
Pharmaceuticals / clinical research — integrated clinical trial, safety and supply chain data to support safety reporting and post-market analytics.
Tech Stack: Apache Spark, Apache Airflow, Hadoop, Hive, SAP
- Built Spark and Airflow pipelines to integrate clinical trial datasets, adverse event reports and SAP supply chain feeds into a Hadoop data lake, processing 10M+ records monthly.
- Rewrote Hive partitioning strategy and optimized SQL/Hive queries for safety reporting, cutting nightly report generation time by 35%.
- Implemented validation and reconciliation checks against source systems to ensure accuracy of adverse event counts and patient exposure metrics for regulatory submissions.
- Enforced HIPAA-compliant data handling with field-level de-identification, audit logging and role-based access controls for PHI datasets.
- Collaborated with data science to produce feature datasets for a post-market drug-safety signal detection model, standardizing variables and timestamps.
- Automated batch scheduling and added Airflow-based monitoring and alerting to improve pipeline SLAs for safety and regulatory reporting.
General Insurance Corporation of India
Hyderabad, India
Associate Data Engineer
Mar 2020 – Dec 2021
Insurance — consolidated reinsurance, premium and claims data to support actuarial, underwriting and finance analytics.
Tech Stack: PySpark, Hive, Power BI, Tableau, Shell Scripting
- Designed PySpark and Hive ETL pipelines to consolidate reinsurance treaty data, premium ledgers and claims records across motor, health and property lines into a central data warehouse.
- Optimized Hive partition design and query plans across 3+ years of claims history, reducing the nightly batch window by 30%.
- Built Power BI and Tableau dashboards tracking loss ratios, claims frequency and premium-to-claims trends used in quarterly reviews by senior underwriters.
- Automated recurring reconciliations and data extract jobs using Shell scripting, saving ~15 hours of manual effort per week for the reporting team.
- Developed idempotent backfill and reprocessing routines to ensure data consistency across historical claims data seasons.
- Provided production support including monitoring, troubleshooting and root-cause analysis to maintain consistent data delivery for actuarial and finance teams.
Projects
Real-Time Fraud Detection Pipeline
Tools Used: Azure Event Hubs, Databricks, PySpark, Delta Lake
- Built end-to-end streaming pipeline on Azure Event Hubs and Databricks using PySpark and Delta Lake to detect fraudulent card transactions, reducing alert latency by ~65%.
- Implemented Gold-layer feature tables and served datasets enabling downstream ML scoring and BI consumption with consistent data freshness guarantees.
Clinical Data Lake on AWS
Tools Used: Amazon S3, AWS Glue, Apache Airflow, Amazon Redshift
- Designed and deployed a scalable S3-based clinical data lake ingesting EHR and claims data using AWS Glue and Airflow to enable population health analytics.
- Built a Redshift data warehouse and optimized distribution and WLM to deliver ~40% faster analytic query performance for clinical reporting.
Education
Georgia State University
Master of Science, Computer Science • Atlanta, GA
Certifications
AWS Certified Data Engineer Associate — Amazon Web Services
Azure Data Engineer Associate (DP-203) — Microsoft
Databricks Certified Data Engineer — Databricks
Powered by Drivetube · Create your own profile at drivetube.ai