Srisurya Chunchu
Professional Summary
Data Engineer with 5+ years of experience designing cloud-native ETL/ELT and streaming pipelines across banking, healthcare, and financial services. Built production-grade medallion-architecture pipelines on Azure Databricks/ADF and AWS Glue/EMR processing multi-TB daily, delivered sub-90s fraud-detection latency, engineered ML feature datasets, and enforced SOX/PCI/PHI compliance for regulated domains.
Technical Skills
Work Experience
- Built Azure Data Factory and Databricks pipelines using Bronze/Silver/Gold medallion architecture to ingest and transform transaction and account data, processing 5+ TB daily to support fraud, payments and finance reporting.
- Developed PySpark streaming jobs on Azure Event Hubs to process real-time card authorization events, reducing fraud-detection alert latency from ~4 minutes to under 90 seconds.
- Designed curated Gold-layer data marts in Synapse and Delta Lake for GL, payments, and Customer 360 workloads, improving downstream BI query performance by 40%.
- Optimized Databricks Spark jobs on 500M+ row transaction tables via partitioning, join reordering, compaction and small-file management—cutting job runtimes by 45% and lowering compute costs.
- Implemented column-level encryption, dynamic data masking and role-based access controls across Synapse and ADLS Gen2 to achieve SOX and PCI-DSS compliance requirements.
- Built an Azure Monitor-backed observability layer and PySpark health checks across 20+ production workflows to track pipeline health and data freshness, reducing incident response time by 30%.
- Built end-to-end ETL pipelines on AWS Glue and Airflow to ingest Epic EHR clinical data from multiple hospitals into a centralized S3 data lake, consolidating patient records for analytics.
- Designed Redshift schemas, tuned WLM and revised distribution keys to improve complex analytics query performance by 40%.
- Developed PySpark jobs on EMR to process structured claims and encounter data, producing encounter-level datasets for revenue cycle and managed care analytics.
- Implemented Kafka producers and Kinesis Firehose to stream admission/discharge events, enabling bed-management dashboards with sub-5-minute data freshness.
- Built a PHI-aware data quality framework with validation and referential-integrity checks that identified a patient-record duplication issue causing 25% of downstream matching failures.
- Partnered with clinical informatics and analytics stakeholders to operationalize Power BI dashboards for care-gap reporting and operational KPIs used by site leadership.
- Built PySpark and Airflow pipelines to ingest ACH, wire and card settlement feeds into Google BigQuery, supporting peak volumes of 50M+ transactions per day.
- Developed a real-time ingestion layer using Pub/Sub and Dataflow, replacing a T+1 batch reconciliation flow with near-real-time payment reporting.
- Redesigned BigQuery schemas with date-based partitioning and clustering to reduce average query costs for payments analytics by 50%.
- Integrated data from multiple core banking and risk systems into a unified transaction model to reconcile discrepancies between operations and finance reporting.
- Delivered Looker and Tableau dashboards tracking chargebacks, authorization declines and interchange revenue for product and commercial teams.
- Maintained SLA adherence for end-of-day settlement pipelines by implementing reprocessing routines and coordinating schema changes with upstream feed providers.
- Built Spark and Airflow pipelines to integrate clinical trial datasets, adverse event reports and SAP supply chain feeds into a Hadoop data lake, processing 10M+ records monthly.
- Rewrote Hive partitioning strategy and optimized SQL/Hive queries for safety reporting, cutting nightly report generation time by 35%.
- Implemented validation and reconciliation checks against source systems to ensure accuracy of adverse event counts and patient exposure metrics for regulatory submissions.
- Enforced HIPAA-compliant data handling with field-level de-identification, audit logging and role-based access controls for PHI datasets.
- Collaborated with data science to produce feature datasets for a post-market drug-safety signal detection model, standardizing variables and timestamps.
- Automated batch scheduling and added Airflow-based monitoring and alerting to improve pipeline SLAs for safety and regulatory reporting.
- Designed PySpark and Hive ETL pipelines to consolidate reinsurance treaty data, premium ledgers and claims records across motor, health and property lines into a central data warehouse.
- Optimized Hive partition design and query plans across 3+ years of claims history, reducing the nightly batch window by 30%.
- Built Power BI and Tableau dashboards tracking loss ratios, claims frequency and premium-to-claims trends used in quarterly reviews by senior underwriters.
- Automated recurring reconciliations and data extract jobs using Shell scripting, saving ~15 hours of manual effort per week for the reporting team.
- Developed idempotent backfill and reprocessing routines to ensure data consistency across historical claims data seasons.
- Provided production support including monitoring, troubleshooting and root-cause analysis to maintain consistent data delivery for actuarial and finance teams.
Projects
- Built end-to-end streaming pipeline on Azure Event Hubs and Databricks using PySpark and Delta Lake to detect fraudulent card transactions, reducing alert latency by ~65%.
- Implemented Gold-layer feature tables and served datasets enabling downstream ML scoring and BI consumption with consistent data freshness guarantees.
- Designed and deployed a scalable S3-based clinical data lake ingesting EHR and claims data using AWS Glue and Airflow to enable population health analytics.
- Built a Redshift data warehouse and optimized distribution and WLM to deliver ~40% faster analytic query performance for clinical reporting.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Everything on Drivetube
Six products, two of them free forever. Start wherever you are.
Your free digital resume
at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept.
Documentation- Free Job BoardFree
Verified jobs, posted in the last 3 days
verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed.
Documentation - Job Hunt ProgramFrom $199.99
We run your job hunt for you
managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you.
Documentation - Resume Writing ServicesFrom $25.99
Written by senior career writers
human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile.
Documentation - Community MembershipFrom $49.99/yr
AI tools, gated filters and a $10,000+ library
from $49.99/year (₹1,999/year in India), or $4.99/month in the iOS app. Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates.
Documentation - Drivetube HireFree tier
For employers — no job postings, no applications
hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do.
Documentation
- Documentation
every product explained in full, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
- Playbooks
tactical job-search guides. Each post is one named tactic with the situation it applies to, the exact steps, a copy-paste message and honest failure modes. Free, no account.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.
Pricing is identical on the web, iOS and Android — no app-store markup. Monthly Community membership is purchased in the iOS app only; the website sells annual plans. The Android app has no in-app purchase, so Android members subscribe on drivetube.ai and then sign in to the app with full access. Membership follows your account, not your device.