Skip to content

SANJANA K

Data Engineer • Austin, TX • s************@gmail.com • 737****074 • drivetube.ai/•••••

Professional Summary

Data Engineer with 4+ years of experience building scalable ETL/ELT pipelines and analytics infrastructure across healthcare, finance, and enterprise sectors. Proficient in Python, PySpark, SQL, Databricks and AWS; proven track record reducing pipeline latency, improving data quality and cutting cloud costs while maintaining compliance.

Technical Skills

Programming Languages: Python,Java
Web Technologies: REST APIs
Frameworks and Libraries: Pandas,NumPy
Databases: SQL,Snowflake
Cloud and DevOps: Amazon S3,Amazon Redshift,AWS Glue,AWS Lambda,EMR,Amazon Athena,Amazon DynamoDB,Amazon Aurora,AWS Kinesis,Step Functions,Lake Formation,CloudFormation,IAM,KMS,CloudTrail,OpenSearch,Terraform,Docker,Amazon EKS,AWS CodePipeline,CI,CD,Pipeline Monitoring
Data and Analytics: ETL,ELT Pipelines,Data Integration,Star Schema Modeling,Data Quality,Pipeline Orchestration,Change Data Capture CDC,Tableau,Power BI,Looker
Tools and Methodologies: Databricks,dbt,Apache Spark,Apache Airflow,Amazon MWAA,Data Lake,Git
Skills: PySpark,T-SQL
Streaming & Processing: Apache Kafka,Kafka Connect,Debezium,Apache Flink,Spark Structured Streaming,Amazon Kinesis

Work Experience

Cardinal Health
Austin, TX
Data Engineer
Feb 2025 – Present
Healthcare supply and services company — built data engineering solutions supporting hospital systems, patient records, reporting and compliance.
Tech Stack: Databricks, PySpark, Apache Kafka, Debezium, Amazon Kinesis, AWS Glue, Amazon S3, Amazon Redshift, Apache Airflow, Power BI, Tableau
  • Engineered HIPAA-compliant batch and real-time data pipelines for 12 hospital systems, processing 50M+ patient records monthly using Databricks, PySpark, AWS Glue and S3 with zero compliance violations.
  • Built real-time ingestion and CDC pipelines using Kafka, Debezium and Kinesis, reducing end-to-end data latency from 4 hours to under 5 minutes for operational analytics.
  • Designed a star-schema data warehouse on Amazon Redshift to power 35+ Power BI and Tableau dashboards, enabling teams to deliver reports 15% faster.
  • Orchestrated batch and streaming workflows in Apache Airflow with automated validation and alerting, achieving 99.5% on-time pipeline completion and faster incident detection.
  • Optimized SQL and Redshift performance through query tuning and partitioning, cutting execution time by 25% and reducing annual data warehouse costs by $120,000.
  • Resolved 30+ data-quality and pipeline incidents through root-cause analysis and monitoring improvements, eliminating 95% of recurring issues.
Mastercard
Austin, TX
AWS Data Engineer
Aug 2023 – Dec 2024
Global payments company — developed AWS-based data platforms and pipelines to process payment and settlement data for analytics and reporting.
Tech Stack: Amazon S3, AWS Glue, AWS Lambda, Step Functions, EMR, Spark, PySpark, Apache Kafka, Apache Flink, Amazon Kinesis, Apache Airflow, Amazon Redshift
  • Designed scalable AWS data architectures processing 15M+ payment and settlement records daily using S3, Glue, Lambda and Step Functions to standardize ingestion pipelines.
  • Built a data lake and ETL solutions on S3 and AWS Glue processing 2TB+ of data daily, reducing end-to-end processing time by 35%.
  • Developed batch processing pipelines with Spark and PySpark on EMR, reducing critical job runtimes from 4 hours to 90 minutes through optimization and parallelism.
  • Implemented real-time streaming and CDC pipelines with Kinesis, Kafka and Flink, cutting transaction-data latency from 30 minutes to under 2 minutes for downstream analytics.
  • Orchestrated 100+ production workflows on Apache Airflow / MWAA, achieving 99% on-time completion and reducing data-quality incidents by 40% through automated checks.
  • Optimized Spark jobs and Redshift workloads via tuning, partitioning and resource adjustments, improving performance by 30% and lowering monthly cloud costs by 20%.
PayPal
Austin, TX
Data Engineer
Sep 2021 – Dec 2022
Online payments company — built high-throughput ETL/ELT pipelines and reporting infrastructure for financial risk and regulatory reporting.
Tech Stack: PySpark, Apache Spark, Apache Airflow, SQL, T-SQL, Amazon S3, Amazon Redshift
  • Built high-performance ETL pipelines handling 1B+ daily transactions to support financial risk analytics and regulatory reporting across multiple regulatory bodies.
  • Optimized complex SQL and T-SQL queries, cutting data retrieval latency by 50% and reducing risk report generation time from 4 hours to 2 hours.
  • Implemented ELT transformations processing 200GB+ of data daily to enable advanced analytics across six business units.
  • Maintained 99.9% system uptime across production pipelines serving 5,000+ daily business users through proactive monitoring and runbook improvements.
  • Implemented data validation and monitoring using Airflow sensors and custom PySpark checks, reducing production incidents and false positives.
  • Collaborated with data science and compliance teams to deliver traceable data lineage and schema evolution support, enabling timely regulatory submissions.

Education

Southern Arkansas University
Master of Science in Computer Science (GPA: 3.67) • Magnolia, AR
Aurora's Institute of Science and Technology
Bachelor of Science in Computer Science • India

Certifications

AWS Certified Cloud Practitioner — Amazon Web Services

Powered by Drivetube · Create your own profile at drivetube.ai