Skip to content

SANJANA K

Data Engineer • Austin, TX • s************@gmail.com • 737****074 • drivetube.ai/•••••

Professional Summary

Data Engineer with 4+ years of experience building scalable ETL/ELT pipelines and analytics infrastructure across healthcare, finance, and enterprise sectors. Proficient in Python, PySpark, SQL, Databricks and AWS; proven track record reducing pipeline latency, improving data quality and cutting cloud costs while maintaining compliance.

Technical Skills

Programming Languages: Python,Java
Web Technologies: REST APIs
Frameworks and Libraries: Pandas,NumPy
Databases: SQL,Snowflake
Cloud and DevOps: Amazon S3,Amazon Redshift,AWS Glue,AWS Lambda,EMR,Amazon Athena,Amazon DynamoDB,Amazon Aurora,AWS Kinesis,Step Functions,Lake Formation,CloudFormation,IAM,KMS,CloudTrail,OpenSearch,Terraform,Docker,Amazon EKS,AWS CodePipeline,CI,CD,Pipeline Monitoring
Data and Analytics: ETL,ELT Pipelines,Data Integration,Star Schema Modeling,Data Quality,Pipeline Orchestration,Change Data Capture CDC,Tableau,Power BI,Looker
Tools and Methodologies: Databricks,dbt,Apache Spark,Apache Airflow,Amazon MWAA,Data Lake,Git
Skills: PySpark,T-SQL
Streaming & Processing: Apache Kafka,Kafka Connect,Debezium,Apache Flink,Spark Structured Streaming,Amazon Kinesis

Work Experience

Cardinal Health
Austin, TX
Data Engineer
Feb 2025 – Present
Healthcare supply and services company — built data engineering solutions supporting hospital systems, patient records, reporting and compliance.
Tech Stack: Databricks, PySpark, Apache Kafka, Debezium, Amazon Kinesis, AWS Glue, Amazon S3, Amazon Redshift, Apache Airflow, Power BI, Tableau
  • Engineered HIPAA-compliant batch and real-time data pipelines for 12 hospital systems, processing 50M+ patient records monthly using Databricks, PySpark, AWS Glue and S3 with zero compliance violations.
  • Built real-time ingestion and CDC pipelines using Kafka, Debezium and Kinesis, reducing end-to-end data latency from 4 hours to under 5 minutes for operational analytics.
  • Designed a star-schema data warehouse on Amazon Redshift to power 35+ Power BI and Tableau dashboards, enabling teams to deliver reports 15% faster.
  • Orchestrated batch and streaming workflows in Apache Airflow with automated validation and alerting, achieving 99.5% on-time pipeline completion and faster incident detection.
  • Optimized SQL and Redshift performance through query tuning and partitioning, cutting execution time by 25% and reducing annual data warehouse costs by $120,000.
  • Resolved 30+ data-quality and pipeline incidents through root-cause analysis and monitoring improvements, eliminating 95% of recurring issues.
Mastercard
Austin, TX
AWS Data Engineer
Aug 2023 – Dec 2024
Global payments company — developed AWS-based data platforms and pipelines to process payment and settlement data for analytics and reporting.
Tech Stack: Amazon S3, AWS Glue, AWS Lambda, Step Functions, EMR, Spark, PySpark, Apache Kafka, Apache Flink, Amazon Kinesis, Apache Airflow, Amazon Redshift
  • Designed scalable AWS data architectures processing 15M+ payment and settlement records daily using S3, Glue, Lambda and Step Functions to standardize ingestion pipelines.
  • Built a data lake and ETL solutions on S3 and AWS Glue processing 2TB+ of data daily, reducing end-to-end processing time by 35%.
  • Developed batch processing pipelines with Spark and PySpark on EMR, reducing critical job runtimes from 4 hours to 90 minutes through optimization and parallelism.
  • Implemented real-time streaming and CDC pipelines with Kinesis, Kafka and Flink, cutting transaction-data latency from 30 minutes to under 2 minutes for downstream analytics.
  • Orchestrated 100+ production workflows on Apache Airflow / MWAA, achieving 99% on-time completion and reducing data-quality incidents by 40% through automated checks.
  • Optimized Spark jobs and Redshift workloads via tuning, partitioning and resource adjustments, improving performance by 30% and lowering monthly cloud costs by 20%.
PayPal
Austin, TX
Data Engineer
Sep 2021 – Dec 2022
Online payments company — built high-throughput ETL/ELT pipelines and reporting infrastructure for financial risk and regulatory reporting.
Tech Stack: PySpark, Apache Spark, Apache Airflow, SQL, T-SQL, Amazon S3, Amazon Redshift
  • Built high-performance ETL pipelines handling 1B+ daily transactions to support financial risk analytics and regulatory reporting across multiple regulatory bodies.
  • Optimized complex SQL and T-SQL queries, cutting data retrieval latency by 50% and reducing risk report generation time from 4 hours to 2 hours.
  • Implemented ELT transformations processing 200GB+ of data daily to enable advanced analytics across six business units.
  • Maintained 99.9% system uptime across production pipelines serving 5,000+ daily business users through proactive monitoring and runbook improvements.
  • Implemented data validation and monitoring using Airflow sensors and custom PySpark checks, reducing production incidents and false positives.
  • Collaborated with data science and compliance teams to deliver traceable data lineage and schema evolution support, enabling timely regulatory submissions.

Education

Southern Arkansas University
Master of Science in Computer Science (GPA: 3.67) • Magnolia, AR
Aurora's Institute of Science and Technology
Bachelor of Science in Computer Science • India

Certifications

AWS Certified Cloud Practitioner — Amazon Web Services

Powered by Drivetube · Create your own profile at drivetube.ai

Explore Drivetube

  • Drivetube Profile — your free digital resume at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
  • Free Job Board verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
  • Job Hunt Program managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
  • Resume Writing Services human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
  • Community Membership from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
  • Drivetube Hire — for employers hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.

Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.

The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.