Skip to content

Sai Teja Talluri

Data Engineer • United States • s***************@gmail.com • 940****242 • linkedin.com/••••• • drivetube.ai/•••••

Professional Summary

Data Engineer with 5+ years of experience building production-scale Databricks and AWS pipelines for banking, healthcare, and retail clients. Specialized in PySpark ETL/ELT, Delta Lake incremental ingestion, and cost-optimized Spark workloads. Reduced a 200GB join from 52 to 11 minutes, cut ingestion time 60% and cloud costs 25% while delivering governed, audit-ready datasets used by analytics, ML, and compliance teams.

Technical Skills

Programming Languages: Python,Scala
Databases: SQL,Redshift,Snowflake
Cloud and DevOps: AWS S3,AWS Glue,AWS EMR,AWS Lambda,AWS Kinesis,AWS Athena,AWS Step Functions,GitHub Actions,Jenkins,Terraform,Git
Testing: Great Expectations,PyDeequ,Data validation,Reconciliation
Data and Analytics: Delta Lake,Delta Live Tables,Databricks Workflows,Auto Loader,Unity Catalog,LakeFlow,Liquid Clustering,Star Schema,Snowflake Schema,Dimensional Modeling,Medallion Architecture,Data Vault,ML-ready datasets,SageMaker integration,Apache Airflow,Dbt
Skills: PySpark,Spark SQL
Streaming & CDC: Kafka,Spark Structured Streaming,Change Data Capture
Migration & Modernization: Hadoop to S3,Legacy data warehouse migration,On-prem to cloud modernization
Monitoring, Security & Governance: Datadog,CloudWatch,IAM,KMS,Column-level masking,RBAC
Compliance: SOX,PCI-DSS,HIPAA,SOC 2,GDPR,CCPA

Work Experience

Capital One
McLean, VA
Data Engineer
Aug 2025 – Present
Built Databricks data pipelines for banking transaction and credit datasets; supported analytics, compliance, and reporting for financial services.
Tech Stack: Databricks, Delta Lake, Delta Live Tables, Databricks Workflows, Auto Loader, LakeFlow, Unity Catalog, PySpark, Spark SQL, Kafka, AWS S3, Redshift, Great Expectations, Datadog, SNS, Claude Code, Cursor, Liquid Clustering, GitHub Actions
  • Developed and maintained Databricks pipelines for banking transaction and credit data using Delta Lake, Delta Live Tables and Databricks Workflows; ingested via Auto Loader, LakeFlow and Kafka with CDC-based MERGE, cutting ingestion throughput time by 60% and loading curated tables to Redshift.
  • Led migration of legacy SQL Server and Informatica warehouse to S3 and Delta Lake: catalogued 20+ ETL jobs, translated transformation logic to PySpark, and validated outputs with automated reconciliation prior to cutover, reducing infrastructure overhead 25%.
  • Optimized Spark workloads using Spark UI profiling and query diagnostics; resolved a cartesian-product root cause to reduce a 200GB financial join from 52 to 11 minutes and applied Liquid Clustering to lower cluster spend 25% and query latency 20%.
  • Utilized AI-assisted development using Claude Code and Cursor to accelerate PySpark pipeline scaffolding and tests while enforcing manual validation and code review practices to prevent logic or security defects and preserve release quality.
  • Implemented data quality at each stage using Great Expectations and automated alerts (Datadog, SNS); discovered a nullable FK that would drop 15% of records and reduced mean time-to-detect issues 35%.
  • Established governance via Unity Catalog with RBAC, column-level masking for financial PII, lineage and audit trails; implemented SOC/PCI controls and supported a SOX/PCI-DSS review with zero findings.
Sentara Health
Bengaluru, India
Data Engineer
Jun 2021 – Jul 2023
Built scalable ETL for clinical and claims data (HL7, FHIR) into S3-backed Delta Lake to support payer reporting, analytics, and ML in healthcare.
Tech Stack: AWS Glue, PySpark, S3, Delta Lake, Great Expectations, PyDeequ, Redshift, CloudWatch, Datadog, IAM, SageMaker, Git
  • Designed and implemented scalable ETL pipelines using AWS Glue and PySpark to process HL7, FHIR, JSON and claims data across three payer reporting domains; integrated clinical and relational sources into S3-backed Delta Lake.
  • Implemented data quality and reconciliation frameworks with Great Expectations and PyDeequ to validate schema, completeness and referential integrity, reducing failures 50% and time-to-detection 30%.
  • Prepared ML-ready datasets and feature stores by performing PySpark preprocessing and feature engineering for patient utilization and readmission models; supported downstream model training and deployment workflows.
  • Built Redshift star-schema data marts for healthcare KPIs and patient utilization dashboards; automated monitoring with CloudWatch and Datadog and authored runbooks for production troubleshooting.
  • Enforced HIPAA and SOC 2 controls via IAM policies, encryption and column-level masking to protect PHI and meet regulatory requirements across ETL processes.
  • Accelerated CI/CD releases by 45% through Git-based pipelines and Agile delivery; migrated key ETL jobs to Delta Lake to increase reliability and reduce runtime variability.
Sephora
Bengaluru, India
Data Engineer
Jan 2020 – May 2021
Migrated retail analytics from on-prem Hadoop to cloud Lakehouse; built streaming ingestion and analytics datasets to support personalization and executive reporting.
Tech Stack: Cloudera Hadoop, S3, Delta Lake, PySpark, AWS Glue, EMR, Kafka, Kinesis, Snowflake, Redshift, Lambda, Step Functions, CloudWatch, Datadog
  • Migrated retail analytics workloads from an on-prem Cloudera Hadoop cluster to S3 and Delta Lake; converted legacy Spark batch jobs to PySpark on AWS Glue and EMR, reducing processing costs 20% at millions of daily transactions.
  • Built streaming ingestion pipelines using Kafka and AWS Kinesis for clickstream and transaction events feeding personalization and behavioral analytics.
  • Delivered analytics-ready datasets to Snowflake and Redshift for executive reporting; automated event-driven workflows using Lambda and Step Functions to streamline downstream consumption.
  • Implemented monitoring and alerting via CloudWatch and Datadog, creating dashboards and alerts to track pipeline SLAs and reduce manual incident response.
  • Applied GDPR and CCPA practices including encryption, access controls and column-level masking to protect customer PII while enabling analytics.
  • Collaborated with analytics and product stakeholders to define data requirements and ETL schedules; tuned Glue and EMR jobs to improve throughput and reduce compute time.

Education

University of North Texas
M.S. Computer Science • Denton, TX, USA • Aug 2023 – May 2025
CMRIT
B.E. Computer Science • Bengaluru, India • Aug 2016 – Jul 2020

Certifications

AWS Certified Data Engineer — Amazon Web Services
Databricks Certified Data Engineer Associate — Databricks

Powered by Drivetube · Create your own profile at drivetube.ai