Skip to content

Sujeeth Reddy

Data Engineer • Sarnia, Ontario, Canada • s**************@gmail.com • +12******282 • drivetube.ai/•••••

Professional Summary

Data Engineer with 5+ years of experience designing and delivering scalable ETL/ELT pipelines, cloud data platforms, and real-time streaming solutions across healthcare, financial services, and insurance domains. Experienced with AWS and Azure services, Databricks, Spark (Scala/PySpark), Snowflake, and orchestration frameworks to improve data availability, reduce pipeline latency, and enable analytics and ML-ready datasets.

Technical Skills

Programming Languages: Scala,Python,Java,R
Databases: Snowflake,Azure SQL,SQL Server,BigQuery,DynamoDB,Data Lake architectures,Star and Snowflake schemas,SQL
Cloud and DevOps: AWS S3,AWS Redshift,AWS Lambda,AWS Glue,AWS Athena,AWS Kinesis,AWS EMR,AWS EKS,Azure Data Factory,Azure Databricks,Azure Synapse Analytics,Azure Blob Storage,Azure Event Hubs,Azure Functions,Azure Data Lake,Jenkins,Docker,Terraform,CloudFormation,AWS CodePipeline,Azure DevOps,GitHub Actions,Kubernetes
Testing: pytest,ScalaTest,Unit and integration testing,Automated pipeline validation,Pull request review workflows
Data and Analytics: Apache Spark,PySpark,Databricks,DBT,Airflow,Snowpipe,Delta Lake,Unity Catalog,Apache Kafka,Spark Structured Streaming,GCP Dataflow,HDFS,YARN,Stream Analytics,Tableau,Power BI
Tools and Methodologies: Git
Security & Compliance: AWS IAM,Azure Active Directory,VPC configuration,Data encryption,HIPAA,SOC 2,Data masking
Business Intelligence & Reporting: AWS QuickSight,Dashboarding and reporting
Monitoring & Observability: AWS CloudWatch,Azure Monitor,CloudTrail,Log Analytics,Alerting and SLO monitoring

Work Experience

Optum BC
Canada
Data Engineer
Sep 2024 – Present
Worked at Optum (healthcare services) on data platforms supporting financial and operational analytics across clinical and business units.
Tech Stack: AWS Glue, AWS EMR, AWS Lambda, AWS Redshift, AWS Kinesis, Databricks, Delta Lake, Unity Catalog, Scala, PySpark, Terraform, AWS CodePipeline, CloudWatch, pytest, ScalaTest
  • Designed and implemented scalable ETL/ELT pipelines using AWS EMR, Glue, Lambda and Redshift to ingest and process 5TB+ of daily healthcare financial and operational data, enabling consolidated analytics across business units.
  • Developed Scala-based Spark transformations with partition tuning, broadcast joins and caching on Databricks, reducing pipeline runtimes by 30% for high-volume datasets.
  • Built real-time streaming pipelines with AWS Kinesis and Spark Structured Streaming to process 2M+ events/minute for near‑real-time operational insights and patient interaction analytics.
  • Implemented Delta Lake-based ELT on Databricks and Unity Catalog governance to provide ACID compliance, fine-grained access control and standardized datasets for analytics and ML teams.
  • Introduced a Retrieval-Augmented Generation pipeline integrating LLM APIs for intelligent data search and automated insight generation over enterprise datasets, improving analyst query times and discoverability.
  • Established CI/CD and observability: authored pytest and ScalaTest suites, automated deployments with Terraform and AWS CodePipeline, and configured CloudWatch alerts—cutting release cycles by 35% and operational disruptions by 30%.
TCS
India
Data Engineer
Nov 2022 – Aug 2024
Worked at Tata Consultancy Services on healthcare data engineering engagements delivering ETL, streaming, and cloud‑native data platforms for provider and payer analytics.
Tech Stack: Apache Spark, Scala, AWS Glue, AWS S3, Apache Kafka, Spark Structured Streaming, Snowflake, Snowpipe, Airflow, Cloud Composer, AWS CodePipeline
  • Built and maintained batch ETL pipelines using Apache Spark, AWS Glue, Lambda and S3 to process 3TB+ of healthcare data monthly, improving data availability for analytics and reporting.
  • Authored Scala Spark jobs for distributed transformations and optimized partitioning strategies to reduce processing time and resource consumption across large clinical datasets.
  • Implemented low‑latency streaming ingestion with Apache Kafka and Spark Structured Streaming to process 1.5M+ events/day for event-driven healthcare integrations and near-real-time analytics.
  • Designed ELT pipelines integrating Snowflake with AWS S3, Glue and Airflow; leveraged Snowpipe for continuous ingestion and Time Travel for data recovery and testing workflows.
  • Optimized Snowflake performance using clustering keys and micro-partitioning to reduce average query execution times by ~35%, improving analyst productivity for clinical and operational reporting.
  • Automated CI/CD and orchestration using Cloud Composer (Airflow) and AWS CodePipeline, reducing release cycle times by 40% and enforcing repeatable deployment practices.
HCL
India
Data Engineer
Sep 2020 – Oct 2022
Delivered data engineering solutions for financial services clients, building cloud migrations, streaming monitors and analytical data products for trading and transaction analytics.
Tech Stack: Azure Data Factory, Azure Databricks, Azure Synapse Analytics, Azure Event Hubs, Stream Analytics, Databricks, Delta Lake, Azure Monitor, Log Analytics
  • Implemented end-to-end data pipelines using Azure Data Factory and Databricks to process 4TB+ of monthly financial data, enabling consolidated reporting and analytics across trading and operations.
  • Optimized storage and query performance in Azure Synapse Analytics and Snowflake through schema tuning and indexing, reducing query execution times by 40% for business reporting.
  • Built streaming solutions with Azure Event Hubs and Stream Analytics to process 500K+ events/day for transaction monitoring and near-real-time fraud detection workflows.
  • Led migration of legacy on-premise ETL to Azure cloud services, improving scalability and data availability while reducing infrastructure downtime by 40%.
  • Implemented monitoring and alerting with Azure Monitor and Log Analytics, cutting troubleshooting time by 25% and improving SLA adherence for data pipelines.
  • Deployed cost-optimization strategies across Azure services and Databricks clusters to reduce cloud spend by ~20% while maintaining required performance and SLAs.

Education

Jawaharlal Nehru Technological University Hyderabad (JNTUH)
Bachelor of Engineering • Hyderabad, India
Lambton College
Post Graduate Diploma in Cloud Infrastructure and Administration • Sarnia, Ontario, Canada

Powered by Drivetube · Create your own profile at drivetube.ai