Skip to content

JASWANTH REDDY BHIMAVARAPU

Sr. Data Engineer • Dallas, TX • b**************@gmail.com • 314****138 • drivetube.ai/•••••

Professional Summary

Sr. Data Engineer with 5+ years of experience designing, building, and optimizing scalable data pipelines and cloud-native data platforms across Healthcare, Insurance, and Financial Services. Proficient in Python, SQL, PySpark, Spark, Databricks, AWS and Azure services, and building ETL/ELT, streaming, and data quality frameworks to deliver HIPAA- and compliance-aware analytics and ML pipelines.

Technical Skills

Programming Languages: Scala,Java,Shell Scripting,R
Frameworks and Libraries: Scikit-learn,TensorFlow,Keras,Matplotlib,Seaborn
Databases: SQL,Snowflake,SQL Server,PostgreSQL,MySQL,Oracle,Azure SQL Database,DB2,MongoDB,DynamoDB,Cosmos DB,Cassandra
Cloud and DevOps: AWS S3, Redshift, Glue, Athena, Lambda, EMR, Kinesis, Step Functions,Azure Data Factory, Synapse, Data Lake, Databricks, HDInsight,GCP BigQuery, Dataflow, Pub,Sub,Docker,Kubernetes,Terraform,Jenkins,GitHub Actions
Testing: Great Expectations,Apache Griffin,Informatica DQ,AWS CloudWatch,Datadog,Azure Monitor,Monte Carlo
Data and Analytics: Apache Spark,PySpark,Hadoop HDFS, Hive, MapReduce,Kafka,Databricks,Delta Lake,Amazon Redshift,Azure Synapse Analytics,Google BigQuery,Teradata,PySpark MLlib,XGBoost,Tableau,Power BI,Plotly,Excel Dashboards
Tools and Methodologies: Git,JIRA,Confluence,Jupyter Notebook,Agile,Data Governance,HIPAA,SOX
Skills: Python Pandas, NumPy, PySpark
ETL and Orchestration: Apache Airflow,Apache NiFi,dbt,AWS Glue,AWS Step Functions,Azure Data Factory,Informatica,SSIS

Work Experience

Blue Cross Blue Shield Association
Remote, USA
Sr. Data Engineer
January 2025 – Present
Worked on insurance data engineering supporting claims, member enrollment, provider network and analytics for a national health insurance association.
Tech Stack: Python, SQL, Tableau, AWS S3, Redshift, Glue, Athena, Lambda, EMR, Apache NiFi, MongoDB, PySpark, Scikit-learn, Great Expectations, CloudWatch, Docker, Jenkins, Git, HIPAA, Agile
  • Designed and implemented end-to-end ETL/ELT workflows to ingest EDI 837/835 claims, eligibility, and provider feeds using AWS Glue, Apache NiFi and PySpark; produced curated Redshift tables for claims and membership analytics.
  • Built PySpark processing jobs on EMR to handle terabyte-scale historical and daily claims datasets, tuning partitioning and shuffle strategies to accelerate batch runs and enable near-real-time fraud analytics.
  • Architected a cloud-native data platform on AWS (S3, Redshift, Glue, Athena, Lambda) applying HIPAA-compliant storage, encryption, and access controls to support secure querying and cost-optimized retention.
  • Developed feature pipelines and productionized Random Forest fraud detection and risk scoring models using Scikit-learn and PySpark; implemented scheduled scoring jobs and feature stores for operational alerts.
  • Implemented data quality checks and observability using Great Expectations and CloudWatch; created schema, null and threshold validations and SLA alerting to improve pipeline freshness visibility.
  • Automated CI/CD for data artifacts with Git, Jenkins and Docker; collaborated with actuarial, claims operations and compliance teams to deliver HIPAA-compliant analytics and reporting in Agile sprints.
JPMorgan Chase
Atlanta, GA
Data Engineer
January 2024 – January 2025
Built data engineering solutions for retail and commercial banking to support transaction, account, loan, fraud and risk analytics.
Tech Stack: Python, SQL, Power BI, AWS S3, EMR, Lambda, Kinesis, Redshift, Step Functions, Apache Airflow, Snowflake, DynamoDB, Spark, Databricks, Kafka, TensorFlow, Keras, XGBoost, Datadog, Apache Griffin, Terraform, GitHub Actions, Agile
  • Engineered streaming and batch pipelines using Spark on Databricks, Kafka, Kinesis and AWS EMR to process billions of transaction records, enabling sub-minute latency for fraud and risk event detection.
  • Orchestrated complex multi-stage workflows with Apache Airflow and AWS Step Functions across core banking, CRM and regulatory sources; implemented retries, SLA alerts and dependency handling.
  • Designed and tuned Snowflake data models for transaction and customer analytics and implemented DynamoDB low-latency lookup pipelines for session and fraud evaluation use cases.
  • Productionized deep learning pipelines with TensorFlow and Keras for churn prediction and transaction anomaly detection; built feature engineering, training and inference workflows integrated into model serving.
  • Established observability using Datadog and Apache Griffin to monitor pipeline health, throughput and data quality KPIs; developed dashboards and alerts for on-call engineers and stakeholders.
  • Implemented infrastructure-as-code with Terraform and CI/CD using GitHub Actions; partnered with risk, compliance and product teams to deliver regulatory reporting and credit/fraud scoring pipelines.
Tata Consultancy Services (TCS)
Hyderabad, India
Data Engineer
February 2020 – August 2023
Delivered data engineering services for an insurance client, consolidating policy, premium and claims data into Azure cloud to support reporting, BI and ML use cases.
Tech Stack: Python, SQL, Power BI, Excel, Azure Data Factory, Synapse, Data Lake, Databricks, Blob, Monitor, HDInsight, Informatica, Azure SQL Database, Cosmos DB, Hadoop, Hive, PySpark MLlib, Agile
  • Designed and implemented an Azure data platform using Data Factory, Synapse Analytics, Data Lake Storage, Databricks and Blob Storage to ingest policy administration, agent portal and legacy mainframe extracts.
  • Built ETL pipelines with Informatica PowerCenter and Azure Data Factory to transform and load transactional policy and claims data into Azure SQL and Synapse for analytics and reporting.
  • Developed PySpark MLlib pipelines to productionize churn prediction, premium optimization and claim severity models; engineered repeatable feature pipelines and scheduled model scoring jobs.
  • Executed batch processing on HDInsight using Hadoop and Hive for long-term trend analysis; optimized partitioning and query patterns to reduce cost and improve query performance.
  • Implemented data quality and monitoring controls using Informatica DQ and Azure Monitor; defined validation rules, freshness alerts and SLA reporting for critical insurance feeds.
  • Collaborated with onshore stakeholders in Agile/Scrum to gather requirements, present technical designs and deliver Power BI dashboards and data deliverables for business users.

Projects

Business Performance Dashboard (Power BI)
Tools Used: Power BI, SQL, Excel, Data Modeling
  • Built an end-to-end Power BI dashboard aggregating operational and performance KPIs to provide stakeholders with visibility into business trends and efficiency metrics.
  • Performed data cleaning and aggregation using SQL and Excel; implemented measures and calculated columns to support consistent KPI calculations.
  • Presented insights and recommended operational improvements to stakeholders and supported data-driven decision making through scheduled dashboard delivery.

Education

Saint Louis University
Master of Science in Information Systems
GITAM University
Bachelor of Technology in Computer Science

Certifications

Data Structures — University of California, San Diego (Coursera)
Software Development Processes and Methodologies — University of Minnesota (Coursera)
Data Analysis with Python — IBM (Coursera)
AWS S3 Basics — Coursera

Powered by Drivetube · Create your own profile at drivetube.ai