Anyutha Reddy
Data Engineer • Charlotte, NC • a**************@gmail.com • 313****087 • drivetube.ai/•••••
Professional Summary
Data Engineer with 5+ years of experience designing and implementing distributed data systems and production-grade ETL/ELT pipelines using cloud-native data platforms and modern orchestration frameworks. Skilled in Python, SQL, PySpark, Snowflake, Databricks and dbt; experienced delivering compliant analytics for finance, healthcare, and insurance domains.
Technical Skills
Programming Language: Python,C#,R,JavaScript,Shell Scripting,SQL
Databases: Snowflake,Microsoft SQL Server,PostgreSQL,Azure Synapse Analytics
Cloud Platforms: S3,Object Storage,Amazon Web Services,Microsoft Azure,Google Cloud Platform
Version Control & Development Tools: Linux,Git,Windows,Unix
DevOps & Infrastructure: GitHub Actions,Jenkins,Terraform,Docker,Kubernetes
Messaging & Monitoring: Apache Kafka,Grafana,Datadog,CloudWatch
Data Engineering & Processing: Data Lake,Apache Spark,PySpark,Spark SQL,Apache Hadoop,HDFS,Apache Hive,Apache Flink,Azure Databricks,Apache Airflow
Data Analysis & Visualization: Tableau,Power BI
AI/ML Frameworks & Libraries: scikit-learn,MLlib
Data Integration & ETL: dbt,Azure Data Factory,Informatica
Work Experience
Bank Of America
Charlotte, NC
Data Engineer
April 2023 – Present
Financial services — designed and maintained data pipelines and analytics datasets used by risk, finance, and compliance teams for loan, credit and account reporting.
Tech Stack: Snowflake, Databricks, Apache Airflow, dbt, Apache Kafka, Terraform, AWS S3, Datadog, AWS CloudWatch, GitHub Actions, Tableau, Python, SQL
- Designed and implemented end-to-end ETL/ELT pipelines in Snowflake and Databricks to process multi-terabyte loan and credit transaction datasets, enabling timely risk reporting and regulatory analytics.
- Built dbt models and standardized transformation layers to produce analytics-ready financial datasets supporting BCBS239 and CCAR compliance frameworks.
- Automated Airflow DAGs to ingest data from internal APIs, AWS S3, and Kafka streams for near-real-time refresh and consistent downstream availability for analytics users.
- Optimized Snowflake performance using materialized views, clustering keys and warehouse sizing to improve query responsiveness for risk and compliance reporting.
- Implemented Infrastructure-as-Code with Terraform to automate Snowflake role provisioning, warehouse scaling and secure environment setup, reducing manual setup time and configuration drift.
- Established PII masking and role-based access control while integrating Datadog and CloudWatch monitoring with alerting to ensure pipeline reliability and audit readiness.
Texas Health Resources
Data Engineer
Feb 2021 – March 2023
Healthcare — built and tuned data ingestion and transformation pipelines for EHR, clinical, and claims systems to support clinical operations, reporting and compliance with HIPAA.
Tech Stack: Azure Databricks, Azure Synapse Analytics, Azure Data Factory, dbt, PySpark, Power BI, Azure Key Vault, Azure Purview, Microsoft SQL Server, Python, SQL
- Designed and maintained ingestion pipelines for EHR, patient encounter and clinical systems using Azure Databricks and Synapse to centralize clinical and operational data.
- Built PySpark and Azure Data Factory pipelines to ingest HL7 and FHIR-formatted healthcare data, delivering timely datasets for operational reporting and clinical analytics.
- Implemented dbt transformation models for claims, patient demographics and billing analytics, enabling consistent schemas for downstream reporting and analytics consumption.
- Collaborated with data governance and security teams to implement HIPAA-compliant PHI masking and integrated Azure Key Vault for secure credential management.
- Implemented data lineage and metadata tracking using Azure Purview to support audit readiness and traceability for clinical and billing datasets.
- Led migration of on-premises SQL Server data warehouse to Azure Synapse preserving schema and history; tuned performance via partitioning, caching and delta optimizations.
CNA Insurance
Chicago, IL
Big Data Engineer
Jan 2020 – Jan 2022
Insurance — developed data pipelines and marts for claims, policy and underwriting analytics used by actuarial and business teams for pricing and risk assessments.
Tech Stack: AWS, Snowflake, dbt, Apache Airflow, Python, SQL, Terraform, S3, AWS Lambda, GitHub Actions, Tableau
- Designed end-to-end data pipelines and curated Snowflake schemas to support claims, policy and customer analytics using AWS-hosted ingestion and Snowflake transformations.
- Ingested raw data from Salesforce, Guidewire and third-party APIs into standardized schemas and built dbt models for underwriting risk assessment and loss analytics.
- Developed reusable Python frameworks for data validation, logging and error handling; integrated Airflow for scheduling, dependency management and monitoring.
- Partnered with actuarial and analytics teams to provide clean, reconciled datasets enabling advanced modeling and pricing simulations.
- Implemented CI/CD pipelines with GitHub Actions and Terraform to automate deployment of dbt models, Airflow DAGs and infrastructure provisioning.
- Tuned Snowflake via clustering, virtual warehouse configuration and query tuning; developed automated reconciliation and backup scripts to support regulatory reporting.
Education
Lindsey Wilson University
Master’s in Computer Science
Powered by Drivetube · Create your own profile at drivetube.ai