Pallavi Goparaju
Senior Databricks Data Engineer • Redmond, WA • p*******@gmail.com • 475****779 • linkedin.com/••••• • drivetube.ai/•••••
Professional Summary
Senior Databricks Data Engineer with 5+ years of experience designing, building, and optimizing Databricks and PySpark ETL/ELT pipelines in Azure lakehouse environments. Skilled in Delta Lake, ADLS Gen2, ADF orchestration, dbt refactoring, performance tuning, monitoring and CI/CD to improve pipeline reliability and reporting SLAs.
Technical Skills
Programming Languages: Python
Databases: SQL,Snowflake
Cloud and DevOps: Microsoft Azure,AWS,Azure DevOps,Azure Monitor,Kusto Query Language
Testing: Data validation,Record-count checks,Lineage,Azure RBAC,Azure Key Vault
Data and Analytics: Azure Databricks,Spark SQL,Databricks Jobs,Cluster tuning,Notebook workflows,dbt,Azure Synapse Analytics,Azure SQL Database,Power BI
Tools and Methodologies: Git
Skills: PySpark
Lakehouse & Storage: Delta Lake,ADLS Gen2,Azure Blob Storage,Partitioning,Schema evolution
Orchestration & ETL: Azure Data Factory,Airflow,AWS Glue,Fivetran
BI & Reporting: DAX,Query optimization
Work Experience
Cigna
Connecticut, CT
Data Engineer - Operations Analytics
May 2025 – Present
Worked at a healthcare insurer supporting enterprise analytics and operations reporting by building and maintaining Azure lakehouse ETL/ELT pipelines.
Tech Stack: Azure Data Factory, Azure Databricks, PySpark, dbt, Azure Monitor, KQL, Azure DevOps, Python, SQL
- Built and supported production ETL/ELT pipelines using Azure Data Factory, Azure Databricks, PySpark, dbt, Python and SQL to ingest and transform enterprise analytics data from multiple source systems for operations analytics.
- Optimized high-volume PySpark transformations and Databricks cluster configurations to improve nightly processing scalability and reduce downstream reporting bottlenecks.
- Refactored dbt models and SQL logic for critical reporting workloads, improving query performance by 30% while preserving business metric definitions across analytics layers.
- Implemented end-to-end validation controls including freshness checks, record-count gates and structured error handling across ADF and Databricks jobs to detect feed issues earlier in the pipeline lifecycle.
- Used Databricks job logs, Azure Monitor, KQL and SQL reconciliation checks to investigate pipeline failures, isolate anomalous source records and reduce repeated production support escalations.
- Established CI/CD patterns with Git and Azure DevOps for Databricks notebooks and ADF artifacts to improve release consistency and reduce deployment-related incidents.
Cognizant
Redmond, WA
Senior Data Engineer
Feb 2024 – May 2025
Worked at an IT services provider building data pipelines to move healthcare member and claims data into ADLS Gen2 and curated lakehouse layers for enterprise reporting.
Tech Stack: Azure Databricks, Azure Data Factory, ADLS Gen2, Delta Lake, PySpark, Azure SQL, Power BI, Azure Monitor, KQL
- Built and maintained Databricks and ADF pipelines to ingest healthcare member and claims data into ADLS Gen2 raw, staged and curated zones for enterprise reporting consumers.
- Profiled slower nightly workloads and migrated complex transformations into PySpark and Delta Lake, improving production reliability and reducing early-morning SLA overruns.
- Defined ADLS folder standards, partitioning patterns and curated storage layer conventions to enable predictable automated consumption by Power BI and SQL-based reporting.
- Troubleshot intermittent ADF and Databricks failures using Azure Monitor, KQL, Databricks job logs and SQL reconciliation checks; identified malformed payloads and integration runtime issues.
- Optimized SQL used by downstream reports and refactored heavier logic into Azure SQL stored procedures to improve Power BI refresh predictability during business hours.
- Partnered with BI and compliance stakeholders to define ADLS/Synapse RBAC, encryption defaults and catalog lineage practices for secure access and traceability.
University of Bridgeport
Bridgeport, CT
Graduate Teaching Assistant - Data Engineering Projects
Aug 2022 – Dec 2023
Supported university data engineering projects consolidating clinical, operational and research datasets for healthcare analytics and research reporting.
Tech Stack: Snowflake, BigQuery, dbt, Fivetran, Airflow, Python, SQL
- Consolidated clinical, operational and research datasets into Snowflake and BigQuery to support healthcare analytics, cohort analysis and longitudinal research reporting.
- Built repeatable ingestion workflows using Fivetran, Airflow and Python to handle EHR schema drift, delayed feeds and source reliability issues before data reached research dashboards.
- Developed dbt models, macros and window-function-heavy SQL with time-aware encounter logic to centralize reusable healthcare metric definitions for researchers.
- Added freshness checks, record-count validations and profiling controls; tuned Snowflake queries and dbt lineage views to improve source-to-report traceability.
- Documented lineage and data quality rules for research consumers and automated profiling to make cohort extraction reproducible and auditable.
- Mentored student teams on pipeline design, code reviews and best practices for data validation and reproducible analytics workflows.
Capgemini
Hyderabad, India
Senior Software Engineer
May 2021 – Aug 2022
Worked at a technology consulting firm on Azure-based data engineering solutions, implementing lakehouse practices and secure data access for analytics teams.
Tech Stack: ADLS Gen2, Azure Data Factory, Delta Lake, Azure Databricks, Python, KQL, Azure SQL
- Managed ADLS data zones by defining folder structures, partition standards, retention and curated layer conventions to support downstream analytics and reporting teams.
- Optimized ADF pipelines ingesting Azure SQL data by adjusting batch sizes, trigger timing and transformation placement between ADF, Azure SQL and Databricks.
- Implemented Delta Lake tables on ADLS with schema evolution and merge patterns to enable transactional updates while preserving existing Power BI reporting compatibility.
- Investigated data quality issues by comparing raw ADLS files with curated Delta layers and writing SQL/KQL checks to isolate problematic pipeline runs.
- Supported cloud data security activities including ADLS RBAC, storage key rotation and encryption practices and participated in row-level security design discussions for shared datasets.
- Automated routine maintenance and ingestion validations with Python scripts and Databricks notebooks to reduce manual remediation and improve pipeline uptime.
Capgemini
Hyderabad, India
Software Engineer
May 2020 – May 2021
Delivered AWS-based ETL solutions at a consulting firm, building Glue and PySpark pipelines to process and curate source files into S3 for analytics.
Tech Stack: AWS Glue, PySpark, S3, AWS SDK, Step Functions, Lambda, CloudWatch, Python
- Developed AWS Glue and PySpark ETL jobs to process CSV source files into partitioned S3 datasets for downstream querying and reporting.
- Built Python utilities leveraging the AWS SDK to validate data, remediate exceptions and automate recurring checks across S3 and Glue workflows.
- Configured Glue Catalog metadata and tuned crawlers to handle schema inference edge cases across changing source files and formats.
- Integrated nightly workflows with AWS Step Functions and Lambda for orchestration and monitored data quality issues through CloudWatch logs.
- Implemented partitioning and compression strategies to reduce query latency and manage S3 storage costs for analytical workloads.
- Collaborated with cross-functional teams to onboard new data sources, define ingestion SLAs and agree format contracts with consumers.
Operations Analytics
Software
Education
University of Bridgeport
Master of Science in Computer Science • Bridgeport, CT • Aug 2022 – Dec 2023
CMR College of Engineering & Technology
Bachelor of Technology in Electrical & Electronics Engineering • India • Aug 2016 – Apr 2020
Certifications
Databricks Certified Data Engineer Associate — Databricks
Databricks Machine Learning Certification — Databricks
Databricks Data Analysis Certification — Databricks
Python Certification
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.