Skip to content

Saikiran Anugula

Data Engineer • Durham, United States • a************@gmail.com • 419****044 • linkedin.com/••••• • drivetube.ai/•••••

Professional Summary

Data Engineer with 4 years of experience building scalable ETL/ELT platforms and production data pipelines for healthcare and analytics use-cases. Strong hands-on experience with PySpark and Databricks for large-scale batch processing, Delta Lake for governed lakehouse architecture, and Python and SQL for transformation and data quality. Experienced in orchestrating reliable workflows with Apache Airflow and migrating pipelines to AWS (S3, Glue, EC2, Redshift). Delivered production platforms processing healthcare claims, EHR and pharmacy data, reduced end-to-end delivery times and query runtimes, and implemented automated data validation and monitoring for downstream ML and BI consumers. Seeking a Data Engineer role where I can apply robust data modeling, pipeline orchestration, and monitoring to improve data reliability and scale analytics.

Technical Skills

Programming Language: Python,Bash,SQL
Databases: PostgreSQL,DuckDB
Cloud Platforms: Amazon Web Services,S3,AWS Glue
Version Control & Development Tools: Git
DevOps & Infrastructure: Docker
Messaging & Monitoring: Apache Kafka
Data Engineering & Processing: Delta Lake,PySpark,Spark SQL,Apache Airflow
Data Warehousing: Data Warehousing,Data Modeling
Data Analysis & Visualization: Pandas,NumPy,Streamlit,Plotly
Data Integration & ETL: ETL,ELT,Data Validation
Compliance & Governance: Data Quality

Work Experience

IQVIA
Durham, United States
Data Engineer (Contract)
Jan. 2026 – Present
Worked on healthcare data engineering: building production ETL/ELT pipelines and lakehouse architectures to support analytics and modeling for claims, EHR, and pharmacy datasets.
Tech Stack: PySpark, Delta Lake, Databricks, Python, SQL, Apache Airflow
  • Designed and maintained scalable ETL pipelines using PySpark and Delta Lake to process 1.5 GB/day of healthcare claims, EHR, and pharmacy data, reducing end-to-end delivery time from 7 hours to 2 hours.
  • Implemented Bronze/Silver/Gold data architecture on Databricks and Delta Lake, applying partitioning and data skipping strategies that reduced analytical query runtime by 35%.
  • Established automated data quality and validation rules using Python and SQL to enforce schema conformance and business checks across 15 production tables, improving downstream accuracy and consistency.
  • Orchestrated production workflows with Apache Airflow by implementing SLA monitoring, retry logic, and alerting to accelerate failure detection and simplify incident response.
  • Automated ingestion monitoring and diagnostic logging with Python and Databricks jobs to surface job-level failures and reduce manual troubleshooting overhead.
  • Partnered with data scientists and analysts to design feature marts using SQL and Delta Lake to support model training and repeatable analytics.
Cognit AI Inc.
Elizabeth City, United States
Intern
Jan. 2025 – Apr. 2025
Built ETL pipelines and data validation for an AI/recommendation workflows to prepare training datasets and increase ingestion reliability.
Tech Stack: Python, SQL, Pandas
  • Built Python and SQL ETL pipelines integrating two internal APIs to automate extraction, transformation, and loading, reducing manual processing steps from 4 to 1 per ingestion run.
  • Cleaned, profiled, and standardized multi-source datasets using Pandas and SQL to produce reliable, training-ready tables for a recommendation model.
  • Implemented automated validation checkpoints using Python and SQL, including row-count checks and schema conformance to reduce recurring ingestion failures.
  • Developed logging and error-handling in pipelines to capture schema drift and invalid records for faster remediation and clearer failure diagnostics.
  • Documented data contracts, pipeline interfaces, and ingestion runbooks to streamline handoffs between engineering and analytics stakeholders.
  • Implemented unit-level data tests and end-to-end smoke checks to validate pipeline integrity prior to deployment.
Wipro
India
Software Engineer
Feb. 2022 – Jun. 2024
Delivered enterprise data warehouse and cloud migration workstreams: built star-schema warehouses, validation routines, and migrated ETL to AWS to support analytics and churn modeling.
Tech Stack: SQL Server, AWS S3, AWS Glue, Bash, Python
  • Designed and optimized star-schema data warehouses in SQL Server using complex SQL and window functions, cutting report query time from 60 seconds to 35 seconds.
  • Developed data validation and reconciliation routines using SQL and Python across production pipelines, reducing data reliability incidents from 15 to 4 per quarter.
  • Migrated legacy ETL workloads to AWS S3 and AWS Glue, redesigning storage and compute patterns for analytics and reducing storage costs by 10%.
  • Automated ingestion and transformation of JSON, XML, and text data using Bash and Python to onboard new data sources without manual preprocessing.
  • Designed incremental load strategies and Redshift distribution/partitioning patterns to improve ETL performance for large fact tables.
  • Implemented SQL-based reconciliation jobs and scheduled monitoring to detect pipeline regressions and accelerate incident resolution.

Projects

MoteIQ – Motel Analytics Platform | 2025 – 2026
Tools Used: Python, Pandas, DuckDB, SQL, Streamlit, Plotly, NumPy
  • Built an end-to-end analytics platform using Python and Pandas to ingest, profile, cleanse, and transform operational motel datasets into reporting-ready tables.
  • Implemented data quality checks, deduplication, and schema standardization with SQL and Pandas to convert raw records into consistent datasets for analysis.
  • Developed interactive Streamlit dashboards with Plotly visualizations to communicate KPIs and operational trends to stakeholders.
Real-Time Data Processing Pipeline | 2025 – 2025
Tools Used: Apache Kafka, PySpark, Spark SQL, Docker, PostgreSQL, Git
  • Built a real-time ingestion and processing pipeline using Apache Kafka and PySpark to support scalable streaming analytics.
  • Implemented Spark SQL transformations, checkpointing, and schema validation to enforce data quality and improve processing reliability.
  • Configured Docker-based reproducible development environments and used Git for version control to maintain engineering workflows.
  • Designed PostgreSQL schemas for downstream storage and low-latency analytical queries.

Education

Bowling Green State University
Master of Science in Computer Science • Bowling Green, OH • Aug. 2024 – May 2026

Powered by Drivetube · Create your own profile at drivetube.ai

Everything on Drivetube

Six products, two of them free forever. Start wherever you are.

Read the docs →
  • Your free digital resume

    at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept.

    Documentation
  • Verified jobs, posted in the last 3 days

    verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed.

    Documentation
  • Job Hunt ProgramFrom $199.99

    We run your job hunt for you

    managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you.

    Documentation
  • Written by senior career writers

    human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile.

    Documentation
  • Community MembershipFrom $49.99/yr

    AI tools, gated filters and a $10,000+ library

    from $49.99/year (₹1,999/year in India), or $4.99/month in the iOS app. Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates.

    Documentation
  • For employers — no job postings, no applications

    hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do.

    Documentation
  • Documentation

    every product explained in full, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.

  • Playbooks

    tactical job-search guides. Each post is one named tactic with the situation it applies to, the exact steps, a copy-paste message and honest failure modes. Free, no account.

The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.

Pricing is identical on the web, iOS and Android — no app-store markup. Monthly Community membership is purchased in the iOS app only; the website sells annual plans. The Android app has no in-app purchase, so Android members subscribe on drivetube.ai and then sign in to the app with full access. Membership follows your account, not your device.