Saikiran Anugula
Professional Summary
Data Engineer with 4 years of experience building scalable ETL/ELT platforms and production data pipelines for healthcare and analytics use-cases. Strong hands-on experience with PySpark and Databricks for large-scale batch processing, Delta Lake for governed lakehouse architecture, and Python and SQL for transformation and data quality. Experienced in orchestrating reliable workflows with Apache Airflow and migrating pipelines to AWS (S3, Glue, EC2, Redshift). Delivered production platforms processing healthcare claims, EHR and pharmacy data, reduced end-to-end delivery times and query runtimes, and implemented automated data validation and monitoring for downstream ML and BI consumers. Seeking a Data Engineer role where I can apply robust data modeling, pipeline orchestration, and monitoring to improve data reliability and scale analytics.
Technical Skills
Work Experience
- Designed and maintained scalable ETL pipelines using PySpark and Delta Lake to process 1.5 GB/day of healthcare claims, EHR, and pharmacy data, reducing end-to-end delivery time from 7 hours to 2 hours.
- Implemented Bronze/Silver/Gold data architecture on Databricks and Delta Lake, applying partitioning and data skipping strategies that reduced analytical query runtime by 35%.
- Established automated data quality and validation rules using Python and SQL to enforce schema conformance and business checks across 15 production tables, improving downstream accuracy and consistency.
- Orchestrated production workflows with Apache Airflow by implementing SLA monitoring, retry logic, and alerting to accelerate failure detection and simplify incident response.
- Automated ingestion monitoring and diagnostic logging with Python and Databricks jobs to surface job-level failures and reduce manual troubleshooting overhead.
- Partnered with data scientists and analysts to design feature marts using SQL and Delta Lake to support model training and repeatable analytics.
- Built Python and SQL ETL pipelines integrating two internal APIs to automate extraction, transformation, and loading, reducing manual processing steps from 4 to 1 per ingestion run.
- Cleaned, profiled, and standardized multi-source datasets using Pandas and SQL to produce reliable, training-ready tables for a recommendation model.
- Implemented automated validation checkpoints using Python and SQL, including row-count checks and schema conformance to reduce recurring ingestion failures.
- Developed logging and error-handling in pipelines to capture schema drift and invalid records for faster remediation and clearer failure diagnostics.
- Documented data contracts, pipeline interfaces, and ingestion runbooks to streamline handoffs between engineering and analytics stakeholders.
- Implemented unit-level data tests and end-to-end smoke checks to validate pipeline integrity prior to deployment.
- Designed and optimized star-schema data warehouses in SQL Server using complex SQL and window functions, cutting report query time from 60 seconds to 35 seconds.
- Developed data validation and reconciliation routines using SQL and Python across production pipelines, reducing data reliability incidents from 15 to 4 per quarter.
- Migrated legacy ETL workloads to AWS S3 and AWS Glue, redesigning storage and compute patterns for analytics and reducing storage costs by 10%.
- Automated ingestion and transformation of JSON, XML, and text data using Bash and Python to onboard new data sources without manual preprocessing.
- Designed incremental load strategies and Redshift distribution/partitioning patterns to improve ETL performance for large fact tables.
- Implemented SQL-based reconciliation jobs and scheduled monitoring to detect pipeline regressions and accelerate incident resolution.
Projects
- Built an end-to-end analytics platform using Python and Pandas to ingest, profile, cleanse, and transform operational motel datasets into reporting-ready tables.
- Implemented data quality checks, deduplication, and schema standardization with SQL and Pandas to convert raw records into consistent datasets for analysis.
- Developed interactive Streamlit dashboards with Plotly visualizations to communicate KPIs and operational trends to stakeholders.
- Built a real-time ingestion and processing pipeline using Apache Kafka and PySpark to support scalable streaming analytics.
- Implemented Spark SQL transformations, checkpointing, and schema validation to enforce data quality and improve processing reliability.
- Configured Docker-based reproducible development environments and used Git for version control to maintain engineering workflows.
- Designed PostgreSQL schemas for downstream storage and low-latency analytical queries.
Education
Powered by Drivetube · Create your own profile at drivetube.ai
Everything on Drivetube
Six products, two of them free forever. Start wherever you are.
Your free digital resume
at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept.
Documentation- Free Job BoardFree
Verified jobs, posted in the last 3 days
verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed.
Documentation - Job Hunt ProgramFrom $199.99
We run your job hunt for you
managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you.
Documentation - Resume Writing ServicesFrom $25.99
Written by senior career writers
human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile.
Documentation - Community MembershipFrom $49.99/yr
AI tools, gated filters and a $10,000+ library
from $49.99/year (₹1,999/year in India), or $4.99/month in the iOS app. Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates.
Documentation - Drivetube HireFree tier
For employers — no job postings, no applications
hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do.
Documentation
- Documentation
every product explained in full, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
- Playbooks
tactical job-search guides. Each post is one named tactic with the situation it applies to, the exact steps, a copy-paste message and honest failure modes. Free, no account.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.
Pricing is identical on the web, iOS and Android — no app-store markup. Monthly Community membership is purchased in the iOS app only; the website sells annual plans. The Android app has no in-app purchase, so Android members subscribe on drivetube.ai and then sign in to the app with full access. Membership follows your account, not your device.