Skip to content

Pavan Udata

Azure Data Engineer • p**********@gmail.com • +91*******158 • linkedin.com/••••• • drivetube.ai/•••••

Professional Summary

Azure Data Engineer with 4+ years of experience designing and operating large-scale data pipelines that ingest 100GB+ of enterprise manufacturing data daily. Experienced in Azure Data Factory, Databricks, Delta Lake and Unity Catalog; proven track record improving pipeline throughput (35%), reducing failures (30%), and implementing medallion architectures and governance for analytics-ready data.

Technical Skills

Programming Languages: Python
Databases: SQL,MySQL,Oracle
Cloud and DevOps: Azure Data Factory,Azure Databricks,ADLS Gen2,AWS S3,CI,CD Pipeline Automation
Data and Analytics: Apache Spark,PySpark,Delta Lake,Medallion Architecture,ETL,ELT Pipeline Design,Data Modeling,Slowly Changing Dimensions SCD,Query Optimization,Unity Catalog,Azure Key Vault,Metadata Management,Window Functions,Aggregations,Broadcast Joins
Tools and Methodologies: Git,Notebook Development,Automated Monitoring,Data Validation Frameworks
Performance & Reliability: Partitioning,Z-order Indexing,Caching,Compaction

Work Experience

Infosys
Data Engineer
Nov 2024 – Present
Data engineering within an IT services environment building manufacturing and enterprise analytics pipelines, governance, and lakehouse solutions for downstream reporting and analytics.
Tech Stack: Azure Data Factory, Azure Databricks, PySpark, ADLS Gen2, Delta Lake, Unity Catalog, Git, Azure Key Vault
  • Engineered scalable end-to-end ETL/ELT pipelines using Azure Data Factory and Databricks to ingest and process 100GB+ of manufacturing data daily, enabling centralized analytics across the enterprise.
  • Improved pipeline throughput by 35% through Spark and Delta Lake tuning—implemented broadcast joins, partitioning, caching, Z-order indexing and targeted compaction to reduce job runtimes.
  • Reduced pipeline failures by over 30% by designing automated data validation, retry logic, and monitoring alerts that decreased incident recurrence and sped up mean time to resolution.
  • Implemented Unity Catalog for fine-grained access control, schema enforcement and lineage tracking, strengthening governance and audit readiness across analytics platforms.
  • Delivered a reusable Delta Lake medallion (bronze/silver/gold) architecture and standardized transformation patterns, improving downstream data reliability and developer onboarding time.
  • Automated CI/CD workflows and Git-based notebook/version control to standardize deployments of Databricks jobs and ADF pipelines, reducing manual promotion errors and accelerating releases.
Infosys
Junior Data Engineer
Jan 2023 – Oct 2024
Supported enterprise data migration and ETL development for inventory and manufacturing datasets within Infosys delivery teams, enabling centralized analytics on cloud storage.
Tech Stack: PySpark, Spark, SQL, Python, AWS S3, MySQL, Oracle, Git
  • Built scalable PySpark ETL pipelines to migrate and transform ~100GB/day of inventory and manufacturing data into centralized cloud storage (AWS S3), enabling cross-team analytics.
  • Accelerated pipeline runtimes by ~20% by applying Spark execution plan tuning, partitioning strategies and SQL optimizations that reduced compute costs and job latency.
  • Developed modular, reusable PySpark and SQL transformation components (joins, aggregations, windowing, business logic) to shorten development cycles for new pipelines.
  • Implemented data validation checks and root-cause analysis processes to improve production reliability and reduce time-to-detect for data quality issues.
  • Collaborated with cross-functional stakeholders to translate business rules into transformation logic and SCD handling, ensuring analytics correctness for inventory use cases.
  • Documented ETL designs, schemas and runbooks to support operational handover and consistent maintenance of production pipelines.
Infosys
Systems Engineer Trainee
June 2022 – Dec 2022
Completed entry-level technical training at Infosys focused on data engineering fundamentals and hands-on ETL pipeline projects using Python, SQL, PySpark and Databricks.
Tech Stack: Python, SQL, PySpark, Databricks, Delta Lake, Git
  • Completed intensive training in Python, SQL, PySpark and Databricks and delivered foundational ETL pipeline projects applying transformation and analysis techniques to enterprise-scale datasets.
  • Implemented PySpark notebooks demonstrating joins, aggregations, window functions and basic SCD logic as part of capstone ETL exercises.
  • Designed Bronze/Silver/Gold medallion layer patterns in Delta Lake during labs to illustrate progressive refinement and query performance improvements.
  • Authored SQL queries and transformation scripts for data validation and basic unit testing of ETL outputs to ensure correctness in training projects.
  • Used Git for source control and documented pipeline architecture, transformation logic and run procedures to support knowledge transfer.
  • Presented capstone ETL solutions to trainers and received positive evaluations for correctness, design clarity and readiness for production handover.

Education

Aditya College of Engineering
Bachelor of Technology in Mechanical Engineering • 2017 – 2020

Certifications

Microsoft Certified: Azure Fundamentals (AZ-900) — Microsoft
Infosys Certified PySpark Professional — Infosys
Infosys Certified Python Programmer — Infosys

Powered by Drivetube · Create your own profile at drivetube.ai

Explore Drivetube

  • Drivetube Profile — your free digital resume at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
  • Free Job Board verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
  • Job Hunt Program managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
  • Resume Writing Services human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
  • Community Membership from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
  • Drivetube Hire — for employers hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.

Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.

The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.