Nagendra Reddy Koppula
Professional Summary
Data Engineer with 5+ years of experience building resilient batch and streaming data platforms. Expert in cloud-native ETL/ELT, Spark-based processing, and warehouse optimization to accelerate analytics and production ML workflows. Proven track record implementing CI/CD and infrastructure-as-code to improve reliability, reduce costs, and scale multi-terabyte pipelines.
Technical Skills
Work Experience
- Designed end-to-end ingestion and processing pipelines using Azure Data Factory and Databricks (PySpark) to move telemetry and business data from ADLS Gen2 and Event Hubs into Azure Synapse, enabling analytics and ML feature stores.
- Implemented scalable ETL/ELT frameworks in Python and SQL with incremental loads, schema evolution, and automated data quality checks, improving batch processing performance by 35% across 20+ multi-terabyte tables.
- Optimized Synapse data warehouse using distribution strategies, partitioning, and query tuning to reduce median analytical query times by 40%, resulting in $8K monthly infrastructure cost savings.
- Partnered with data science to productionize ML data pipelines using Azure Machine Learning and MLflow; built real-time feature engineering paths with Azure Stream Analytics and integrated model outputs into Power BI dashboards.
- Introduced CI/CD and IaC using Azure DevOps, Terraform, and GitHub Actions to automate deployments and environment provisioning, cutting deployment time to 15 minutes and improving pipeline reliability to 99.5%.
- Mentored and onboarded 3 junior data engineers on ETL design patterns, dimensional modeling, and performance tuning, increasing team delivery velocity and reducing critical incident recovery time.
- Architected and optimized ETL pipelines with Azure Data Factory and Databricks to process ~50,000 daily transactions from ADLS Gen2, reducing end-to-end processing time by 40% through PySpark automation.
- Built a customer insights platform leveraging Event Hubs and Databricks for near real-time processing using Python and PySpark; integrated outputs into Power BI to deliver a 25% lift in key engagement metrics.
- Led implementation of data governance with Purview and Synapse, establishing metadata, lineage, and automated validation in Data Factory and Databricks to ensure regulatory compliance and data accuracy.
- Engineered ML-ready pipelines to support risk analytics using PySpark and MLflow; containerized scoring with Azure Container Instances and orchestrated runs via Data Factory, cutting credit risk data processing time by 60%.
- Orchestrated multi-source integration using Azure Cosmos DB and Azure SQL Database and enforced CI/CD using Azure DevOps and Git to automate testing and deployment, reducing data inconsistencies by 35% across systems.
- Performed query tuning and table design in Synapse and Databricks to accelerate dashboard refreshes and reduce report latency for business stakeholders.
- Led on-premises to AWS data warehouse migration, consolidating 500GB+ of insurance data into Amazon S3 and Redshift with star schema designs, cutting reporting query times by 50%.
- Built serverless ETL pipelines using AWS Glue (PySpark) and Lambda to automate claims ingestion and validation, reducing average processing time from 4 hours to 15 minutes.
- Designed dimensional models and optimized Redshift fact/dimension tables to support fraud detection analytics, improving fraud identification response by 20% through faster queries.
- Implemented event-driven ingestion using S3 event notifications, Glue jobs, and Lambda triggers to reduce data latency from 6 hours to 1 hour and eliminate manual uploads.
- Enforced data security and compliance for insurance workloads using S3 encryption (AES-256), IAM role-based access, and VPC isolation to meet HIPAA and SOC2 requirements.
- Established CI/CD for ETL using AWS CodePipeline and Git to automate testing and deployment of Glue jobs, reducing deployment cycles by 70% and minimizing manual errors.
Projects
- Prepared and standardized an emotion-labeled text dataset by cleaning raw text, normalizing labels, and validating inputs to ensure consistent model training.
- Built a reusable Python preprocessing pipeline to produce feature-consistent inputs for both training and inference, enabling reproducible ML workflows.
- Trained a lightweight DistilBERT classification model to validate dataset quality and preprocessing, demonstrating multi-class emotion prediction capability.
- Developed a standalone prediction script that reused training preprocessing logic for real-time inference on new text inputs.
- Designed a batch pipeline to ingest Twitter API streams and external datasets, handling schema differences and normalizing records for analytics.
- Orchestrated ingestion and transformation workflows using Apache Airflow on an AWS EC2 host to provide scheduled, reliable data movement.
- Implemented Python ETL logic to clean and enrich tweets, storing processed datasets in organized S3 paths to support downstream analysis.
- Explored industrial time-series data to assess quality and performed cleaning to address missing values and outliers common to sensor datasets.
- Executed correlation analysis and feature selection to prepare inputs for regression and tree-based models, demonstrating the impact of preprocessing on forecast accuracy.
- Evaluated baseline ML models to benchmark prediction performance and identified preprocessing strategies that improved model stability.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.