Skip to content

Naveen Nuukarapu

Data Engineer • Hyderabad, India • n***************@gmail.com • +91*******989 • drivetube.ai/•••••

Professional Summary

Data Engineer with 0 years of experience designing and operating ETL pipelines, cloud data platforms, and analytics solutions. Skilled in Python, SQL, Spark and AWS; delivered automated pipelines and interactive dashboards that improved reporting efficiency and data quality for financial, healthcare, and retail clients.

Technical Skills

Programming Language: Python,Shell Scripting,Scala,R,SQL
Databases: Amazon Redshift,Apache HBase,Materialized Views
Cloud Platforms: S3,Amazon Web Services,Lambda,EC2,AWS Glue,Azure Blob Storage,Google Cloud Platform
Messaging & Monitoring: Apache Kafka
Data Engineering & Processing: Amazon EMR,Apache Hadoop,HDFS,MapReduce,Apache Spark,Spark SQL,Apache Hive,Apache Oozie,Apache Airflow,Sqoop,Apache Pig
Data Warehousing: Dimensional Modeling
Data Analysis & Visualization: Power BI
Data Integration & ETL: Azure Data Factory,Apache NiFi,Informatica,Talend

Work Experience

Hopkins Software Pvt Ltd
Hyderabad, India
Data Engineer
Jul 2024 – Present
Worked on data engineering supporting financial, healthcare, and retail clients; built ETL pipelines, cloud migrations, and analytics solutions to improve reporting and data quality.
Tech Stack: Python, Pandas, SQL, Apache Spark, AWS EMR, Amazon Redshift, Amazon S3, AWS Lambda, Power BI, Apache Kafka, Apache Hive
  • Designed and maintained ETL pipelines processing 5M+ records across financial, healthcare, and retail datasets using Python, Apache Spark, and AWS EMR to increase data availability for analytics consumers.
  • Analyzed complex datasets with Python, SQL, and Pandas to extract actionable insights that supported revenue optimization efforts, contributing to a 15% improvement in identified revenue opportunities.
  • Built and delivered Power BI dashboards and automated reporting workflows, reducing manual reporting effort by 60% and accelerating executive decision cycles.
  • Developed ingestion and parsing solutions for unstructured sources (PDFs, external APIs) using Python and Spark, improving data accuracy by 25% and decreasing manual cleansing needs.
  • Led migration from on-premise Microsoft SQL Server to Amazon Redshift, performing schema conversion, ETL validation, and performance tuning to ensure continuity of analytics and reporting.
  • Constructed dimensional data models and materialized views and optimized SQL queries and Python scripts to accelerate report generation and improve query performance and processing throughput.

Projects

Finance Data ETL Pipeline Automation
Tools Used: Python, AWS Lambda, Amazon S3, Amazon Redshift
  • Orchestrated serverless data ingestion and validation workflows using AWS Lambda to handle financial datasets and ensure reliable daily delivery.
  • Automated manual Excel-based processes into repeatable ETL pipelines, reducing reporting SLA from 24 hours to 30 minutes for critical finance reports.
  • Leveraged Amazon S3 for reliable cloud storage and staging to improve pipeline reliability and data availability for downstream analytics.
  • Eliminated manual intervention through end-to-end automation and monitoring, increasing on-time delivery and reducing human error.
  • Administered and validated the daily ETL pipeline to ensure consistent and accurate delivery of financial information to reporting consumers.

Education

Sangai International University
Bachelor of Technology (B.Tech) • Manipur, India • Jul 2015 – May 2019

Powered by Drivetube · Create your own profile at drivetube.ai