Sunil Kumar Beesu
Junior Data Engineer • Andhra Pradesh, India • s*******************@gmail.com • +91*******330 • linkedin.com/••••• • drivetube.ai/•••••
Career Objective
Junior Data Engineer with 0 years of experience building ETL pipelines and processing large datasets using Python, SQL, Apache Spark and Hadoop. Familiar with AWS (S3, EC2, EMR), Hive, and relational databases; experienced in data cleansing, transformations, and Spark optimization. Seeking an entry-level data engineering role to design scalable pipelines and support analytics-driven decision making.
Education
Sree Vidyanikethan Engineering College, Tirupati
Bachelor of Technology (B. Tech) in Electrical and Electronics Engineering • Tirupati, India • 2019 – 2023
Sri Chaitanya Junior College, Vijayawada
Intermediate (12th) • Vijayawada, India • 2017 – 2019
D.A.V High School, R.T.P.P
Secondary Education (10th SSC) • India • 2010 – 2017
Technical Skills
Programming Languages: Python,Java
Databases: SQL,MySQL,Oracle SQL
Cloud and DevOps: AWS S3,AWS EC2,AWS EMR,AWS Step Functions,Azure Storage,Azure Data Factory,Google Cloud Storage,BigQuery,Linux
Data and Analytics: Apache Spark,PySpark,Spark SQL,Hadoop HDFS,Hive,Sqoop,ETL,Data Pipelines,Data Cleaning,Data Transformation,Data Validation,Data Warehousing,Data Modeling
Tools and Methodologies: Git,GitHub,IntelliJ IDEA,PuTTY,MobaXterm,Spark partitioning and optimization
Projects
Big Data Analytics Using Apache Spark
Tools Used: Apache Spark, PySpark, Spark SQL, Hadoop HDFS, Hive, AWS S3, AWS EMR, Git, Python
- Built a scalable Spark-based data processing pipeline to ingest and process large structured datasets stored on Hadoop HDFS and AWS S3 for analytical reporting.
- Implemented data extraction, cleansing, transformation, and aggregation using PySpark DataFrames and Spark SQL to prepare datasets for downstream analytics.
- Optimized job performance by applying partitioning strategies, selective column pruning, and caching which reduced iterative processing time in development.
- Integrated Hive tables for structured query access and used Spark SQL to perform complex joins and aggregations for business reports.
- Deployed processing workloads on AWS EMR to enable distributed execution and leveraged S3 for durable storage of raw and processed data.
ETL Pipeline Using Python and SQL
Tools Used: Python, Pandas, SQL, MySQL
- Designed and implemented an end-to-end ETL pipeline to extract CSV data, perform cleansing and transformations using Python and Pandas, and load results into MySQL.
- Automated validation and preprocessing steps including null handling, type conversion, and normalization to improve data quality before ingestion.
- Wrote parameterized SQL scripts to load transformed data into MySQL tables and implemented primary key and indexing strategies for query performance.
- Packaged reusable Python modules and scripts to standardize data loading and reporting tasks across multiple datasets.
- Established logging and error-handling in ETL processes to enable easier debugging and operational stability during runs.
Certifications
HackerRank - SQL — HackerRank
HackerRank - Python — HackerRank
Project Management — NPTEL
Salesforce Developer Virtual Internship — Smart Internz
Powered by Drivetube · Create your own profile at drivetube.ai