Prajwal Reddy Nakka
Data Engineer • Newark, Delaware • p***************@gmail.com • +13******628 • drivetube.ai/•••••
Professional Summary
Data Engineer with 0 years of experience designing, building, and optimizing ETL/ELT pipelines and cloud data platforms for banking and insurance use cases. Skilled in Python, SQL, PySpark and AWS data services to deliver reliable ingestion, transformation, validation and warehousing. Experience automating CI/CD-driven data deployments, implementing monitoring and security controls, and enabling downstream analytics and ML integrations.
Technical Skills
Programming Languages: Python,Shell,Bash,PowerShell
Frameworks and Libraries: TensorFlow,PyTorch,Scikit-learn,Pandas,NumPy
Databases: SQL,MySQL,PostgreSQL,MongoDB
Cloud and DevOps: Amazon Web Services EC2, S3, Lambda, Glue, Redshift, ECS, EKS, RDS, DynamoDB,AWS SageMaker integration-ready,AWS CodePipeline,AWS CodeBuild,AWS CodeDeploy,Jenkins,GitHub Actions,GitLab CI,CD,Docker,Kubernetes,Amazon EKS,Amazon ECS,Helm,Argo CD,Terraform
Data and Analytics: AWS Glue,PySpark,Apache Airflow,Apache Kafka,AWS DMS,Amazon Redshift,Amazon Athena,Amazon S3,Amazon RDS,Amazon DynamoDB,Amazon Aurora
Tools and Methodologies: Git,GitHub,GitLab,Bitbucket,Jira
Infrastructure as Code: AWS CloudFormation,AWS CDK
Monitoring & Observability: Amazon CloudWatch,AWS X-Ray,Prometheus,Grafana,ELK,OpenSearch,Datadog
Version Control & Collaboration: Confluence,Slack,Microsoft Teams
Work Experience
ETech Solutions (Client: Avis Insurance)
Data Engineer (Cloud & ETL Focus)
April 2023 – May 2024
Contract data engineering for Avis Insurance — built cloud ETL/ELT pipelines and ML-ready data delivery for insurance analytics and enterprise applications.
Tech Stack: Python, PySpark, AWS Glue, Amazon S3, Amazon Redshift, TensorFlow, Scikit-learn, AWS Lambda, AWS CodePipeline, AWS CodeBuild, AWS CodeDeploy, Docker, Amazon ECS, Amazon EKS, CloudWatch, AWS X-Ray, AWS Secrets Manager, AWS Systems Manager Parameter Store
- Designed and developed end-to-end ETL/ELT pipelines using PySpark and AWS Glue to ingest, clean and transform structured and unstructured insurance data into S3 and Redshift for analytics and ML consumption.
- Built data pipelines that exposed curated datasets and model outputs via REST-ready integrations, supporting downstream reporting and ML inference workflows using Python and Scikit-learn/TensorFlow artifacts.
- Implemented automated data profiling and validation routines with Python and PySpark to detect schema drift and data defects, improving data accuracy and trust for analytics consumers.
- Automated CI/CD for data pipeline code and deployments using AWS CodePipeline, CodeBuild and CodeDeploy, enabling repeatable testing and faster production rollouts while reducing manual steps.
- Containerized data processing applications with Docker and deployed scalable workloads to Amazon ECS/EKS to standardize environments and support parallel processing of large insurance datasets.
- Integrated CloudWatch and AWS X-Ray monitoring and alerting for pipeline health and data quality drift, and applied IAM, Secrets Manager and Parameter Store best practices to enforce secure data handling.
PRR Technologies (Client: Chase Bank)
Data Operations Engineer (ETL Support)
January 2022 – July 2022
Supported production ETL operations for Chase Bank — monitored and maintained pipelines, automated workflows and improved pipeline reliability for banking reporting and analytics.
Tech Stack: Python, SQL, AWS EC2, AWS CodeBuild, Amazon CloudWatch, Docker, IAM, Git
- Monitored and maintained production ETL/ELT pipelines running on AWS EC2 and cloud services; used CloudWatch logs and alerts to troubleshoot runtime failures and ensure pipeline availability for banking reports.
- Optimized pipeline performance via SQL and PySpark tuning, data partitioning and parallelism adjustments to reduce job runtime and improve throughput for batch reporting workloads.
- Built CI pipelines using AWS CodeBuild to run unit tests and validation checks for data pipeline code, improving deployment quality and reducing rollback incidents.
- Automated recurring data workflows using Python scripting to eliminate manual processing steps, accelerating data refresh cycles for downstream analytics teams.
- Created dashboards and routine reports from pipeline outputs to support stakeholders and operational decision-making for banking use cases.
- Packaged processing tasks using Docker and enforced IAM-based access control for deployed jobs to standardize runtime environments and secure data access.
Education
Wilmington University
Master of Science in Information Systems Technology • Wilmington, Delaware
Powered by Drivetube · Create your own profile at drivetube.ai