Skip to content

Rupa Jhade

Data Engineer • Pittsburgh, PA • s***********@gmail.com • 224****003 • linkedin.com/••••• • drivetube.ai/•••••

Professional Summary

Data Engineer with 4 years of experience building and operating production data pipelines in banking and enterprise environments. Strong in PySpark, Apache Airflow, and AWS, focused on data reliability, performance tuning, and compliance-aware ETL design. Experienced owning end-to-end pipelines from ingestion to validated analytics-ready datasets and collaborating with analysts and compliance stakeholders.

Technical Skills

Programming Languages: Python,Bash,R,JavaScript,C++
Frameworks and Libraries: Matplotlib,Seaborn,Streamlit,Flask,Jupyter,MLflow
Databases: SQL,PostgreSQL,Amazon Redshift,Snowflake,NoSQL
Cloud and DevOps: Apache Airflow,Docker,CI,CD,Terraform,AWS S3,AWS Glue,AWS Lambda,AWS EMR,AWS Redshift,AWS RDS,AWS EC2,AWS IAM,AWS CloudWatch,Azure Databricks,Azure Data Factory,GCP
Data and Analytics: ETL,ELT Development,Data Modeling,Schema Design,Pipeline Automation,Data Integration,Data Warehousing,Data Validation,Data Quality,Data Governance,dbt,PySpark,Spark Scala,Hive,Hadoop,Tableau,Power BI,Excel
Tools and Methodologies: Git
Messaging & Stream Processing: Kafka,Kinesis,Flink

Work Experience

EXL Services
Pittsburgh, PA
Data Engineer
May 2024 – Present
Worked at EXL Services supporting retail banking analytics and regulatory reporting; built and operated ETL pipelines delivering analytics-ready datasets for BI, analysts, and compliance teams.
Tech Stack: PySpark, Apache Airflow, AWS S3, RDS, Parquet, Docker, Git
  • Built and maintained daily batch PySpark pipelines processing retail banking transactions and customer data from 4 to 10 source systems, delivering analytics-ready datasets consumed by BI and analyst teams.
  • Designed ETL workflows ingesting data from AWS S3, RDS, and internal APIs; transformed and validated raw banking feeds into structured datasets for compliance and business reporting.
  • Reduced runtime of a critical Spark pipeline by ~25–30% through repartitioning skewed joins, converting intermediates to Parquet, and broadcasting small lookup tables to stabilize daily SLAs.
  • Automated batch orchestration in Apache Airflow with modular DAGs, sensors, retries, and SLA-based alerting to ensure reliable daily delivery and fast failure detection.
  • Implemented PySpark and SQL validation checks including schema enforcement, null handling, duplicate detection, and source-to-target row count reconciliation, lowering downstream data issues and manual reconciliation effort.
  • Contributed to secure pipeline practices and CI/CD for sensitive customer and transaction data: field-level masking, access-controlled datasets, Git-based versioning, Dockerized jobs, and controlled promotion across dev/UAT/prod.
Rochester Institute of Technology, GCCIS
Rochester, NY
Data Engineer
Aug 2023 – Apr 2024
Built end-to-end pipelines integrating academic, co-op, and financial systems; produced centralized analytics-ready datasets used by academics, finance, and operations teams.
Tech Stack: PySpark, Apache Airflow, Docker, Tableau, AWS S3, Git
  • Built end-to-end PySpark and SQL pipelines to integrate academic, co-op, and financial data from 3 institutional sources into centralized, analytics-ready datasets used by academia and finance.
  • Designed and modularized 5+ Apache Airflow DAGs with Docker-based execution and Git versioned code, improving deployment efficiency by ~40% versus previous manual processes.
  • Implemented validation rules for schema mismatches, nulls, duplicates, and source-to-target counts, significantly reducing manual reconciliation and increasing dataset trust across teams.
  • Designed partitioning and indexing strategies and optimized data models for reporting and predictive analytics, noticeably reducing query latency for heavier reporting workloads.
  • Developed Tableau dashboards over pipeline outputs to surface operational metrics (latency, throughput, load success) and enable proactive monitoring by analysts.
  • Owned UAT support, requirement gathering, and production rollouts with reproducible Docker deployments and staging verification to ensure reliable dataset delivery.
Sree Rayalaseema Hi-Strength Hypo Ltd.
Hyderabad, India
Data Engineer
Jun 2022 – Apr 2023
Supported production and quality reporting for a chemical manufacturing plant by consolidating DCS, lab, and operational data into reporting databases and analytics outputs.
Tech Stack: Python, SQL, cron, Power BI, Task Scheduler
  • Built automated Python and SQL ETL scripts to ingest plant data from Distributed Control System (DCS) exports and laboratory systems into centralized operational and reporting databases.
  • Productionized ETL workflows using scheduled Python jobs (cron and Task Scheduler), automating recurring ingestion, transformation, and reporting tasks that were previously manual.
  • Implemented data quality and reconciliation checks including null handling and range checks on sensor readings, improving reliability of production and inventory datasets used for reporting.
  • Integrated Statistical Process Control (SPC) logic and hypothesis testing routines into the ETL layer to programmatically flag anomalies in chemical process data for process engineers.
  • Optimized batch sizes and transformation steps to reduce ETL runtime and improve availability of near-real-time operational reports for plant teams.
  • Documented ETL processes, created runbooks, and collaborated with operations and engineering teams to align reporting outputs with business requirements.
Accenture
Hyderabad, India
Data Analyst
Jan 2021 – May 2022
Delivered HR analytics and supported migration of HR reporting to cloud for enterprise HR stakeholders; standardized definitions and automated recurring reports.
Tech Stack: Python, SQL, Power BI, AWS
  • Performed HR analytics using Python, SQL, and Power BI to analyze performance, recruitment, attrition, and diversity data, producing interactive dashboards that reduced recurring manual reporting.
  • Supported migration of legacy HR reports to AWS and automated recurring reporting workflows using Python, improving consistency and reducing manual effort across teams.
  • Built standardized data definitions and consistency checks across HR data sources to improve reporting accuracy and governance for workforce planning.
  • Partnered with HR business partners and IT stakeholders to gather requirements, deliver workforce planning dashboards, and provide UAT support for report releases.
  • Optimized SQL queries and reporting schedules to speed up report refreshes and lower run-time for recurring HR reports.
  • Documented analytics processes and trained HR analysts on dashboard usage and underlying query logic to enable self-service reporting.

Projects

NL2SQL Banking Analytics Chatbot
Tools Used: Python, OpenAI API, Streamlit, SQLite, Pandas
  • Built an end-to-end natural-language-to-SQL application that converts analyst questions into executable queries against a synthetic retail banking schema (customers, accounts, transactions) using the OpenAI API with prompt-engineered schema context.
  • Implemented a SQL guardrail layer enforcing SELECT-only execution and a table allowlist to block destructive or out-of-scope queries; added unit tests to cover SQL injection and prompt injection scenarios.
  • Delivered a Streamlit interface showing generated SQL and live query results; packaged the app with synthetic data generation using Faker for reproducible demos.

Education

Rochester Institute of Technology
Master of Science in Data Science, CGPA: 3.6 • Rochester, NY
Amrita Vishwa Vidyapeetham
Bachelor of Technology in Computer Science and Engineering, CGPA: 4.0 • Coimbatore, India

Certifications

AWS Certified Data Engineer, Associate
Snowflake Data Engineer Certification
IBM Data Science Professional Certificate — IBM

Powered by Drivetube · Create your own profile at drivetube.ai

Explore Drivetube

  • Drivetube Profile — your free digital resume at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
  • Free Job Board verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
  • Job Hunt Program managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
  • Resume Writing Services human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
  • Community Membership from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
  • Drivetube Hire — for employers hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.

Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.

The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.