Skip to content

Jithendra Chava

Data Engineer • United States • j**************@gmail.com • +15******919 • linkedin.com/••••• • drivetube.ai/•••••

Professional Summary

Data Engineer with 4+ years of experience designing scalable data pipelines, cloud-based data platforms, machine learning infrastructure, and Generative AI solutions across banking, healthcare, and enterprise domains. Improved data processing efficiency by up to 42% by building AI-ready ETL pipelines, real-time analytics, and Retrieval-Augmented Generation solutions. Delivers secure, high-quality, and governed data platforms that accelerate analytics, machine learning, and enterprise AI initiatives.

Technical Skills

Programming Languages: Python,R
Libraries & Frameworks: Pandas,NumPy,Matplotlib,Scikit-learn,TensorFlow,PyTorch,Hugging Face Transformers
Databases & Warehousing: SQL,PostgreSQL,Oracle,Microsoft SQL Server,MySQL,Snowflake,Google BigQuery,Redshift
Big Data & ETL Development: Apache Spark,PySpark,Spark SQL,Delta Lake,Apache Airflow,dbt,AWS Glue,Azure Data Factory,Data Pipelines,Databricks
Cloud Platforms & DevOps: AWS,Microsoft Azure,Google Cloud Platform,Terraform,Docker,CI/CD Pipelines,Azure DevOps
Data Analytics & Visualization: Power BI,Tableau,Microsoft Excel,KPI Reporting,Data Profiling
Machine Learning & Generative AI: Deep Learning,Predictive Analytics,Natural Language Processing,Large Language Models,Prompt Engineering,Retrieval-Augmented Generation,LangChain,LangGraph,OpenAI API,Hugging Face
MLOps & AI Infrastructure: MLflow,Kubeflow,Feature Engineering,Model Training,Model Deployment,Model Monitoring,Vector Databases,Embeddings
APIs & Data Formats: RESTful APIs,JSON,XML,Parquet,Avro
Version Control & Methodologies: Git,GitHub,Agile,Scrum,JIRA,Confluence,ServiceNow

Work Experience

US Bank
United States
Data Engineer
Jan 2025 – Present
Major U.S. banking and financial-services institution where data engineering supported high-volume transactions, real-time fraud analytics, and customer insight platforms.
Tech Stack: Python, PySpark, SQL, Snowflake, Apache Airflow, AWS, ETL, Machine Learning, LLMs, AI, Data Validation, Anomaly Detection, Data Quality, Data Governance
  • Engineered scalable ETL pipelines using PySpark, SQL, and Snowflake to integrate high-volume banking data, improving data availability by 40% for enterprise analytics, regulatory reporting, and AI-driven fraud detection.
  • Orchestrated automated data ingestion workflows using Apache Airflow, enabling reliable processing for consumer banking, lending, and real-time analytics while preparing LLM-ready datasets for AI applications.
  • Optimized large-scale financial data transformations using SQL and PySpark, reducing query execution time by 45% and accelerating reporting, predictive analytics, and machine learning feature engineering.
  • Collaborated with data engineers, data scientists, business analysts, and compliance teams to deliver secure, AI-enabled data solutions, ensuring on-time delivery of strategic banking initiatives.
  • Implemented cloud-native ETL pipelines using Snowflake, Python, AWS, and Apache Airflow, reducing infrastructure and data processing costs by $240K annually through workflow automation, resource optimization, and scalable data architecture.
  • Developed automated data validation and anomaly detection frameworks using Python and ML techniques, reducing production data defects by 30% and improving data quality for AI models and business intelligence reporting.
HCA Healthcare
United States
Data Engineer
Feb 2024 – Dec 2024
Nationwide hospital and healthcare system where data pipelines enabled analytics on electronic health records, billing, and clinical operations within compliance constraints.
Tech Stack: Azure Data Factory, Delta Lake, Apache Kafka, Spark SQL, Azure Databricks, Azure OpenAI, LangChain, FAISS, Retrieval-Augmented Generation, dbt, ETL, Machine Learning, Data Governance, HIPAA
  • Architected scalable healthcare data pipelines using Azure Data Factory and Delta Lake to integrate EHR, patient encounters, and clinical datasets, improving data accessibility by 42% for enterprise analytics, AI applications, and regulatory reporting.
  • Engineered real-time streaming pipelines with Apache Kafka and Spark Structured Streaming, enabling AI-driven patient monitoring, predictive healthcare analytics, and operational dashboards for clinical decision-making across hospital systems.
  • Optimized large-scale healthcare data transformations using Spark SQL and Azure Databricks, improving processing performance by 38% while delivering high-quality datasets for machine learning models, reporting, and advanced analytics.
  • Developed AI-ready data platforms by integrating Azure OpenAI, LangChain, and FAISS vector databases to build Retrieval-Augmented Generation solutions for intelligent clinical document search and automated knowledge retrieval.
  • Automated scalable ETL/ELT workflows using Azure Databricks and dbt, reducing manual processing effort by 40% while improving data reliability for enterprise reporting, predictive analytics, and AI model training pipelines.
  • Implemented automated data quality, governance, and HIPAA compliance frameworks using dbt and validation pipelines, reducing data exceptions by 28% while ensuring trusted datasets for AI, business intelligence, and healthcare analytics.
Capgemini
India
Data Engineer
Feb 2021 – Jun 2023
Global consulting and IT services firm; built data platforms for diverse enterprise clients, integrating disparate sources to power centralized reporting and analytics.
Tech Stack: Azure Data Factory, PySpark, SQL, Databricks, dbt, Azure OpenAI, MLflow, Delta Lake, ETL, Machine Learning, Retrieval-Augmented Generation, Generative AI, Data Governance, Anomaly Detection
  • Modernized enterprise data integration using Azure Data Factory and AI-driven data enrichment, improving processing efficiency by 40% while enabling trusted datasets for analytics and Generative AI applications.
  • Engineered scalable PySpark ETL pipelines to process structured and semi-structured data, creating AI-ready datasets for machine learning models and Retrieval-Augmented Generation solutions.
  • Optimized SQL and Spark transformation workflows through intelligent performance tuning, reducing pipeline execution time by 43% and improving SLA compliance across enterprise data platforms.
  • Collaborated with data scientists, ML engineers, architects, and stakeholders to build AI-ready data platforms that accelerated machine learning development and enterprise analytics initiatives.
  • Implemented automated Databricks workflows with MLflow and Delta Lake, increasing pipeline automation efficiency by 36% while supporting AI model training and batch inference.
  • Strengthened data governance using dbt and AI-assisted anomaly detection, reducing data inconsistencies by 30% and improving data quality for BI, analytics, and AI workloads.

Education

University of Cincinnati
Master of Science in Information Technology • Cincinnati, OH • Aug 2023 – Dec 2024

Powered by Drivetube · Create your own profile at drivetube.ai