Parthavanaidu Jyotha
Machine Learning Engineer • San Jose, CA • p*************@gmail.com • 513****193 • linkedin.com/••••• • drivetube.ai/•••••
Professional Summary
Machine Learning Engineer with 3+ years of experience building production ML and LLM systems for fraud, risk, personalization, and RAG applications. Experienced across the ML lifecycle using Python, Spark, Databricks, and SQL for EDA, feature engineering, model training, threshold tuning, containerized deployment, experiment tracking, and scheduled serving. Recent work includes explainable fraud and anomaly detection, LLM fine-tuning and retrieval-augmented systems, and distributed feature and scoring pipelines on Databricks and AWS.
Technical Skills
Programming Languages: Python,Java,Bash
Frameworks and Libraries: scikit-learn,PyTorch,LangChain,Hugging Face,Pandas,NumPy,FastAPI
Databases: SQL,Snowflake
Cloud and DevOps: Docker,Kubernetes,MLflow,Experiment Tracking,Model Serving,Drift Monitoring,CI,CD,AWS,GCP Vertex AI
Data and Analytics: XGBoost,LightGBM,Classification,Regression,Anomaly Detection,Recommendation Systems,Feature Engineering,Hyperparameter Tuning,A,B Testing,Spark,PySpark,Databricks,Airflow,Flyte,Kafka,ETL,Feature Stores,Batch Inference,Streaming Inference
Tools and Methodologies: Git
GenAI and LLMs: LLMs,LLM Fine-Tuning,Retrieval Augmented Generation,MCP,Agentic Workflows,Embeddings,Vector Search,FAISS,Prompt Engineering,LLM Evaluation
Work Experience
Expedia Group
San Jose, CA
Machine Learning Engineer
Mar 2026 – Present
Worked on travel and hospitality personalization and CRM systems processing booking, trip, property, and engagement signals to serve recommendations and marketing workflows.
Tech Stack: Python, Spark, PySpark, Databricks, SQL, Docker, Flyte, MLflow, Git
- Built CRM personalization workflows in Python and PySpark on Databricks, converting booking, trip, property, and engagement signals across millions of traveler records into model-ready features for downstream ML pipelines.
- Designed batch scoring and feature engineering logic (Databricks, Spark, SQL) to produce stable features consumed by personalization models and downstream reporting, improving pipeline reliability for marketing campaigns.
- Integrated LLM-based recommendation components using prompt templates, structured JSON output contracts, ranking constraints, and MLflow experiment tracking to add contextual recommendations to line-of-business workflows.
- Converted notebook prototypes into reusable Dockerized Flyte workflows with Git-based CI, cutting manual deployment effort by ~30% and enabling reproducible scheduled batch scoring.
- Instrumented models and workflows with MLflow experiment tracking and drift monitoring to detect data shifts and accelerate model iteration and rollback.
- Optimized Databricks job configurations and Spark transformations to improve runtime and cost efficiency of scheduled scoring jobs while preserving feature consistency.
PayPal
San Jose, CA
AI/ML Engineer (Contract)
Sep 2024 – Mar 2026
Worked on payments fraud, anomaly detection, and customer-service automation for a global payments platform, delivering models and LLM-based tools used by investigators and analysts.
Tech Stack: Python, PySpark, XGBoost, LightGBM, scikit-learn, LangChain, FAISS, FastAPI, Airflow, Kafka, Databricks, MLflow
- Built and evaluated fraud, anomaly, and behavioral risk models (XGBoost, LightGBM, scikit-learn) end-to-end from EDA and hyperparameter tuning to threshold optimization and deployment, improving fraud-detection precision by ~27%.
- Fine-tuned open-source LLMs to automate contact and complaint categorization across multiple global markets, scaling inference to 10M+ contacts per month for customer service compliance reporting.
- Designed retrieval-augmented workflows (LangChain, FAISS) for policy lookup, dispute summarization, and analyst copilots; exposed RAG endpoints via FastAPI and reduced analyst lookup time by ~45%.
- Developed an MCP server using agentic workflows to automate root-cause diagnosis of merchant conversion and authorization rate drops, accelerating triage processes for payment volume recovery.
- Engineered unified Python, PySpark, and SQL pipelines to join transaction, device, and behavioral signals for batch and streaming scoring on Airflow, Kafka, and Databricks; instrumented pipelines with MLflow tracking and drift monitoring to reduce data-prep time by ~35%.
- Performed decision-threshold tuning, explainability outputs for investigator review, and operationalized model alerts to improve detection precision/recall tradeoffs in production monitoring.
Tata Consultancy Services
Bengaluru, India
Machine Learning Intern
Dec 2022 – Dec 2023
Delivered churn modeling and preprocessing automation for client analytics projects, focusing on data preparation, model training, and evaluation.
Tech Stack: Python, Pandas, NumPy, scikit-learn, XGBoost, Random Forest
- Trained and evaluated churn classification models (Logistic Regression, Random Forest, XGBoost) on a 10,000+ record dataset, achieving 84% accuracy across iterative experiments.
- Automated preprocessing pipelines for missing values, outlier handling, encoding, and train-test splits using Pandas and NumPy, reducing manual data-prep effort by ~30%.
- Performed exploratory data analysis and feature engineering to identify high-signal predictors and improve model performance and interpretability.
- Implemented cross-validation and hyperparameter tuning workflows to select robust model variants and track experiment outcomes.
- Produced evaluation reports and visualizations to communicate model performance and recommendations to project stakeholders.
- Collaborated with team engineers to package model artifacts and prepare initial deployment scripts and documentation to support handoff.
Projects
RecruitEdge ATS - AI Resume Matching
Tools Used: Python, LangChain, FAISS, FastAPI, Vector Search, Embeddings
- Built a retrieval-based resume-to-job matching service using FAISS and embeddings with LLM-based reranking, served via FastAPI and evaluated across 100+ resume profiles.
- Improved retrieval quality by tuning chunking strategy, selecting embedding models, and optimizing reranking thresholds against a labeled relevance set.
Machine Learning for Spatial Data
Tools Used: Python, SVM, CNN, Computer Vision
- Benchmarked SVM and CNN architectures for spatial image classification across preprocessing, feature extraction, and cross-validation to establish a reproducible pipeline.
- Increased classification accuracy over the SVM baseline through data augmentation and hyperparameter tuning and performed error analysis to isolate confused classes.
Fraud Detection in Online Retail Transactions
Tools Used: Python, SQL, Isolation Forest, Local Outlier Factor, K-Means
- Built an anomaly-detection framework using Isolation Forest, LOF, and K-Means over engineered transaction-level features, flagging 768+ suspicious transactions.
- Benchmarked against a rule-based baseline and improved precision on flagged transactions by ~22%.
Education
University of Cincinnati
MS in Information Technology • Cincinnati, OH • Jan 2024 – May 2025
SCSVMV University
BE in Computer Science Engineering • Tamil Nadu, India • Aug 2019 – Jun 2023
Powered by Drivetube · Create your own profile at drivetube.ai