Movva Sandeep
Data Scientist • Plano, Texas • s************@gmail.com • +19******561 • linkedin.com/••••• • drivetube.ai/•••••
Professional Summary
Data Scientist with 6+ years of experience building predictive models, end-to-end ML pipelines, and scalable data platforms on cloud and big-data stacks to drive retail, healthcare, insurance, and pharma analytics.
Technical Skills
Programming Languages: Python,Scala,Java,C++,Shell Scripting
Frameworks and Libraries: Scikit-Learn,TensorFlow,PyTorch,Keras,Matplotlib,Seaborn,Pandas,NumPy
Databases: SQL,Oracle,Microsoft SQL Server,MySQL,MongoDB,HBase,Cassandra,DynamoDB,Snowflake
Cloud and DevOps: AWS EC2, S3, EMR,Azure ADLS Gen2, ADF, Databricks, Azure DevOps,Azure Cognitive Services,Azure DevOps,Jenkins,Maven,Docker,Kubernetes
Data and Analytics: XGBoost,LightGBM,MLlib,NLTK,Apache Spark,PySpark,Hadoop,MapReduce,Hive,Sqoop,Flume,Kafka,Storm,Tableau,Spotfire,Power BI,Plotly,SciPy,Statsmodels
Tools and Methodologies: Git
Skills: PL
Work Experience
Ahold Delhaize
Texas, USA
Data Scientist
May 2024 – Current
Worked for a global food retail & grocery company building data products to improve retail operations, inventory visibility, e-commerce analytics, and customer experience.
Tech Stack: Python, PySpark, Databricks, Spark SQL, XGBoost, Hive, Sqoop, Kafka, Apache Storm, AWS S3, EC2, EMR, DynamoDB, Snowflake, NLTK
- Built and evaluated production-grade predictive models (XGBoost, ensemble methods) using cross-validation, log-loss and ROC/AUC for customer and inventory forecasting; integrated feature selection into model pipelines.
- Prepared master modeling datasets by combining multiple transactional and semi-structured JSON sources, performing feature engineering and scaling on datasets totaling ~2 petabytes for downstream training on Databricks.
- Developed ETL and feature pipelines using PySpark, Spark SQL and Databricks; automated ingestion from Teradata via Sqoop and created Hive tables to serve analytics and ML scoring.
- Implemented real-time clickstream ingestion using Kafka + Apache Storm and persisted events to HDFS/S3 to enable web analytics and personalized recommendations.
- Led a team of 4 developers under Agile/Scrum, instituted code reviews and documentation (source-to-target mappings, unit test cases), and translated business requirements into technical specs.
- Deployed models and data workflows on AWS (EC2, S3, EMR), used DynamoDB and Snowflake for processed outputs, and applied NLP techniques (NLTK) for text analytics and sentiment extraction.
Walgreens
Houston, Texas, USA
Data Scientist
Apr 2023 – Apr 2024
Worked for a pharmacy-led healthcare and retail company, building an Azure enterprise data platform and ML solutions to improve pharmacy operations, inventory and patient engagement.
Tech Stack: Azure Data Factory, Azure Databricks, ADLS Gen2, Azure DevOps, LightGBM, PyCaret, Spark, Hive, Pig, Kafka, Flume, Spotfire, Tableau
- Architected and implemented an Azure enterprise data platform connecting ADF, Databricks and ADLS Gen2 to centralize data ingestion and enable analytics for pharmacy operations.
- Implemented CI/CD pipelines using Azure DevOps to automate deployments of Databricks notebooks, ADF pipelines and model artifacts, improving release cadence and traceability.
- Built and tuned predictive models using LightGBM and PyCaret for manufacturing and demand forecasting use cases; integrated model scoring into Databricks jobs.
- Designed and optimized ETL pipelines using Hive, Pig and Spark; tuned Hive and Pig scripts and resolved performance bottlenecks to reduce job runtimes and improve throughput.
- Developed streaming solutions by integrating Kafka with Spark Streaming and used Flume for log collection to ingest large volumes of behavioral and transaction data into HDFS/ADLS.
- Produced interactive dashboards and visualizations in Spotfire and Tableau; collaborated with business stakeholders to convert analytics insights into technical requirements.
Virtusa / American International Group (AIG)
Mumbai, India
Data Scientist
Jun 2022 – Dec 2022
Consulted for AIG (finance & insurance) and related smart-city initiatives to evaluate ML/AI approaches and design IoT and computer-vision platform components.
Tech Stack: TensorFlow, PyTorch, Keras, MXNet, Azure Cognitive Services, Hive, Sqoop, Storm, Kafka, HDFS, Python, scikit-learn
- Researched and benchmarked ML and deep-learning algorithms for IoT and smart-cities use cases; compared models by accuracy and deployment feasibility to recommend optimal approaches.
- Designed and prototyped computer vision models (object detection, facial recognition) using TensorFlow, PyTorch, MXNet and Keras for video/image analytics.
- Integrated Azure Cognitive Services (Custom Vision, OCR) to create handwriting recognition and Sketch2Code pipelines for converting sketches to HTML components.
- Authored functional requirements and RFP components for ML modelling, data warehouse and data aggregation architectures supporting IoT platform design.
- Implemented data ingestion and processing pipelines using Hive, Sqoop, Storm, Kafka and HDFS to collect sensor and telemetry data for analytics.
- Applied supervised and unsupervised methods (clustering, SVM, neural nets) and evaluated models using scikit-learn and MATLAB to support insurance forecasting and segmentation.
Infosys / Nike
Mumbai, India
Data Scientist
Feb 2021 – Jun 2022
Delivered data science and engineering work for Nike (retail & consumer analytics), focusing on personalization, user engagement and backend data services on cloud and big-data platforms.
Tech Stack: AWS S3, Glue, CodeCommit, Spark, PySpark, Scala, Snowflake, Hadoop, Kafka, Kinesis, TensorFlow, PyTorch, C#, Python
- Built data pipelines and analytics in AWS using CodeCommit, S3, DynamoDB and Glue; implemented API layers and business logic with C# and Python for data access and scoring.
- Developed computer vision and ML algorithms implemented on CPU/GPU targets; authored simulation and evaluation code in Python and C++ for performance testing.
- Used Spark, PySpark and Scala along with Snowflake and Hadoop to process large-scale behavioral and transaction datasets for personalization models.
- Applied ensemble and deep-learning models (GBM, XGBoost, Random Forest, TensorFlow, PyTorch) to classification and recommendation tasks; improved model effectiveness.
- Worked with streaming platforms (Kafka, Kinesis, Spark Streaming) for real-time analytics and used HBase/Cassandra for low-latency storage patterns.
- Collaborated with product teams to translate ML outcomes into business metrics, contributing to initiatives that increased user lifetime by 45% and tripled conversions for target categories.
Astellas Pharma
Mumbai, India
Jr Data Scientist
May 2019 – Jan 2021
Supported analytics for a multinational pharmaceutical company by integrating clinical, lab and operational data to build predictive models and dashboards for outcomes analysis.
Tech Stack: Python, PySpark, MLlib, Random Forest, XGBoost, Tableau, SQL Server, NumPy, Pandas, Scikit-Learn
- Integrated and standardized diverse clinical and operational datasets, performed preprocessing (imputation, normalization, outlier handling) and curated analysis-ready datasets.
- Developed predictive models using Random Forest and XGBoost to forecast customer behavior and retention; implemented end-to-end analytics pipelines in PySpark and Python.
- Delivered dashboards and automated reports using Tableau and Python visualization libraries to present findings to management and technical stakeholders.
- Applied classification and clustering techniques (SVM, KNN, K-means) and used MLLib for scalable model training on large datasets.
- Predicted returning-customer likelihood using ensemble models, contributing to a solution that achieved 90% monthly customer retention in tracked cohorts.
- Collaborated with domain teams to translate clinical/business questions into analytic tasks and supported deployment and monitoring of models in production.
Education
Southeast Missouri State University
Masters / Computer Science • Missouri, USA • Jan 2023 – May 2024
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.