Skip to content

D N RAKESH

Data & Business Analyst • Bangalore, India • r***********@gmail.com • +91*******343 • linkedin.com/••••• • drivetube.ai/•••••

Professional Summary

Data Analyst with less than 1 year of experience delivering end-to-end analytics and machine learning solutions — from data extraction and exploratory analysis to model development and deployment. Experienced in SQL, Python, MySQL, Power BI and Streamlit, with practical project work in forecasting, classification, and prediction for marketing, finance, retail and operations use cases. Strong at translating complex datasets into actionable dashboards and clear stakeholder presentations.

Technical Skills

Programming Languages: Python,
Frameworks and Libraries: Pandas,NumPy,Scikit-learn
Databases: SQL,MySQL
Cloud and DevOps: Streamlit,Jupyter Notebook
Data and Analytics: Excel,Power BI,Exploratory Data Analysis,Dashboarding,XGBoost,LightGBM,Random Forest,ARIMA,Prophet,Feature Engineering,SMOTE,StandardScaler,One-Hot Encoding,Cross-Validation,Hyperparameter Tuning,Threshold Tuning
Tools and Methodologies: Git,GitHub
Model Evaluation & Metrics: ROC-AUC,Precision,Recall,F1-score

Work Experience

Rubixe
Bangalore, India
Data Scientist Intern
Feb 2026 – Present
Worked at Rubixe, a technology company providing AI and automation solutions; supported end-to-end data science workflows and analytics for business use cases.
Tech Stack: Python, Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, ARIMA, Prophet, SMOTE, MySQL, Streamlit, Git, Jupyter Notebook, Power BI, Excel
  • Built and maintained data extraction and transformation pipelines using Python (Pandas) and MySQL to produce model-ready datasets for forecasting and classification tasks.
  • Conducted exploratory data analysis and feature engineering to identify predictive signals and reduce data leakage, improving model inputs and interpretability.
  • Developed, validated, and tuned machine learning models (LightGBM, XGBoost, Random Forest, ARIMA/Prophet) addressing class imbalance with SMOTE and scale_pos_weight.
  • Deployed interactive prototypes and demos using Streamlit and maintained reproducible code and version control with Git/GitHub for stakeholder review.
  • Prepared business-facing dashboards and reports with Power BI and Excel to translate model outputs into actionable insights for operations, marketing, and finance teams.
  • Documented modeling decisions, evaluation metrics, and deployment instructions in Jupyter notebooks and README files to ensure reproducibility and handover readiness.

Projects

Bajaj Two-Wheeler Spare Parts Demand Forecasting
Tools Used: MySQL, Python, ARIMA, Prophet, Random Forest, Feature Engineering
  • Built a full analytics pipeline on a MySQL dataset of 28,482 records spanning ~19 months; performed EDA and time-series feature engineering to support demand forecasting.
  • Evaluated ARIMA, Prophet, Moving Average, Linear Regression and Random Forest models and delivered the best-performing model as a business-ready forecasting solution for spare-parts planning.
Portuguese Bank Marketing Prediction
Tools Used: Python, LightGBM, SMOTE, One-Hot Encoding, Streamlit, GitHub
  • Engineered 53 features and applied SMOTE, StandardScaler and One-Hot Encoding to prepare the dataset for modeling.
  • Selected LightGBM as the top model achieving ROC-AUC of 0.9478; deployed an interactive Streamlit app and published documentation and model comparison report to GitHub.
Santander Customer Transaction Prediction
Tools Used: Python, LightGBM, SMOTE, Threshold Tuning, Streamlit
  • Tackled a 90:10 class imbalance using SMOTE and scale_pos_weight; tuned classification threshold to 0.35 to optimize LightGBM performance for business objectives.
  • Delivered a deployable Streamlit application with session-state handling and model evaluation (ROC-AUC 0.7642) hosted via GitHub for stakeholder access.
Home Loan Default Prediction
Tools Used: Python, XGBoost, Feature Engineering, Threshold Tuning, Data Preprocessing
  • Processed a multi-table dataset of 307,511 records across 7 CSVs and engineered domain features such as Credit-Income Ratio and Risk Score.
  • Trained XGBoost with class imbalance handling (scale_pos_weight=11) and prioritized recall and F1-score; improved recall from 3% to 39% through threshold tuning to reduce costly false negatives.
Flight Fare Prediction
Tools Used: Python, Random Forest, Model Debugging, Streamlit
  • Rebuilt and debugged an existing notebook (fixed five critical bugs) and trained a Random Forest model that achieved R² of 0.84.
  • Deployed the predictive model to Streamlit Community Cloud and maintained a public GitHub repository for reproducibility and review.

Education

Sri Venkateshwara College of Engineering and Technology
B.Tech, Computer Science & Engineering (Data Science) — CGPA: 70% • 2022 – 2026
Narayana Junior College, Hindupur
Class XII (Senior Secondary) — 66% • Hindupur, India • 2020 – 2022
Panchajanya Brillant's High School, Hindupur
Class X (Secondary School) — 98% • Hindupur, India • 2020

Certifications

Python Programming for Data Analytics — L&T EduTech
Foundation of Data Engineering — Course Completion Certificate

Powered by Drivetube · Create your own profile at drivetube.ai