D N Rakesh
Data & Business Analyst • Bangalore, India • r***********@gmail.com • +91*******343 • github.com/••••• • drivetube.ai/•••••
Professional Summary
Analytical and detail-oriented Data Science professional with hands-on experience delivering end-to-end analytics and machine learning solutions — from data extraction and exploratory analysis to model deployment. Proven ability to translate complex datasets into actionable business insights across marketing, finance, retail, and operations domains. Skilled in SQL, Python, and BI storytelling, with a track record of building forecasting, classification, and prediction pipelines deployed via Streamlit. Seeking to apply strong analytical and stakeholder-communication skills in a Data Analyst / Business Analyst / Data Scientist role.
Technical Skills
Programming Languages: Python
Frameworks and Libraries: Pandas,NumPy,Scikit-learn,XGBoost,LightGBM,Matplotlib,Seaborn
Databases: SQL,MySQL,SQL Querying
Cloud and DevOps: Streamlit,Jupyter Notebook
Data and Analytics: Classification,Regression,Ensemble Methods,Feature Engineering,Imbalanced Learning,Hyperparameter Tuning,Model Evaluation,Exploratory Data Analysis,Data Visualization,Power BI,Data Cleaning,One-Hot Encoding,Standard Scaling,SMOTE
Tools and Methodologies: Git,GitHub
Time Series & Forecasting: ARIMA,Prophet,Moving Average,Time Series Feature Engineering
Reporting & Business Skills: Dashboarding,Business Reporting,Stakeholder Communication,Data Storytelling
Work Experience
Rubixe
Bangalore
Data Scientist Intern
Feb 2026 – Present
Worked at Rubixe, an AI/ML and automation technology company; contributed to end-to-end data science workflows supporting analytics, model development, and deployment for business use cases.
Tech Stack: Python, Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, ARIMA, Prophet, MySQL, Streamlit, Git, Power BI, SMOTE
- Implemented end-to-end ML pipelines (data ingestion, cleaning, feature engineering) using Python, Pandas and MySQL on datasets up to 307,511 rows to prepare production-ready training datasets and reproducible notebooks.
- Engineered time-series forecasting solutions for spare-parts demand using ARIMA, Prophet and Random Forest on a MySQL dataset of 28,482 records; selected best model and delivered a business-ready forecasting pipeline for demand planning.
- Built and tuned classification models with LightGBM and XGBoost for marketing and credit-risk use cases; addressed class imbalance with SMOTE and scale_pos_weight and optimized decision thresholds to prioritize recall.
- Improved credit-risk model recall from 3% to 39% through threshold tuning and targeted feature engineering, increasing model sensitivity to low-frequency default events.
- Developed and deployed interactive Streamlit applications and published reproducible code to GitHub to enable stakeholder demos and self-serve model exploration; managed session-state and deployed to Streamlit Cloud.
- Produced model comparison reports and dashboards using Power BI; evaluated models with ROC-AUC (up to 0.9478 for a bank-marketing model) and regression metrics (R² 0.84 for a fare-prediction model) to drive model selection decisions.
Projects
Bajaj Two-Wheeler Spare Parts Demand Forecasting
Tools Used: MySQL, Python, ARIMA, Prophet, Random Forest, Feature Engineering, EDA
- Built a full analytics pipeline on a MySQL dataset of 28,482 records spanning ~19 months, performing cleaning, time-series decomposition and feature engineering for demand forecasting.
- Compared ARIMA, Prophet, Moving Average, Linear Regression and Random Forest models and selected the best-performing model for spare parts demand planning.
- Packaged forecasting pipeline for business consumption and documented assumptions and model selection rationale for planners.
Portuguese Bank Marketing Prediction (PRCP Project)
Tools Used: Python, LightGBM, SMOTE, Feature Engineering, One-Hot Encoding, Streamlit
- Engineered 53 features and preprocessed data with standard scaling and one-hot encoding to prepare inputs for tree-based models.
- Addressed imbalance using SMOTE and identified LightGBM as the top model achieving ROC-AUC of 0.9478.
- Deployed an interactive Streamlit app and published model comparison documentation and README on GitHub for reproducibility.
Santander Customer Transaction Prediction (PRCP-1003)
Tools Used: Python, LightGBM, SMOTE, Threshold Tuning, Model Evaluation, Streamlit
- Handled a 90:10 class imbalance with SMOTE and scale_pos_weight and tuned classification threshold to 0.35 to optimize LightGBM performance for business priorities.
- Achieved ROC-AUC of 0.7642 and built a Streamlit app with session-state to demonstrate model behavior and edge-case handling.
- Documented preprocessing steps and model selection decisions to support reproducibility and stakeholder review.
Home Loan Default Prediction (Personal Project)
Tools Used: Python, XGBoost, Feature Engineering, Threshold Tuning, Data Processing
- Processed multi-table dataset of 307,511 records across 7 CSVs and engineered domain features such as Credit-Income Ratio and Risk Score.
- Trained XGBoost with scale_pos_weight=11 and focused on recall and F1-score to reduce costly false negatives, improving recall from 3% to 39% via threshold tuning.
- Packaged model and inference pipeline for evaluation and stakeholder review.
Flight Fare Prediction (PRCP-1025)
Tools Used: Python, Random Forest, Debugging, Model Deployment, Streamlit
- Rebuilt and debugged an existing notebook, addressing five critical bugs to restore pipeline correctness and reproducibility.
- Trained a Random Forest regression model achieving R² of 0.84 and deployed the model to Streamlit Community Cloud with public GitHub repository.
- Documented fixes and model assumptions to support future maintenance.
Education
Sri Venkateshwara College of Engineering and Technology
B.Tech, Computer Science & Engineering (Data Science) • 2022 – 2026
Narayana Junior College, Hindupur
Class XII (Senior Secondary), State Board • Hindupur • 2020 – 2022
Panchajanya Brillant's High School, Hindupur
Class X (Secondary School), State Board • Hindupur • 2020 – 2020
Certifications
Python Programming for Data Analytics — L&T EduTech
Foundation of Data Engineering — Course Completion Certificate
Powered by Drivetube · Create your own profile at drivetube.ai