Shajahan Sayyed
Data Analyst • West Haven, CT • s*****************@gmail.com • +12******690 • drivetube.ai/•••••
Professional Summary
Data Analyst with 0 years of experience applying SQL, Python, ETL and BI tools to transform operational data into stakeholder-ready insights. M.S. Data Science candidate with internship experience building automated ETL pipelines, interactive Tableau dashboards, and predictive models to reduce manual effort and inform business decisions.
Technical Skills
Programming Language: Python,SQL
Databases: MySQL,Relational Databases
Cloud Platforms: Amazon Web Services,S3
Data Warehousing: Data Warehousing
Data Analysis & Visualization: Pandas,NumPy,Matplotlib,Seaborn,Data Wrangling,Data Cleaning,Statistical Analysis,Tableau,Power BI,Microsoft Excel,boto3
Machine Learning & AI: K-Means,Clustering,Random Forest,Logistic Regression,Predictive Modeling,SHAP
AI/ML Frameworks & Libraries: scikit-learn
MLOps: Feature Engineering
Data Integration & ETL: ETL
Work Experience
Avishkar Tech Solutions
Hyderabad, India
Data Analyst Intern
May 2023 – Dec 2023
Supported manufacturing and plant operations analytics; delivered reporting and insights used by plant management and business-unit stakeholders.
Tech Stack: SQL, Python, Pandas, NumPy, Tableau, ETL Pipeline Design
- Queried and analyzed manufacturing and operational datasets using SQL (multi-table JOINs, GROUP BY, window functions) and Python (Pandas) to surface process gaps and performance drivers; analysis contributed to a 20% improvement in campaign performance insights used by management.
- Built and maintained 5+ interactive Tableau dashboards for plant management and business-unit stakeholders, reducing manual reporting time by 30% and improving operational visibility for daily decision-making.
- Automated recurring reporting workflows with Python (Pandas, NumPy), eliminating 8–10 hours of manual work per week and increasing report freshness for stakeholders.
- Implemented ETL steps for ingestion, cleaning, and validation of multi-source operational data to standardize formats and improve downstream reporting reliability across pipelines.
- Documented data definitions, transformation logic, and dashboard requirements; presented findings and prioritized recommendations that informed operational and scheduling adjustments.
- Worked cross-functionally with operations and business teams to translate technical results into stakeholder-ready visualizations and action items, accelerating adoption of data-driven process improvements.
Projects
E-Commerce Sales & Customer Behavior Analysis
Tools Used: Python, boto3, AWS S3, MySQL, Tableau, SQL, Pandas
- Designed an end-to-end analytics pipeline on 100k+ e-commerce records: ingested CSVs from AWS S3 using boto3, cleaned and transformed data with Pandas, and loaded processed tables into MySQL for analysis.
- Used SQL aggregations and window functions to identify top revenue product categories and map regional performance; found a category that contributed 10.6% of total revenue across 27 states.
- Analyzed delivery delay impact on customer satisfaction, linking orders delayed >7 days with average review score drops from 4.5 to 3.1; recommended prioritized logistics fixes to reduce churn.
Retail Sales KPI Dashboard
Tools Used: SQL, Python, Tableau, Power BI, Pandas
- Processed 50,000+ transactional records to extract revenue trends, seasonality, and regional performance gaps using SQL and Python.
- Developed KPI-driven dashboards in Tableau and Power BI visualizing revenue growth, margin, and regional sales, reducing ad-hoc reporting requests by 40%.
- Standardized data collection and validation steps and delivered documentation for management use in monthly business reviews.
Customer Segmentation & Churn Analysis
Tools Used: Python, SQL, Tableau, K-Means Clustering, Pandas
- Segmented 80,000+ customer records by recency, frequency, and monetary metrics using SQL and K-Means, identifying high-risk churn cohorts and priority retention targets.
- Built visualizations in Tableau to communicate cohort behavior and recommended targeted campaigns projected to reduce churn by ~15%.
- Cleaned and reconciled CRM fields using Pandas to ensure segmentation accuracy and repeatability across analyses.
Employee Attrition Prediction Model
Tools Used: Python, Scikit-learn, Random Forest, Logistic Regression, Power BI, SHAP
- Built classification models (Random Forest, Logistic Regression) on 15,000+ HR records to predict attrition with 87% accuracy; performed feature engineering to surface top predictors (tenure, compensation, engagement).
- Used SHAP values to explain model drivers and created a Power BI dashboard showing attrition risk by department and role level for HR decision-making.
- Delivered recommendations that informed workforce planning scenarios and projected turnover mitigation strategies.
Automated ETL Pipeline for Sales Reporting
Tools Used: Python, Pandas, SQL, AWS S3, Tableau, ETL Pipeline Design
- Designed and implemented an automated ETL pipeline to ingest sales data from CSVs, databases, and S3; transformed and validated 100k+ records weekly into a centralized reporting store.
- Scheduled pipeline jobs and automated data quality checks, eliminating 12+ hours of manual preparation per week and reducing reporting errors by 35%.
- Built downstream Tableau dashboards to provide near real-time sales KPIs and documented the pipeline architecture for engineering and analytics teams.
Education
University of New Haven
M.S., Data Science • West Haven, CT • 2025
Coursework: Statistical Analysis, Data Visualization, Machine Learning, Data Warehousing, Business Intelligence, Database Management
KL University
B.Tech., Computer Science • India • 2019 – 2023
Certifications
AWS Certified Developer – Associate — Amazon Web Services
Cisco Data Analytics Essentials — Cisco
Data Visualization: Storytelling
Powered by Drivetube · Create your own profile at drivetube.ai