ANIRUDH SURABI
Professional Summary
Data Engineer with 0 years of professional experience and an MS in Computer Science (Loyola University Chicago, 2026). Hands-on experience building reproducible data pipelines, preparing clinical and observational datasets, and deploying ML-ready feature stores. Skilled in Python, SQL, PostgreSQL, AWS, FAISS, and data visualization (Power BI). Seeking an entry-level Data Engineer role to apply data pipeline design, ETL, and cloud-based data infrastructure skills to business and analytics problems.
Technical Skills
Work Experience
- Led experimental design and end-to-end reproducible workflows for evaluating physician-rated quality of synthetic medical images, coordinating dataset preprocessing and experiment tracking.
- Implemented and trained CNN classifiers (AlexNet, ResNet50, DenseNet201, VGG16) on remote Linux GPU servers using transfer learning to predict physician-assigned quality labels.
- Developed a Pix2Pix image-to-image translation pipeline to synthesize high-fidelity medical images, expanding labeled training data for downstream model experiments.
- Applied data augmentation, targeted fine-tuning, Dropout, and L2 regularization to reduce overfitting and improve model generalization across validation folds.
- Designed evaluation and visualization artifacts (confusion matrices, loss/accuracy curves) and translated results into technical summaries for faculty and cross-disciplinary collaborators.
- Managed dataset curation, annotation verification, and reproducible codebase versioning using Git and Jupyter notebooks to support manuscript preparation.
- Completed an intensive ML certification covering supervised and unsupervised algorithms, model evaluation methods, and practical ML workflows.
- Built end-to-end predictive models using Python, Pandas, and Scikit-learn for classification and regression exercises, applying cross-validation and hyperparameter tuning.
- Implemented clustering and unsupervised pipelines using KMeans and HDBSCAN and engineered TF-IDF text features for analysis assignments.
- Evaluated models using precision, recall, F1, and confusion matrices and documented reproducible experiments in Jupyter notebooks.
- Participated in peer code review and Git-based workflows to iterate on model implementations and ensure code quality during the program.
- Translated theoretical ML concepts into applied pipelines and reusable notebooks that formed the basis for follow-on academic projects.
Projects
- Built a retrieval-augmented generation (RAG) pipeline that ingests clinical reports (PDF/TXT), chunks and embeds content with OpenAI embeddings, and serves answers to natural language clinical queries via GPT-3.5-turbo.
- Designed a custom prompt template and grounding strategy to reduce hallucinations and ensure responses remain strictly within report context for safer clinical Q&A.
- Integrated FAISS vector search and evaluated retrieval quality to improve answer relevance for clinical information retrieval workflows.
- Trained and compared ResNet-18 and Vision Transformer (ViT-B/16) models on the CIFAKE dataset for real vs. AI-generated image classification.
- Achieved strong test accuracy (97.43% for ResNet-18; 98.34% for ViT) and conducted precision, recall, F1, and confusion matrix analysis to validate model robustness.
- Analyzed architectural differences to demonstrate ViT's benefits in capturing global image features over CNN baselines.
- Implemented serial, OpenMP-parallelized, and CUDA GPU-accelerated versions of Conway's Game of Life and benchmarked across grids up to 16,384×16,384.
- Achieved significant CPU and GPU speedups versus the serial baseline, identified memory bandwidth bottlenecks, and documented performance characteristics across implementations.
- Developed transfer learning pipelines for binary and multi-class MRI brain tumor classification using ResNet50, DenseNet201, and VGG16 backbones.
- Applied feature engineering, augmentation, and hyperparameter tuning to achieve consistent test accuracy above 85% across model variants.
- Built a CNN-based real-time pothole detector with sub-100ms inference latency, achieving high operational accuracy during field tests across varied lighting and road conditions.
- Integrated detection output into vehicle navigation pipeline and validated performance with empirical field testing.
- Designed and trained a deep learning model to detect construction materials and site objects, achieving high accuracy and average precision in live deployments.
- Built a worker-facing UI to surface detections and streamline inventory monitoring, reducing material tracking errors in pilot deployments.
- Applied HDBSCAN and KMeans clustering with TF-IDF text features to LinkedIn job posting data to surface hiring demand patterns and potential ghost postings.
- Built an interactive dashboard and validated cluster quality using silhouette and lift metrics to provide actionable hiring trend insights.
- Designed a normalized relational database schema with referential integrity constraints for order, customer, and feedback workflows.
- Implemented complex SQL queries and ORM models to ensure end-to-end data consistency for restaurant operations.
- Developed a React-based task tracker with animated progress visuals and badge mechanics to increase user engagement.
- Led UI debugging, form validation, and collaborative GitHub workflows for deployment across the full development lifecycle.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.