Vijay Shankar Chamakuri
Professional Summary
Data Analyst with 3.5 years of experience building reproducible analytics, data models, and product-facing pipelines across consumer platforms, higher-education research, and pharmaceutical operations. Proficient in Python, SQL, TypeScript, dbt, DuckDB and DynamoDB; I design ETL/ELT, governed marts, and metric contracts to produce reconciled KPI pipelines and operational dashboards. I have implemented ML-ready feature pipelines and RNA‑seq analysis workflows using scikit-learn, PyTorch and FastAPI, and automated CI/CD with GitHub Actions, Docker and Terraform. I deliver stakeholder-ready QuickSight and Tableau dashboards and operational anomaly alerts that inform planning and finance decisions. I seek roles that combine data engineering and analytics ownership—building reliable data products, data governance, and instrumented pipelines that drive product and business outcomes.
Technical Skills
Work Experience
- Designed DynamoDB data models, global secondary indexes, and paginated query patterns to support 15 domain tables for product, engagement, content, and moderation data.
- Built asynchronous Node.js aggregation services that merged recipe, author, rating, and recency signals into ranked discovery datasets while excluding private accounts.
- Developed TypeScript pipelines that standardized inconsistent ingredient records, consolidated compatible measurement units, removed duplicates, and preserved source lineage for grocery-list generation.
- Delivered React Native mobile features for content discovery and saved-content workflows to improve product-facing data flow.
- Engineered Express.js backend endpoints and integrated Stripe to enable subscription checkout and billing workflows.
- Configured AWS S3 and CloudFront to optimize media storage and CDN delivery for mobile and web clients.
- Applied Amazon Rekognition to automate image-safety scanning in moderation pipelines and implemented JWT-based authentication for client sessions.
- Integrated Google OAuth for third-party sign-in and produced Postman collections for API testing and developer onboarding.
- Led faculty-supervised analysis of bulk RNA‑seq responses to ten PFAS exposures, standardizing precomputed contrast tables into a unified analytical universe of 13,852 genes.
- Devised a consensus k‑space clustering workflow to align 30 independently seeded fits into five modules, improving module stability for downstream interpretation.
- Applied leave-one-chemical-out cross-validation, empirical null testing, permutation methods, and Benjamini‑Hochberg correction to control false discoveries across chemical contrasts.
- Performed Gene Ontology enrichment across namespaces to surface biologically coherent module annotations for faculty interpretation.
- Built a FastAPI planning service that enforces dataset-compatibility checks, retrieves evidence, validates citations, and routes unsupported requests to human-review gates.
- Automated CI and runtime checks — Ruff, mypy, pytest, Docker verification, PyTorch and TensorFlow smoke tests, and Terraform validation — using GitHub Actions to maintain reproducibility.
- Built Amazon QuickSight dashboards for inventory, demand patterns, and operational KPIs so stakeholders could identify exceptions during planning reviews.
- Conducted statistical analysis and forecasting on inventory history using Python and statsmodels to inform inventory-planning and cost-management decisions.
- Developed near-real-time anomaly detection pipelines in Python to alert operational stakeholders to unusual inventory conditions for rapid follow-up.
- Automated ETL validation and reconciliation checks to ensure KPI accuracy prior to dashboard refreshes.
- Documented metric definitions and maintained a versioned metric contract to standardize reporting across planning workflows.
- Translated stakeholder requirements into dashboard filters, drill paths, and authored training materials to accelerate stakeholder adoption.
- Designed reusable SQL reporting workflows that converted fragmented operational data into standardized KPI datasets for recurring management analysis in a pharmaceutical environment.
- Owned end-to-end reporting delivery from stakeholder requirements and metric definitions through SQL transformation, output validation, and presentation of findings.
- Standardized KPI calculations and reporting logic across recurring analyses to provide management with consistency for performance evaluation.
- Translated business questions into targeted SQL analyses that surfaced performance drivers, exceptions, and decision-ready findings for cross-functional stakeholders.
- Automated scheduled report generation and produced Excel workbooks for management review and archival.
- Implemented validation checks and reconciliation scripts using pandas and SQL to ensure data accuracy in recurring reports.
Projects
- Architected a TypeScript monorepo with a side-effect-free domain core, SQLite persistence via Drizzle, a CLI, and a server-rendered recruiter workspace for candidate triage and audit history.
- Separated probabilistic extraction from deterministic evaluation: LLM-compatible adapters return verbatim evidence spans while locked rubrics and scoring functions calculate reviewable candidate results.
- Implemented evidence-span relocation, confidence scoring, explicit escalation tasks, human correction flows, immutable audit events, and versioned supersession so unsupported claims fail visibly.
- Built a hermetic seven-candidate proving corpus covering scored, rejected, missing-evidence, and hallucinated-quote paths; verified the monorepo with type checks and 1,718 passing Vitest tests across suites.
- Built a tested SQL and dbt warehouse with 34 models, 80 dbt tests, and a DuckDB star schema to explain MRR movement across 36 months, reconciling invoices, payments, refunds, and recognized revenue to $0.00 variance.
- Detected modeled failed-payment exposure of $1.46M across 638 invoices and surfaced prioritized billing exceptions for Finance Operations to remediate.
- Authored a four-dashboard Tableau workbook tied to governed marts and implemented 128 automated checks validating displayed measures.
- Extended the analytical stack with PySpark and a Hive-compatible table for 600,000 synthetic usage events and orchestrated pipelines with an Airflow DAG.
- Automated more than 100 pytest cases, Ruff and mypy checks, and GitHub Actions across Python 3.11/3.12 to enforce metric correctness and prevent regressions.
- Analyzed NHANES 2017–2018 data using survey weights and complex design to estimate population-level diagnostic gaps and subgroup differences.
- Validated Python survey estimates and standard errors against R’s survey package within numerical tolerance and halted acquisition when CDC source files changed via SHA‑256 integrity checks.
- Audited how screening-model label choice affected subgroup errors using scikit-learn and statsmodels and produced sensitivity analyses for stakeholders.
- Built a reproducible SQL/DuckDB, Python, and R pipeline with 82 tests and CI that reruns the analysis on real CDC data and fails if reported numbers drift.
- Delivered a stakeholder-ready dashboard, an Excel quality-review workbook with live formulas, and a concise quality brief documenting limitations and small-sample warnings.
- Modeled 5.59M synthetic Medicare claims and 13.16M claim lines in a DuckDB star schema with dimension and fact marts, documenting grain and source-to-target mappings.
- Reconciled the warehouse through blocking SQL checks from raw files to staging, facts, and marts and validated payment metrics with 42 independent pandas recomputations.
- Migrated the validated SQL mart layer to 35 dbt models and proved row-for-row equivalence against legacy SQL before production switch-over.
- Defined a governed KPI contract recording owner, grain, numerator, denominator, inclusions, exclusions, and limitations for every metric.
- Published a five-dashboard Tableau workbook whose extracts tie to SQL marts and whose payment measures reconcile within half a cent through independent recomputation.
Education
Powered by Drivetube · Create your own profile at drivetube.ai
Everything on Drivetube
Six products, two of them free forever. Start wherever you are.
Your free digital resume
at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept.
Documentation- Free Job BoardFree
Verified jobs, posted in the last 3 days
verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed.
Documentation - Job Hunt ProgramFrom $199.99
We run your job hunt for you
managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you.
Documentation - Resume Writing ServicesFrom $25.99
Written by senior career writers
human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile.
Documentation - Community MembershipFrom $49.99/yr
AI tools, gated filters and a $10,000+ library
from $49.99/year (₹1,999/year in India), or $4.99/month in the iOS app. Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates.
Documentation - Drivetube HireFree tier
For employers — no job postings, no applications
hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do.
Documentation
- Documentation
every product explained in full, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
- Playbooks
tactical job-search guides. Each post is one named tactic with the situation it applies to, the exact steps, a copy-paste message and honest failure modes. Free, no account.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.
Pricing is identical on the web, iOS and Android — no app-store markup. Monthly Community membership is purchased in the iOS app only; the website sells annual plans. The Android app has no in-app purchase, so Android members subscribe on drivetube.ai and then sign in to the app with full access. Membership follows your account, not your device.