TEJA GOPAL REDDY VAKA
Professional Summary
AI Engineer with 6+ years of experience designing, deploying, and operationalizing production ML and LLM systems across fintech, healthcare, and consumer mobile products. Proven track record reducing GPU inference costs by 33%, cutting support ticket turnaround by 87%, and shipping LLM-backed features used by 10k+ daily users. Seeking to apply production-grade MLOps, LLM safety, and end-to-end AI product skills at a fast-moving startup.
Technical Skills
Work Experience
- Built and launched a cross-platform React Native app spanning four AI-powered modules (Care, Things, Family, Journal); owned frontend, backend APIs, and cloud infrastructure end-to-end as the sole AI Engineer.
- Architected Ask Rivulet conversational assistant using Gemini Flash Lite, aggregating context across modules with parallel Firestore queries and guardrails that prompt clarifying questions to prevent hallucinations.
- Designed a multi-step AI vision scanner using Groq Llama 3.2 Vision with a dual-pass self-review loop and atomic Firestore writes to ensure consistent multi-step scan sessions and reduce missed/occluded items.
- Researched household cognitive labor (Daminger framework) and iterated UX to introduce proactive nudges and shared cognitive-load visibility based on private beta user feedback.
- Instrumented Firebase Analytics and event tracking to establish baseline engagement, reached 23 monthly active users in private beta, and created a retention-first product roadmap.
- Implemented production backend patterns (Firestore rules, Cloud Functions, CI/CD pipelines and monitoring) to ensure app reliability and faster iteration during early launch.
- Built and deployed a LoRA-optimized BERT classifier and summarizer for customer intent routing, improving F1 from 0.76 to 0.89 and increasing CSAT by 15% through automated routing.
- Reduced GPU inference costs by 33% via ONNX export and INT8 quantization, enabling the system to scale to 5x query volume within budget constraints.
- Rebuilt model deployment pipelines in Airflow with automated rollback hooks, reducing model release cycle from 3 days to 4 hours and enabling safer, faster releases.
- Implemented statistical drift detection using KS test and PSI with CloudWatch-triggered rollbacks and a monthly retraining loop to maintain SLA-aligned model performance.
- Led a 5TB migration from S3 to Snowflake to centralize transactional data, which decreased support ticket turnaround from 4 hours to 30 minutes (87% reduction).
- Delivered an LLM-backed feature serving 10,000+ daily active users; managed AI backend APIs (TensorFlow/BERT) and mentored a junior scientist through deployment and monitoring.
- Designed ingestion pipelines using AWS Glue to onboard diverse client feeds and built REST API endpoints for external data producers to standardize data collection.
- Added retry and circuit-breaker logic across ingestion pipelines to improve availability and reduce downstream pipeline failures.
- Implemented PySpark and pandas transformation logic to normalize and validate feeds, reducing data processing errors by 15% and enabling real-time analytics across 20+ client systems.
- Built auto-refreshing KPI dashboards in Amazon QuickSight, cutting manual reporting cycles for Fortune 500 clients and improving stakeholder visibility.
- Maintained SLA-bound ETL workflows using AWS Step Functions and Kinesis, processing millions of records daily with consistent throughput.
- Instrumented logging and operational metrics for ingestion jobs to improve incident response and support SLA-driven uptime targets.
- Curated and validated a dataset of 2,000+ question-answer pairs with manual accuracy checks to create a high-quality ground truth for model training.
- Fine-tuned a BERT-based transformer via transfer learning on the curated dataset, improving contextual understanding and response relevance for the platform's Q&A system.
- Built ETL pipelines using Google Cloud Dataflow and orchestrated workflows with Cloud Composer to ingest healthcare records from 100+ medical facilities for near-real-time analytics.
- Processed and transformed 10M+ patient records into a BigQuery warehouse, streamlining reporting and reducing query latency for downstream analytics.
- Optimized BigQuery schemas, partitioning, and clustering to decrease query latency and improve report generation times for clinical dashboards.
- Implemented data quality checks (KS test, z-score) to detect anomalies and reduce source data errors by 15%.
- Developed anonymization and tokenization workflows applying k-anonymity to protect patient privacy and maintain compliance with data-protection requirements.
- Automated ingestion monitoring and alerting to ensure SLA compliance and rapid incident response for data delivery to analytics teams.
Projects
- Built a greenfield RAG pipeline with 3-tier hybrid retrieval (FAISS → cosine → keyword fallback) using GPT-4o-mini to ground retirement advice in IRS and SSA documents with traceable citations across Member and Advisor modes.
- Implemented an intent router with prompt-injection detection and financial compliance guardrails that enforce safe fallbacks and disclaimer-injected outputs to prevent misaligned LLM responses.
- Optimized inference cost to ~$0.0005 per query using structured JSON outputs validated by Pydantic, JSONL interaction logging, and a production migration path across FastAPI, AWS Lambda, pgvector, S3, and CloudWatch.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.