Vishnuvardhan Reddy Jakku
Professional Summary
Senior AI/ML Engineer with 6+ years of experience designing and delivering production-grade Generative AI, LLM, and machine learning systems. Strong hands-on expertise in LangChain, Hugging Face Transformers and Pinecone for Retrieval-Augmented Generation, plus model fine-tuning with LoRA/PEFT and LlamaIndex for contextual grounding. Experienced building cloud-native MLOps and LLMOps pipelines using MLflow, TensorBoard, Kubeflow, Terraform and AWS services (Bedrock, Lambda, S3) to automate retraining, A/B and shadow deployments, and drift detection. Worked across healthcare, insurance, telecom and CPG/retail domains to deploy secure, explainable AI (SHAP/LIME) and real-time inference pipelines using PyTorch, TensorFlow and ONNX. Proven at taking prototypes to production, improving model reliability and reducing inference latency while controlling cloud costs and adoption risk.
Technical Skills
Work Experience
- Architected and deployed a Retrieval-Augmented Generation (RAG) system using LangChain and LlamaIndex with Pinecone vector search for internal enterprise Q&A, achieving 92% answer accuracy and reducing hallucinated responses.
- Orchestrated continuous feedback and infra provisioning using Amazon SQS, Amazon Bedrock, Step Functions and Terraform to capture runtime signals and provision reproducible LLM inference pipelines.
- Implemented an abstractive summarization pipeline using Hugging Face Transformers to automate policy-document summarization and reduce manual review workload through human-in-the-loop validation.
- Fine-tuned domain-specific models (BERT, GPT-2) with LoRA/PEFT and managed distributed training runs on Kubeflow to improve intent recognition and reduce fine-tuning compute cost.
- Automated model retraining and observability with MLflow and TensorBoard; instrumented drift detection triggers and automated rollback policies via AWS Lambda.
- Optimized real-time inference by converting LightGBM and PyTorch models to ONNX and enabling GPU-accelerated inference, improving latency 2.5x for critical prediction endpoints.
- Built an LLM-powered contract assistant with LangChain and OpenAI API to extract clauses, obligations and renewal terms, delivering 90%+ extraction accuracy on production documents.
- Implemented a Whisper-based audio transcription and summarization pipeline to process call-center recordings and accelerate compliance review through automated summaries.
- Applied LoRA fine-tuning to adapt base LLMs to insurance-specific language under constrained compute budgets, improving downstream extraction robustness.
- Developed a secure RAG document QA interface using Pinecone for contextual retrieval and Streamlit for a controlled user surface with metadata-based grounding.
- Packaged inference workloads as Docker container images and deployed serverless endpoints using AWS Lambda and S3 with DynamoDB for session persistence.
- Standardized CI/CD and monitoring using GitHub Actions, TeamCity, UDeploy and CloudWatch to improve deployment reliability and reduce cloud spend.
- Conducted causal ML experiments and uplift modeling for CPG/retail clients using causal inference frameworks and revenue-forecast techniques, improving uplift measurement accuracy 15%.
- Implemented NLP-based text analytics (TF-IDF, Word2Vec, LSTM) to derive customer sentiment and product feedback signals for merchandizing decisions.
- Built demand-forecasting and churn-prediction models with TensorFlow and PyTorch on Azure Databricks, improving retraining throughput and model stability for batch workloads.
- Developed deep-learning recommendation models to support personalized marketing campaigns that contributed to observable retention gains.
- Engineered real-time streaming ingestion and ETL pipelines with Apache Kafka and Hive to enable near-real-time analytics and feature generation for models.
- Investigated production logs and dashboards in Kibana and optimized MapReduce and feature-extraction pipelines to shorten failure recovery cycles.
Education
Certifications
Achievements
- Enterprise RAG Q&A accuracy improvement: Achieved 92% answer accuracy on an enterprise RAG Q&A system (LangChain + Pinecone), reducing hallucinated responses.
- Generative summarization reduced review workload: Shortened manual document review time by building a Hugging Face Transformers summarization pipeline with human-in-the-loop validation.
- High-accuracy contract extraction: Delivered 90%+ extraction accuracy for an LLM-powered contract assistant, greatly reducing manual review effort.
- Call-center compliance automation: Trimmed compliance review effort by automating transcription and summarization of call-center audio using Whisper pipelines.
- Model latency and inference optimization: Improved inference latency 2.5x for real-time models by converting to ONNX and enabling GPU-accelerated inference.
- Cloud cost and release-velocity improvements: Reduced cloud infrastructure costs and improved release velocity through CI/CD improvements and compute/storage optimizations.
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.