Skip to content

Sai Krishna

Senior Data Engineer • n**************@gmail.com • 848****540 • drivetube.ai/•••••

Professional Summary

Senior Data Engineer with 6.5+ years of experience designing and operating cloud-native batch and streaming data platforms for retail, healthcare, and financial-services domains. Strong hands-on expertise in GCP (BigQuery, Pub/Sub, Dataflow) plus hybrid integrations with Azure and AWS to deliver production ETL/ELT and streaming pipelines. Delivered scalable PySpark/Databricks and Apache Beam data processing, migrated large Hadoop/Hive and Teradata workloads to BigQuery and Snowflake, and implemented MERGE-based CDC and medallion architectures to produce consumption-ready datasets. Built CI/CD and IaC pipelines with Cloud Build and Terraform and enforced governance using IAM, Data Catalog and Collibra. Integrated Vertex AI and vector embeddings to operationalize ML and GenAI workflows. Proven owner of end-to-end data products that improve reliability, reduce query latency, and enable self-service analytics.

Technical Skills

Programming Language: Python,SQL
Databases: Snowflake
Cloud Platforms: Google Cloud Platform,Amazon Web Services,Microsoft Azure,Cloud Build,Cloud Composer,IAM
DevOps & Infrastructure: Terraform,Continuous Integration,Continuous Deployment
Messaging & Monitoring: Apache Kafka
Data Engineering & Processing: PySpark
Data Warehousing: Star Schema,Dimensional Modeling
Data Analysis & Visualization: Looker
Generative AI & LLMs: BERT,Vector Embeddings
MLOps: Google Vertex AI
Data Integration & ETL: Azure Data Factory,Change Data Capture,Collibra
Compliance & Governance: Data Catalog

Work Experience

Walmart
Data Engineer
June 2024 – Present
Worked on retail data platform engineering for transaction, IoT and vendor data to support analytics and ML use cases on GCP and Snowflake.
Tech Stack: BigQuery, Snowflake, Sub, Dataflow, Vertex AI
  • Designed event-driven ingestion pipelines using Pub/Sub and Dataflow to stream transactions, IoT, and vendor feeds into the cloud, improving end-to-end pipeline reliability for downstream consumers.
  • Built enterprise serving layers in BigQuery and Snowflake using partitioning, clustering, and materialized views to accelerate dashboards and reduce query latency by 40%.
  • Migrated legacy Hadoop/Hive workloads to BigQuery by implementing schema evolution, SQL reconciliation, and automated validation checks to preserve data accuracy after cutover.
  • Developed metadata-driven ingestion frameworks in Python to standardize schemas, route semi-structured JSON sources, and simplify onboarding for new data providers.
  • Implemented incremental processing patterns in BigQuery using MERGE-based CDC and watermarking to handle late-arriving data and reduce reprocessing work.
  • Integrated Vertex AI into production pipelines to enable document classification and LLM-powered enrichment while preserving data lineage and governance controls.
Cigna Group
Data Engineer
April 2021 – June 2022
Supported healthcare data modernization and real-time HL7-to-FHIR transformation on GCP to enable clinical reporting and operational analytics.
Tech Stack: BigQuery, Sub, Dataflow, Cloud Composer, Cloud Healthcare API, Apache Kafka, Cloud Build, Terraform, Looker, IAM, Data Catalog, CI, CD
  • Implemented HL7-to-FHIR migration using Cloud Healthcare API and Dataflow to transform low-latency HL7 messages into FHIR resources for downstream clinical systems.
  • Engineered Cloud Composer workflows and parameterized SQL deployments to automate Dev/QA/Prod rollouts and reduce release effort by 30%.
  • Optimized BigQuery datasets with partitioning, clustering, and materialized views to lower query cost and improve response time for clinical reporting.
  • Built fault-tolerant streaming integrations by connecting Apache Kafka with Pub/Sub to ensure continuous ingestion during the HL7-to-FHIR migration.
  • Automated CI/CD and infrastructure provisioning using Cloud Build and Terraform while enforcing IAM and Data Catalog controls for HIPAA compliance.
  • Developed Looker semantic models and dashboards on BigQuery to provide clinicians and operations teams with self-service monitoring of patient and hospital KPIs.
Core Logic
Data Engineer
February 2018 – March 2021
Built hybrid cloud data platforms for property and risk analytics across AWS and Azure, delivering batch and streaming ETL for analytics and reporting.
Tech Stack: Azure Databricks, PySpark, Delta Lake, Apache Kafka, Azure Event Hubs, AWS S3, AWS CloudWatch, Azure Monitor
  • Designed a hybrid cloud data platform spanning AWS S3/Glue/Redshift and Azure Synapse to process large-scale property and risk datasets for analytics.
  • Engineered Azure Databricks PySpark and Delta Lake pipelines to transform structured and semi-structured property records into curated analytics datasets.
  • Built batch and real-time ingestion pipelines using Apache Kafka and Azure Event Hubs to integrate property records, transactions, and third-party risk feeds.
  • Optimized Spark workloads through partitioning, caching, and broadcast join strategies to reduce ETL execution time and improve query performance.
  • Authored reusable Databricks notebooks and a parameterized ETL framework to accelerate source onboarding and reduce maintenance overhead.
  • Implemented monitoring and alerting with AWS CloudWatch and Azure Monitor to provide proactive pipeline health checks and reduce incident resolution time.

Projects

Enterprise LLM Reliability & Hallucination Detection Framework
Tools Used: Python, Azure Bing Search API, BeautifulSoup, BERT, Vector Embeddings
  • Engineered a scalable Python pipeline to evaluate 50,000+ GPT-3.5 responses and 12,000+ factual queries to measure hallucination and response consistency for healthcare and legal use cases.
  • Automated fact verification by querying external sources using Azure Bing Search API and extracting evidence with BeautifulSoup to validate model outputs.
  • Fine-tuned a BERT sentence-transformer to compute cosine similarity scores for semantic consistency and to flag low-confidence responses.
  • Developed analytics and evaluation metrics to surface hallucination patterns and improve LLM trustworthiness for downstream applications.

Education

George Mason University
Master of Science in Data Analytics • August 2022 – May 2024

Certifications

Google Cloud Certified – Professional Data Engineer — Google Cloud
Databricks Certified Data Engineer Professional — Databricks
SnowPro Advanced – Data Engineer — Snowflake
Microsoft Certified: Azure Data Engineer Associate — Microsoft
AWS Certified Solutions Architect – Associate — Amazon Web Services
AWS Certified Cloud Practitioner — Amazon Web Services

Powered by Drivetube · Create your own profile at drivetube.ai

Explore Drivetube

  • Drivetube Profile — your free digital resume at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
  • Free Job Board verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
  • Job Hunt Program managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
  • Resume Writing Services human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
  • Community Membership from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
  • Drivetube Hire — for employers hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.

Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.

The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.