Sai Krishna
Professional Summary
Senior Data Engineer with 6.5+ years of experience designing and operating cloud-native batch and streaming data platforms for retail, healthcare, and financial-services domains. Strong hands-on expertise in GCP (BigQuery, Pub/Sub, Dataflow) plus hybrid integrations with Azure and AWS to deliver production ETL/ELT and streaming pipelines. Delivered scalable PySpark/Databricks and Apache Beam data processing, migrated large Hadoop/Hive and Teradata workloads to BigQuery and Snowflake, and implemented MERGE-based CDC and medallion architectures to produce consumption-ready datasets. Built CI/CD and IaC pipelines with Cloud Build and Terraform and enforced governance using IAM, Data Catalog and Collibra. Integrated Vertex AI and vector embeddings to operationalize ML and GenAI workflows. Proven owner of end-to-end data products that improve reliability, reduce query latency, and enable self-service analytics.
Technical Skills
Work Experience
- Designed event-driven ingestion pipelines using Pub/Sub and Dataflow to stream transactions, IoT, and vendor feeds into the cloud, improving end-to-end pipeline reliability for downstream consumers.
- Built enterprise serving layers in BigQuery and Snowflake using partitioning, clustering, and materialized views to accelerate dashboards and reduce query latency by 40%.
- Migrated legacy Hadoop/Hive workloads to BigQuery by implementing schema evolution, SQL reconciliation, and automated validation checks to preserve data accuracy after cutover.
- Developed metadata-driven ingestion frameworks in Python to standardize schemas, route semi-structured JSON sources, and simplify onboarding for new data providers.
- Implemented incremental processing patterns in BigQuery using MERGE-based CDC and watermarking to handle late-arriving data and reduce reprocessing work.
- Integrated Vertex AI into production pipelines to enable document classification and LLM-powered enrichment while preserving data lineage and governance controls.
- Implemented HL7-to-FHIR migration using Cloud Healthcare API and Dataflow to transform low-latency HL7 messages into FHIR resources for downstream clinical systems.
- Engineered Cloud Composer workflows and parameterized SQL deployments to automate Dev/QA/Prod rollouts and reduce release effort by 30%.
- Optimized BigQuery datasets with partitioning, clustering, and materialized views to lower query cost and improve response time for clinical reporting.
- Built fault-tolerant streaming integrations by connecting Apache Kafka with Pub/Sub to ensure continuous ingestion during the HL7-to-FHIR migration.
- Automated CI/CD and infrastructure provisioning using Cloud Build and Terraform while enforcing IAM and Data Catalog controls for HIPAA compliance.
- Developed Looker semantic models and dashboards on BigQuery to provide clinicians and operations teams with self-service monitoring of patient and hospital KPIs.
- Designed a hybrid cloud data platform spanning AWS S3/Glue/Redshift and Azure Synapse to process large-scale property and risk datasets for analytics.
- Engineered Azure Databricks PySpark and Delta Lake pipelines to transform structured and semi-structured property records into curated analytics datasets.
- Built batch and real-time ingestion pipelines using Apache Kafka and Azure Event Hubs to integrate property records, transactions, and third-party risk feeds.
- Optimized Spark workloads through partitioning, caching, and broadcast join strategies to reduce ETL execution time and improve query performance.
- Authored reusable Databricks notebooks and a parameterized ETL framework to accelerate source onboarding and reduce maintenance overhead.
- Implemented monitoring and alerting with AWS CloudWatch and Azure Monitor to provide proactive pipeline health checks and reduce incident resolution time.
Projects
- Engineered a scalable Python pipeline to evaluate 50,000+ GPT-3.5 responses and 12,000+ factual queries to measure hallucination and response consistency for healthcare and legal use cases.
- Automated fact verification by querying external sources using Azure Bing Search API and extracting evidence with BeautifulSoup to validate model outputs.
- Fine-tuned a BERT sentence-transformer to compute cosine similarity scores for semantic consistency and to flag low-confidence responses.
- Developed analytics and evaluation metrics to surface hallucination patterns and improve LLM trustworthiness for downstream applications.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.