Skip to content

Venkatesh Katragadda

Data Engineer • Dallas, TX • v*******************@applywizard.ai • +14******894 • drivetube.ai/•••••

Professional Summary

Data Engineer with 4+ years of experience building ETL/ELT pipelines, cloud data warehouses, streaming flows, and BI-ready datasets across product, healthcare, finance, and enterprise analytics. Skilled in Python, SQL, PySpark, Databricks, Snowflake, Airflow, and data modeling to deliver reliable reporting layers, data quality checks, and automated pipeline monitoring.

Technical Skills

Programming Languages: Python,Bash
Frameworks and Libraries: PySpark,Apache Spark,Delta Lake,dbt
Databases: SQL,Snowflake,PostgreSQL,MySQL,Oracle,SQL Server
Cloud and DevOps: AWS,Azure,Google Cloud Platform,Jenkins,Datadog
Testing: Dimensional Modeling,SCD Type 1,2,Data Validation
Data and Analytics: ETL,Power BI,Tableau,DAX
Tools and Methodologies: GitHub
Skills: Spark SQL
ETL & Orchestration: Apache Airflow,Azure Data Factory,Informatica,ELT
Warehousing & Storage: AWS Redshift,Azure Synapse Analytics,Google BigQuery
Streaming & Messaging: Kafka

Work Experience

Cintram Inc
Dallas, TX
Data Engineer
Feb 2026 – Present
Built product, billing, and workflow reporting pipelines and models to support operations and analytics for product usage and billing analysis.
Tech Stack: Python, SQL, PostgreSQL, AWS S3, Apache Airflow, AWS CloudWatch, GitHub, GitHub Actions, Power BI
  • Built incremental ETL pipelines using Python, SQL and PostgreSQL with AWS S3 staging to provide operations teams dependable daily product, billing and workflow datasets for reporting and reconciliation.
  • Integrated 12 external feeds (REST API, CSV, SFTP) via Python ingestion and SQL scheduling logic to centralize product usage and billing sources into cleaner reporting layers.
  • Redesigned PostgreSQL reporting models into fact/dimension tables, SQL views and indexes, reducing repeated query effort across product analytics reports by 30%.
  • Implemented SQL-based schema validation, duplicate and null checks and referential integrity rules to prevent invalid records from reaching dashboard datasets.
  • Automated validation and pipeline schedules in Apache Airflow and AWS CloudWatch with structured logging and monitoring, cutting manual pipeline checks by 35%.
  • Controlled production releases with GitHub and GitHub Actions, improved CI/CD traceability and incident troubleshooting for scheduled ETL updates, accelerating support resolution.
Tata Consultancy Services
Data Engineer Associate
Feb 2021 – Dec 2023
Provided data engineering and analytics platform services as part of IT consulting engagements, standardizing ingestion, warehousing and reporting for enterprise clients.
Tech Stack: Python, PySpark, Apache Spark, Informatica, Apache Airflow, Azure Data Factory, AWS Redshift, Snowflake, Azure Synapse Analytics, Kafka, GitHub, Jenkins, Datadog
  • Engineered ingestion pipelines with Python, PySpark, SQL, Informatica and Apache Airflow to standardize CSV, JSON, Parquet and Avro loads, cutting report refresh times by 50%.
  • Tuned heavy warehouse workloads across AWS Redshift, Azure Synapse Analytics and Snowflake using Spark SQL, indexing and query optimization to decrease long-running query runtimes by 40%.
  • Designed dimensional models (star/snowflake), fact and dimension tables and KPI logic in SQL to deliver consistent reporting across 10 business dashboards for stakeholders.
  • Streamed operational event data using Kafka, Apache Spark, PySpark and Delta Lake to create curated near real-time event layers for analytics and monitoring use cases.
  • Orchestrated ETL/ELT workflows in Apache Airflow and Azure Data Factory with retry logic, dependency checks and logging, reducing SLA breaches by 30%.
  • Implemented CI/CD and monitoring workflows with GitHub, Jenkins and Datadog to improve deployment traceability and accelerate incident troubleshooting for support teams.
ITC Infotech
ETL Data Engineer
Sep 2019 – Dec 2020
Delivered ETL and migration projects for enterprise reporting, consolidating file, API and relational sources into cloud warehouses for analytics.
Tech Stack: Informatica, Python, SQL, Apache Airflow, PySpark, Spark SQL, Snowflake, Google BigQuery, AWS S3, AWS Redshift
  • Streamlined Informatica, Python, SQL and Apache Airflow workflows by correcting transformation logic and batch schedules, reducing recurring pipeline failures by 30%.
  • Consolidated SFTP files, REST API extracts and relational source tables through Informatica and SQL to produce a cleaner warehouse layer for sales and customer analytics.
  • Built Snowflake and Google BigQuery reporting datasets with star schema and reusable SQL logic, increasing dashboard readiness by 25%.
  • Validated batch load outputs using PySpark and SQL with null checks, duplicate checks and referential rules to keep inaccurate records out of downstream reporting tables.
  • Migrated on-prem extracts into AWS S3 and AWS Redshift with secure staging and ETL/ELT pipelines, lowering warehouse processing costs by 35%.
  • Optimized Apache Spark batch loads with Spark SQL, partitioning and optimized joins, shortening transaction processing cycles by 20%.

Projects

Healthcare Data Lakehouse Pipeline
Tools Used: Azure Data Factory, Azure Data Lake Storage Gen2, Azure Databricks, PySpark, Delta Lake, Azure Synapse Analytics, dbt, Power BI
  • Designed ingestion in Azure Data Factory from SFTP, REST APIs and CDC feeds into ADLS Gen2 to organize patient encounters, claims, provider and eligibility data for governed healthcare analytics.
  • Transformed raw healthcare files in Azure Databricks with PySpark, Delta Lake and schema validation to create curated layers with PHI handling and duplicate checks.
  • Modeled Synapse and dbt tables for claim denials, readmissions and eligibility trends and connected datasets to Power BI for finance and care team reporting.
Real-Time Financial Transaction Analytics Pipeline
Tools Used: Kafka, PySpark, Apache Spark, AWS S3, Snowflake, dbt, SQL
  • Developed Kafka-based streaming processing with PySpark and Apache Spark to capture transaction events, checkpoint streams and persist clean data to AWS S3.
  • Applied SQL validation, duplicate and null checks and reference-data rules across batch and streaming flows to maintain reliable transaction datasets for risk reporting.
  • Structured Snowflake models using dbt and dimensional modeling to enable Power BI reporting for transaction volume, flagged activity and failed-payment patterns.
Cloud Data Warehouse and BI Data Mart Modernization
Tools Used: Snowflake, PostgreSQL, SQL Server, Python, Apache Airflow, dbt, GitHub
  • Consolidated PostgreSQL, SQL Server and CSV source data into Snowflake using Python, SQL and Airflow to create a centralized warehouse for business reporting.
  • Built reusable data marts with star schema, dimension tables, fact tables and SCD Type 1/2 to provide consistent datasets for sales, billing and operations dashboards.
  • Managed pipeline deployment with GitHub and CI/CD workflows and added logging and monitoring to help teams trace data loads and maintain production reporting cycles.

Education

University of North Texas
Master of Science in Data Science • Denton, TX • Jan 2024 – Dec 2025

Certifications

IBM Data Engineering Professional Certificate — Coursera
DeepLearning.AI Data Engineering Professional Certificate — Coursera
IBM Data Warehouse Engineer Professional Certificate — Coursera
Preparing for Google Cloud Certification: Cloud Data Engineer — Coursera
Data Engineering Foundations Professional Certificate — LinkedIn Learning

Powered by Drivetube · Create your own profile at drivetube.ai

Explore Drivetube

  • Drivetube Profile — your free digital resume at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
  • Free Job Board verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
  • Job Hunt Program managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
  • Resume Writing Services human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
  • Community Membership from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
  • Drivetube Hire — for employers hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.

Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.

The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.