Venkata Sai Vikram Karthik Krovi
Data Engineer • Hyderabad, India • v**********@gmail.com • +91*******697 • drivetube.ai/•••••
Professional Summary
Data Engineer with 2 years of experience building ETL/ELT pipelines, automating web ingestion, and delivering clean, production-ready datasets using Python, PySpark, AWS, and PostgreSQL. Skilled in pipeline design, data validation, and API-backed delivery to support analytics and ML workflows.
Technical Skills
Programming Languages: Python,C++
Frameworks and Libraries: Pandas,NumPy,FastAPI
Databases: SQL,PostgreSQL,Databricks SQL
Cloud and DevOps: AWS S3,AWS IAM,Docker
Testing: Selenium,Data cleaning,Data validation,Schema enforcement
Data and Analytics: ETL,ELT design,Airbyte,AWS Glue,PySpark,Power BI,Tableau
Tools and Methodologies: Git,GitHub
Orchestration & Automation: Job scheduling,Data cataloging
Work Experience
NAFA Barter
Hyderabad, India
Python Developer Intern
2025 – 2025
Fintech/financial analytics work processing real-time OHLC market data; built backend ingestion and delivery services for downstream analysis.
Tech Stack: Python, FastAPI, PostgreSQL, Pandas, PySpark, Docker, GitHub
- Designed and implemented a real-time ingestion pipeline for OHLC market feeds using Python and FastAPI to receive, validate, and persist ticks into PostgreSQL.
- Built ETL transformations with Pandas and PySpark to normalize timestamps, handle missing values, and enforce a consistent schema for downstream consumers.
- Implemented record-level data quality checks and anomaly flagging to prevent malformed or outlier ticks from corrupting analytics datasets.
- Exposed processed market data via FastAPI endpoints to enable analytics and reporting workflows without manual extraction steps.
- Documented pipeline architecture, data contracts, and runbooks to improve handoffs and maintainability for future engineers.
- Containerized services with Docker and managed code and CI via GitHub to ensure reproducible deployments and collaboration.
AICTE Virtual Internship
Hyderabad, India
AWS Data Engineering Intern
2024 – 2024
Virtual data engineering internship focused on building ETL/ELT pipelines on AWS to process structured and semi-structured academic/technical datasets.
Tech Stack: AWS S3, AWS Glue, PySpark, Databricks SQL, SQL, Python
- Built ETL pipelines on AWS using S3, AWS Glue and PySpark to ingest, transform, and store structured and semi-structured datasets.
- Developed PySpark transformations to clean, standardize fields, and implement partitioning strategies for improved query performance.
- Applied secure access patterns using IAM and implemented S3 lifecycle practices to protect sensitive data and optimize storage.
- Added data quality checks and validation steps within Glue jobs to prevent propagation of corrupt records downstream.
- Cataloged datasets and metadata with Glue Data Catalog to enable discoverability and reproducible consumption by analysts.
- Used SQL and Databricks SQL for validation queries and exploratory checks to verify pipeline outputs against source expectations.
Prodigy InfoTech
Hyderabad, India
Data Science Intern
2024 – 2024
IT/analytics engagement focused on preparing and analyzing business datasets to surface insights that supported product and business decisions.
Tech Stack: Python, Pandas, PostgreSQL, SQL, Tableau, Power BI
- Performed end-to-end data cleaning and transformation using Python, Pandas, and SQL to prepare multiple business datasets for analysis.
- Authored SQL queries and views in PostgreSQL to validate joins, aggregations, and derived KPI metrics used by stakeholders.
- Conducted exploratory data analysis and built visualizations in Tableau and Power BI to identify trends and anomalies for stakeholders.
- Automated repetitive preprocessing tasks with parameterized Python scripts to reduce manual effort and accelerate analyses.
- Documented data lineage and preprocessing steps to ensure reproducibility and ease of handoff to analytics teams.
- Collaborated with product and business stakeholders to translate requirements into data specifications and KPI definitions.
Deccan AI Experts
Hyderabad, India
Human-Robot Interaction Data Specialist (Freelance)
2025 – Present
Freelance AI/robotics engagement evaluating and annotating human-robot interaction datasets for model training and research.
Tech Stack: Python, PostgreSQL, AWS S3, Selenium, Pandas
- Evaluated and annotated human-robot interaction datasets against defined quality standards to ensure label accuracy and consistency for model training.
- Built validation pipelines in Python to run consistency checks and remove noisy annotations prior to handoff to modeling teams.
- Created annotation guidelines and sample review workflows to increase annotator agreement and reduce rework during labeling rounds.
- Managed dataset versioning and storage workflows, using S3 and PostgreSQL to maintain traceability across annotation iterations.
- Implemented automated sanity checks (format, range, completeness) to catch issues before downstream consumption by engineers.
- Produced QA reports and feedback loops for annotators and data engineers to continuously improve dataset quality and readiness.
Projects
Financial Data Pipeline with Risk Flagging
Tools Used: Python, SQL, PySpark, PostgreSQL, Pandas
- Built an end-to-end pipeline to extract, clean, and transform financial market data using Python and PySpark.
- Implemented rule-based risk-flagging logic to mark suspicious records and maintain an audit trail for analysts.
- Validated processed data using SQL queries against PostgreSQL and documented each processing stage for reproducibility.
Automated Web Data Collection Pipeline (Self-Directed)
Tools Used: Python, Selenium, Airbyte, PostgreSQL
- Developed Selenium-based scrapers to extract data from dynamic websites and integrated extraction into Airbyte ETL workflows.
- Stored cleaned, canonical records in PostgreSQL and implemented retry, deduplication, and normalization logic for robustness.
- Scheduled and monitored pipelines and implemented error-handling to reduce manual intervention during collection.
Business Requirements & KPI Dashboard
Tools Used: SQL, Power BI, Tableau
- Gathered business requirements and translated them into KPIs, data models, and dashboard wireframes for stakeholders.
- Prepared datasets with SQL-based ETL transformations and automated refresh schedules for dashboard consumption.
- Delivered interactive dashboards in Power BI and Tableau to support market intelligence and business decision-making.
Education
Malla Reddy University
B.Tech, Computer Science (Data Science) • Hyderabad, India • 2026
Coursework: Database Management Systems, Data Structures & Algorithms, Operating Systems, Software Engineering
Certifications
Cloud Computing Fundamentals — AWS
Data Analytics Professional Certificate — Google / Coursera
Data Structures & Algorithms — NPTEL
Virtual Job Simulation — Accenture (Forage)
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.