Skip to content

Lakshmi Narasimha

Data Engineer (AI/ML) • Atlanta, GA • n*********************@gmail.com • +14******629 • linkedin.com/••••• • drivetube.ai/•••••

Professional Summary

Data Engineer (AI/ML) with 4+ years of experience designing and shipping production-grade data pipelines, real-time lakehouse platforms, and ML systems. Experienced across telecom, healthcare, financial services, and manufacturing; skilled with Azure Databricks, Apache Spark, Delta Lake/Iceberg, Snowflake, Collibra, MLflow, RAG/FAISS pipelines, Terraform, Kubernetes, and CI/CD to deliver governed, analytics-ready data and deployed ML services.

Technical Skills

Programming Language: Shell Scripting,SQL,Python
Databases: Snowflake,Amazon Redshift,Oracle,MySQL,Azure Synapse
Cloud Platforms: Amazon Web Services,S3,AWS Glue,Lambda
DevOps & Infrastructure: Terraform,Kubernetes,Docker,GitHub Actions,Jenkins,Azure DevOps
Messaging & Monitoring: Apache Kafka,Vector,Data Lineage,ControlM
Data Engineering & Processing: Azure Databricks,Apache Spark,Delta Lake,Apache Iceberg,Apache Airflow,Prefect,PySpark,Spark SQL
Data Warehousing: Dimensional Modeling,Star Schema
Data Analysis & Visualization: Power BI,Tableau,Pandas,NumPy
Machine Learning & AI: SHAP,LIME
Vector Databases & RAG: RAG,FAISS,FastAPI,GitLab
MLOps: MLflow,Feature Engineering,LLM
Data Integration & ETL: Azure Data Factory,Collibra,Alation,Great Expectations,dbt

Work Experience

T-Mobile
Overland Park, KS
Data Engineer (AI/ML)
July 2025 – Present
Telecom — supported retail and pharmacy data platforms and operations by building lakehouse and ML-powered services for internal analytics and operations.
Tech Stack: Azure Databricks, Apache Iceberg, PySpark, Spark SQL, MLflow, FastAPI, FAISS, Great Expectations, Amazon CloudWatch, Docker, GitHub Actions
  • Designed and implemented a production lakehouse on Azure Databricks using Apache Iceberg for ACID transactions, schema evolution, and time-travel querying — replaced an aging batch warehouse and reduced data availability latency from 24 hours to under 45 minutes.
  • Built distributed Spark (PySpark, Spark SQL) transformation workflows processing 100M+ records daily across retail and pharmacy datasets; applied partitioning, broadcast joins, and caching to meet performance SLAs under high volume.
  • Led end-to-end development and deployment of a binary classification churn model (98% ROC-AUC) including feature engineering, hyperparameter tuning, MLflow experiment tracking and model registry, and FastAPI-based REST serving for live dashboard consumption.
  • Implemented a retrieval-augmented generation (RAG) search pipeline: chunked and embedded internal documentation, stored vectors in FAISS, and integrated LLM APIs to deliver instant Q&A for operations teams, replacing 30+ minutes of manual search.
  • Integrated model explainability (SHAP, LIME) and automated data quality checks (Great Expectations) to meet compliance and audit requirements across feature and serving layers.
  • Instrumented pipeline observability and alerting using Amazon CloudWatch custom metrics and freshness dashboards; standardized deployments with Docker and GitHub Actions CI/CD across dev, staging, and production.
BNY
New York, NY
Data Engineer
Dec 2024 – May 2025
Financial services — built automated ELT and governance solutions to support cross-functional financial analytics and audit reporting.
Tech Stack: Talend, Azure Data Factory, Prefect, Collibra, Alation, Terraform, Kubernetes, GitLab, Jenkins, Amazon Redshift, Azure Monitor
  • Engineered enterprise ELT pipelines in Python, SQL, and Talend processing 20M+ records per batch; automated workflows with Azure Data Factory and Prefect to reduce manual batch effort by 50% across 20+ managed workflows.
  • Designed a reusable REST API-based source onboarding framework to standardize ingestion patterns for third-party vendors and internal teams, reducing average onboarding time from weeks to days.
  • Implemented enterprise data governance using Collibra for business glossary/policy management and Alation for cataloging and lineage, providing audit teams a single governed view across 20+ pipelines.
  • Built feature engineering pipelines in Python (pandas, NumPy) and SQL to produce model-ready datasets; ensured reproducibility and versioning for downstream ML training runs.
  • Provisioned cloud infrastructure with Terraform using environment-specific configurations; containerized ETL workloads for Kubernetes and integrated pipelines into GitLab and Jenkins CI/CD for zero-touch deployments.
  • Implemented Azure Monitor dashboards and alerting across pipelines, cutting average troubleshooting time by ~40% by surfacing anomalies before downstream consumers were impacted.
Birla Soft
India
Data Engineer
Mar 2023 – Dec 2023
IT services for enterprise clients — built ingestion and warehouse layers to support analytics and BI for clients using cloud-native storage and Snowflake.
Tech Stack: Python, Apache Airflow, AWS S3, Parquet, AWS Glue, AWS Lambda, Snowflake, Oracle, MySQL
  • Built Python and SQL ETL pipelines extracting from Oracle, MySQL, and flat files into Amazon S3 as Parquet to create a schema-consistent landing zone for downstream Snowflake models.
  • Designed Snowflake warehouse (fact/dimension tables, surrogate key strategy, incremental loads) enabling BI teams to migrate from ad hoc queries to structured self-service reporting.
  • Orchestrated batch jobs using Apache Airflow DAGs with dependency management, retry logic, and SLA alerting to replace undocumented manual scripts and eliminate on-call failures.
  • Integrated AWS Glue for centralized data cataloging and schema discovery; developed serverless preprocessing with AWS Lambda for lightweight event-driven transforms.
  • Performed query profiling and optimized Snowflake SQL using clustering keys, materialized views, and result caching to reduce reporting query times on core dashboards.
  • Established data landing, transformation, and publishing patterns with documentation and versioning to ensure model-ready datasets for analytics consumers.
Honda Motors
India
Data Engineer
Jan 2022 – Feb 2023
Automotive manufacturing — delivered ETL solutions to extract SAP S/4HANA transactional data and populate analytics data marts for manufacturing analytics.
Tech Stack: Informatica PowerCenter, SAP S, 4HANA, Workflow Manager, ControlM, UNIX shell scripting, MS Visio, SharePoint
  • Developed Informatica PowerCenter mappings and sessions to extract SAP S/4HANA transactional data (via ABAP programs) into staging, integration HUB, and data mart layers, delivering audit-documented data for manufacturing analytics.
  • Implemented complex Informatica transformations including unconnected lookups, aggregators, and routers to handle merging, deduplication, and conditional routing across 10+ source feeds.
  • Built reusable mapplets for Row ID generation, batch control flags, and standard audit columns used across mapping teams to reduce development duplication and accelerate delivery.
  • Automated job scheduling using Workflow Manager with Command and Decision tasks, created ControlM dependency diagrams in Visio, and authored UNIX shell scripts for parameterization and environment switching.
  • Led requirement analysis with business and technical stakeholders to define ETL specifications, error handling, and validation rules; produced documentation that became the team standard.
  • Generated data validation reports during system and UAT testing, performed root-cause analysis on ETL defects, and maintained traceability (daily load logs, documentation) to ensure reliable data mart delivery.

Education

Kennesaw State University
Master of Science, Computer Science • 2025
Coursework: Big Data, NLP, Cloud Computing, Artificial Intelligence, Data Structures & Algorithms
Vel Tech University, India
Bachelor of Technology, Computer Science • 2023

Powered by Drivetube · Create your own profile at drivetube.ai