Lakshmi Narasimha
Data Engineer (AI/ML) • Atlanta, GA • n*********************@gmail.com • +14******629 • linkedin.com/••••• • drivetube.ai/•••••
Professional Summary
Data Engineer (AI/ML) with 4+ years of experience designing and shipping production-grade data pipelines, real-time lakehouse platforms, and ML systems. Experienced across telecom, healthcare, financial services, and manufacturing; skilled with Azure Databricks, Apache Spark, Delta Lake/Iceberg, Snowflake, Collibra, MLflow, RAG/FAISS pipelines, Terraform, Kubernetes, and CI/CD to deliver governed, analytics-ready data and deployed ML services.
Technical Skills
Programming Language: Shell Scripting,SQL,Python
Databases: Snowflake,Amazon Redshift,Oracle,MySQL,Azure Synapse
Cloud Platforms: Amazon Web Services,S3,AWS Glue,Lambda
DevOps & Infrastructure: Terraform,Kubernetes,Docker,GitHub Actions,Jenkins,Azure DevOps
Messaging & Monitoring: Apache Kafka,Vector,Data Lineage,ControlM
Data Engineering & Processing: Azure Databricks,Apache Spark,Delta Lake,Apache Iceberg,Apache Airflow,Prefect,PySpark,Spark SQL
Data Warehousing: Dimensional Modeling,Star Schema
Data Analysis & Visualization: Power BI,Tableau,Pandas,NumPy
Machine Learning & AI: SHAP,LIME
Vector Databases & RAG: RAG,FAISS,FastAPI,GitLab
MLOps: MLflow,Feature Engineering,LLM
Data Integration & ETL: Azure Data Factory,Collibra,Alation,Great Expectations,dbt
Work Experience
T-Mobile
Overland Park, KS
Data Engineer (AI/ML)
July 2025 – Present
Telecom — supported retail and pharmacy data platforms and operations by building lakehouse and ML-powered services for internal analytics and operations.
Tech Stack: Azure Databricks, Apache Iceberg, PySpark, Spark SQL, MLflow, FastAPI, FAISS, Great Expectations, Amazon CloudWatch, Docker, GitHub Actions
- Designed and implemented a production lakehouse on Azure Databricks using Apache Iceberg for ACID transactions, schema evolution, and time-travel querying — replaced an aging batch warehouse and reduced data availability latency from 24 hours to under 45 minutes.
- Built distributed Spark (PySpark, Spark SQL) transformation workflows processing 100M+ records daily across retail and pharmacy datasets; applied partitioning, broadcast joins, and caching to meet performance SLAs under high volume.
- Led end-to-end development and deployment of a binary classification churn model (98% ROC-AUC) including feature engineering, hyperparameter tuning, MLflow experiment tracking and model registry, and FastAPI-based REST serving for live dashboard consumption.
- Implemented a retrieval-augmented generation (RAG) search pipeline: chunked and embedded internal documentation, stored vectors in FAISS, and integrated LLM APIs to deliver instant Q&A for operations teams, replacing 30+ minutes of manual search.
- Integrated model explainability (SHAP, LIME) and automated data quality checks (Great Expectations) to meet compliance and audit requirements across feature and serving layers.
- Instrumented pipeline observability and alerting using Amazon CloudWatch custom metrics and freshness dashboards; standardized deployments with Docker and GitHub Actions CI/CD across dev, staging, and production.
BNY
New York, NY
Data Engineer
Dec 2024 – May 2025
Financial services — built automated ELT and governance solutions to support cross-functional financial analytics and audit reporting.
Tech Stack: Talend, Azure Data Factory, Prefect, Collibra, Alation, Terraform, Kubernetes, GitLab, Jenkins, Amazon Redshift, Azure Monitor
- Engineered enterprise ELT pipelines in Python, SQL, and Talend processing 20M+ records per batch; automated workflows with Azure Data Factory and Prefect to reduce manual batch effort by 50% across 20+ managed workflows.
- Designed a reusable REST API-based source onboarding framework to standardize ingestion patterns for third-party vendors and internal teams, reducing average onboarding time from weeks to days.
- Implemented enterprise data governance using Collibra for business glossary/policy management and Alation for cataloging and lineage, providing audit teams a single governed view across 20+ pipelines.
- Built feature engineering pipelines in Python (pandas, NumPy) and SQL to produce model-ready datasets; ensured reproducibility and versioning for downstream ML training runs.
- Provisioned cloud infrastructure with Terraform using environment-specific configurations; containerized ETL workloads for Kubernetes and integrated pipelines into GitLab and Jenkins CI/CD for zero-touch deployments.
- Implemented Azure Monitor dashboards and alerting across pipelines, cutting average troubleshooting time by ~40% by surfacing anomalies before downstream consumers were impacted.
Birla Soft
India
Data Engineer
Mar 2023 – Dec 2023
IT services for enterprise clients — built ingestion and warehouse layers to support analytics and BI for clients using cloud-native storage and Snowflake.
Tech Stack: Python, Apache Airflow, AWS S3, Parquet, AWS Glue, AWS Lambda, Snowflake, Oracle, MySQL
- Built Python and SQL ETL pipelines extracting from Oracle, MySQL, and flat files into Amazon S3 as Parquet to create a schema-consistent landing zone for downstream Snowflake models.
- Designed Snowflake warehouse (fact/dimension tables, surrogate key strategy, incremental loads) enabling BI teams to migrate from ad hoc queries to structured self-service reporting.
- Orchestrated batch jobs using Apache Airflow DAGs with dependency management, retry logic, and SLA alerting to replace undocumented manual scripts and eliminate on-call failures.
- Integrated AWS Glue for centralized data cataloging and schema discovery; developed serverless preprocessing with AWS Lambda for lightweight event-driven transforms.
- Performed query profiling and optimized Snowflake SQL using clustering keys, materialized views, and result caching to reduce reporting query times on core dashboards.
- Established data landing, transformation, and publishing patterns with documentation and versioning to ensure model-ready datasets for analytics consumers.
Honda Motors
India
Data Engineer
Jan 2022 – Feb 2023
Automotive manufacturing — delivered ETL solutions to extract SAP S/4HANA transactional data and populate analytics data marts for manufacturing analytics.
Tech Stack: Informatica PowerCenter, SAP S, 4HANA, Workflow Manager, ControlM, UNIX shell scripting, MS Visio, SharePoint
- Developed Informatica PowerCenter mappings and sessions to extract SAP S/4HANA transactional data (via ABAP programs) into staging, integration HUB, and data mart layers, delivering audit-documented data for manufacturing analytics.
- Implemented complex Informatica transformations including unconnected lookups, aggregators, and routers to handle merging, deduplication, and conditional routing across 10+ source feeds.
- Built reusable mapplets for Row ID generation, batch control flags, and standard audit columns used across mapping teams to reduce development duplication and accelerate delivery.
- Automated job scheduling using Workflow Manager with Command and Decision tasks, created ControlM dependency diagrams in Visio, and authored UNIX shell scripts for parameterization and environment switching.
- Led requirement analysis with business and technical stakeholders to define ETL specifications, error handling, and validation rules; produced documentation that became the team standard.
- Generated data validation reports during system and UAT testing, performed root-cause analysis on ETL defects, and maintained traceability (daily load logs, documentation) to ensure reliable data mart delivery.
Education
Kennesaw State University
Master of Science, Computer Science • 2025
Coursework: Big Data, NLP, Cloud Computing, Artificial Intelligence, Data Structures & Algorithms
Vel Tech University, India
Bachelor of Technology, Computer Science • 2023
Powered by Drivetube · Create your own profile at drivetube.ai