A. Akhila
Professional Summary
Data Engineer with 5+ years of experience managing the data lifecycle from ETL to data governance in SQL Server on-premises and Azure cloud environments. Experienced building scalable ingestion and transformation pipelines for clinical, financial, and platform telemetry datasets using Azure Data Factory, Azure Databricks, Azure Synapse, Snowflake and Apache Spark. Proven strengths in T-SQL, PySpark, data quality and HIPAA-compliant data handling, delivering analytic datasets for research and regulatory reporting while improving performance, observability, and cost efficiency.
Technical Skills
Work Experience
- Designed and operated ETL ingestion and governance pipelines across SQL Server on-prem and Azure, ingesting CSV, Parquet, JSON and fixed-width clinical files to support research analytics, processing 5+ TB of clinical data daily using ADF, Databricks and Synapse.
- Authored advanced T-SQL queries, views and stored procedures; monitored and tuned SQL Server execution plans to improve performance for multi-table clinical aggregations and analytic dataset creation.
- Built and maintained Azure Data Factory and Microsoft Fabric workflows to standardize ingestion and transformation patterns, reducing pipeline failures and ensuring timely delivery of research-ready datasets to investigators.
- Applied PySpark on Azure Databricks with partition pruning and execution-plan optimization to accelerate large-scale clinical transformations and improve job throughput across high-volume workloads.
- Served as Honest Broker for sensitive clinical data: implemented RBAC, data masking, audit logging and documented HIPAA-compliant procedures to ensure secure access and audit readiness for research datasets.
- Created Power BI semantic models and reporting datasets layered on Synapse and Snowflake; documented ETL specs, data flow diagrams and dataset definitions to support reproducibility and compliance audits.
- Profiled upstream core banking, credit risk and customer master systems to identify 30+ schema inconsistencies before pipeline build, preventing downstream reporting failures during month-end regulatory cycles.
- Engineered ADF batch pipelines and Azure Event Hubs streaming ingest from 8+ source systems to standardize ingestion patterns and ensure reliable data availability for risk monitoring and compliance reporting.
- Developed PySpark transformation workflows in Azure Databricks to apply credit risk business rules and portfolio aggregation logic across 8+ business units, producing audit-ready datasets for reporting.
- Authored and maintained dbt models across staging, intermediate and mart layers with automated tests and reusable macros to enforce data quality and accelerate reportable dataset delivery.
- Prepared ML-ready feature datasets for credit risk scoring and churn models, defining feature engineering logic and refresh cadences to support production model training and scoring.
- Implemented cost-monitoring dashboards and optimized Databricks cluster auto-scaling and spot strategies, reducing pipeline compute costs and unbudgeted cloud spend by ~20% while meeting SLAs.
- Mapped field-level business rules and schema change patterns across loan servicing, mortgage and credit bureau systems to inform ingestion architecture and prevent regulatory reporting failures.
- Built AWS-native batch and near-real-time ingestion pipelines using S3, EMR and Kinesis, reducing fraud detection signal availability from 15 minutes to under 2 minutes via streaming and Lambda triggers.
- Engineered ETL/ELT pipelines with PySpark on EMR and Amazon Redshift to apply credit risk rules and compliance calculations, delivering audit-ready datasets consumed by risk and reporting teams.
- Integrated data observability checks across 50+ critical financial pipelines to monitor freshness, volume anomalies and schema drift, enabling early detection of upstream issues prior to regulatory submissions.
- Implemented schema validation, source freshness monitoring and row-level data quality checks to improve dataset reliability and reporting accuracy for compliance stakeholders.
- Developed a Python-based data quality scorecard aggregating pipeline health metrics and delivering automated weekly reports to compliance teams, eliminating manual Excel-based audits.
- Developed distributed PySpark pipelines to ingest and transform platform telemetry, system metrics and user behavior datasets to support operational monitoring and product analytics.
- Designed optimized Spark batch workflows with Google Cloud Storage partitioning and aggregation strategies, improving recurring analytics processing efficiency by ~30%.
- Built BigQuery serving layer with partitioning, clustering and materialized views that improved query performance for operational dashboards by ~40%.
- Implemented Cloud Composer-managed Airflow orchestration for DAG scheduling, dependency management and failure recovery, reducing unplanned pipeline downtime by ~25%.
- Experimented with Apache Flink for stateful stream processing, producing a proof-of-concept real-time aggregation pipeline that reduced analytics latency for internal dashboards.
- Collaborated with internal engineering stakeholders to deliver reliable telemetry datasets and documented transformation logic, contributing to cross-team observability and analytics adoption.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.