Sriteja Kandimalla
Professional Summary
Cloud Operations & DevOps Resilience Engineer with 9+ years of experience designing, automating, and operating large-scale cloud infrastructure on AWS and Azure — specializing in observability, connectivity, high availability, and platform resilience. Experienced in building CI/CD pipelines, container orchestration (Docker, Kubernetes), IaC (Terraform, Ansible, CloudFormation), multi-account AWS environments, and automated disaster recovery. AWS-certified across Solutions Architect, DevOps Professional, and Security Specialty.
Technical Skills
Work Experience
- Led end-to-end AWS infrastructure automation using Terraform and CloudFormation to provision EC2, S3, RDS, VPC, ELB and Lambda resources, improving provisioning repeatability and reducing manual errors.
- Designed and enforced CI/CD pipelines with Jenkins and AWS CodePipeline, integrating Docker, SonarQube, Nexus and Git to accelerate release cadence and ensure quality gates before production deploys.
- Implemented observability using CloudWatch dashboards, centralized log groups, and alarms to monitor live traffic, resource utilization, and application health across environments for faster incident detection.
- Deployed and managed containerized applications with Docker and Kubernetes, implementing rolling updates and blue/green deployment patterns to achieve zero-downtime application upgrades.
- Automated server provisioning and configuration management with Ansible playbooks and Puppet modules to eliminate configuration drift across dev, staging, and production environments.
- Implemented cloud security controls—IAM roles/policies, GuardDuty, VPC security groups and network ACLs—and built disaster recovery and backup frameworks with automated failover/failback runbooks to meet business continuity SLAs.
- Engineered multi-region active-active failover architecture using Route53 health checks, cross-zone ELB balancing and Auto Scaling policies to meet organizational uptime and regulatory requirements.
- Built and maintained infrastructure CI/CD using Jenkins, Terraform and AWS CodePipeline to automate infrastructure changes with complete audit trails and controlled rollouts across environments.
- Provisioned and managed EKS clusters using Terraform and Helm charts, enabling automated, repeatable deployments and zero-touch application rollouts for the VIRTRU app.
- Implemented observability and alerting pipelines with Splunk, CloudWatch and PagerDuty, creating dashboards and automated incident response workflows to reduce mean time to resolution.
- Provided on-call platform support and automated deployment runbooks, cutting deployment time by 40% and reducing human error during production change windows.
- Led region-level failover exercises and chaos engineering drills to validate disaster recovery playbooks and identify single points of failure in distributed platform components.
- Migrated on-premises BI applications to AWS, designing VPC architectures with public/private subnets, NACLs, NAT Gateways and route tables to meet networking and security needs.
- Automated infrastructure lifecycle using Terraform and CloudFormation stacks to enable repeatable, auditable environment provisioning aligned to corporate governance.
- Built CI/CD pipelines with Jenkins, integrating Docker, Git, and AWS CodePipeline/CodeBuild/CodeDeploy to deploy containerized services to ECS and improve release consistency.
- Implemented AWS CloudWatch monitoring for traffic, logs and resource metrics and integrated SNS for automated alerting to operations teams.
- Automated configuration management across server fleets using Ansible playbooks to reduce configuration drift and manual remediation work.
- Designed and deployed disaster recovery and backup environments with automated failover/failback runbooks to meet RPO/RTO targets for critical applications.
- Designed multi-AZ AWS deployments, Auto Scaling groups and ELB configurations to achieve 99.9% availability SLAs for a distributed financial data platform.
- Defined and architected a new data analytics platform for global Solutions Architects, including scalable Data Lake and Data Warehouse designs to support growing analytics needs.
- Provisioned Airflow infrastructure to automate ETL and data pipeline deployments and onboarded 10+ teams to the platform to standardize pipeline execution and scheduling.
- Implemented a comprehensive observability stack using Splunk, CloudWatch and PagerDuty for log aggregation, anomaly detection and on-call incident workflows.
- Executed multi-region failover procedures including Route53 DNS failover, cross-region S3 replication and RDS read-replica promotion to validate disaster recovery readiness.
- Automated deployment runbooks with Python and Bash, reducing manual deployment time by 40% and enabling self-service release capabilities for multiple development teams.
- Managed enterprise CI/CD and build pipelines using Jenkins, Maven and ANT across customer projects to improve deployment reliability and consistency.
- Provisioned highly available EC2 infrastructure with Terraform and CloudFormation and extended providers via custom Terraform plugins to meet customer-specific needs.
- Deployed and operated Kubernetes clusters for containerized customer workloads, managing pods, services, ConfigMaps and rolling deployments to enable zero-downtime upgrades.
- Participated in disaster recovery exercises to create cross-AZ replicas and configured CloudWatch alarms for comprehensive infrastructure monitoring.
- Configured Splunk add-ons and DB integrations for centralized log aggregation and real-time analytics across customer environments.
- Mentored customer engineering teams on DevOps best practices, IaC patterns and container strategies aligned to the AWS Well-Architected Framework.
- Designed Azure infrastructure including AKS clusters, VNets, Azure AD integration and ExpressRoute connectivity to support secure, high-performance workloads.
- Automated Azure resource provisioning using Terraform modules for resource groups, networking, compute and RBAC policies to ensure consistent deployments.
- Deployed containerized workloads with Docker Compose and Helm, implementing pod autoscaling and zero-downtime deployment strategies for production services.
- Built Splunk architecture for centralized logging and real-time monitoring of cloud control plane services, creating dashboards and anomaly detection alerts.
- Automated infrastructure operations using Ansible Tower and Python SSH-based playbooks for configuration management and continuous deployment orchestration.
- Enforced security and governance through RBAC policies, network controls and IaC best practices to meet internal compliance standards.
- Built and operated RHEL Linux infrastructure fleet, performing system configuration, performance tuning and prioritized incident resolution across the environment.
- Designed and implemented enterprise CI/CD using Jenkins and Puppet to enable continuous integration, automated testing and consistent deployments for .NET and Java teams.
- Developed Puppet modules and manifests to automate deployment and lifecycle management across 500+ nodes, enforcing consistency and reducing manual intervention.
- Managed container adoption with Docker and Kubernetes, supporting teams in containerization strategy, image lifecycle and registry operations.
- Administered DNS, web, mail and DHCP services and managed user directories using LDAP/NIS to support global enterprise operations.
- Implemented HAProxy load balancer configurations and platform automation to improve service availability and scaling for production web services.
- Provisioned AWS infrastructure with CloudFormation to automate EC2, ECS, ELB, S3, RDS, DynamoDB, VPC and IAM resource lifecycles for application workloads.
- Deployed Nagios monitoring with custom checks across 2,000+ Linux servers and integrated monitoring with Jenkins and Puppet for continuous operations.
- Managed Puppet infrastructure upgrades and refactored modules to leverage new capabilities while integrating HAProxy load balancing for high-availability services.
- Configured AWS GuardDuty within a centralized master account and integrated alerting with Slack for real-time security notifications.
- Performed large-scale Linux administration tasks including VMware VM provisioning, Solaris LDOM/ZFS management and network troubleshooting with IPTables/NMAP.
- Implemented authentication and access controls using Centrify and performed system hardening to improve the overall security posture of production systems.
- Designed and built Linux system infrastructure for Health Systems environments including new server builds, VM provisioning and firewall/network configurations.
- Authored Ansible playbooks to automate Tomcat and Apache installation and configuration, reducing manual setup time across environment tiers.
- Built real-time monitoring dashboards using the ELK Stack to ensure high availability of cloud control plane services and detect infrastructure anomalies.
- Deployed EC2 instances using Jenkins-driven CI/CD pipelines with Chef and AWS CloudFormation to standardize environment provisioning.
- Managed end-to-end deployments and code promotions across DEV to PROD and coordinated with firewall teams on DMZ server rules and network segments.
- Provided operational support for applications and collaborated with cross-functional teams to meet compliance and uptime requirements.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.