Srujay M
Professional Summary
DevOps and Site Reliability Engineer with 5+ years designing, automating, and operating hybrid cloud platforms across Azure, AWS and GCP. Proven experience building Infrastructure as Code with Terraform and CloudFormation, container platforms with Docker and Kubernetes, and CI/CD automation using GitHub Actions and Jenkins. Strong background in monitoring and observability with Prometheus, Grafana and Splunk, scripting with Python, Bash and PowerShell, and secure infrastructure practices using Vault. Focused on improving reliability, deployment velocity, and operational efficiency through automation and SRE best practices.
Technical Skills
Work Experience
- Designed and implemented Azure infrastructure using Terraform and HashiCorp Vault to standardize secure environment provisioning across enterprise production systems, reducing provisioning time by 40% and improving scalability.
- Built GitHub Actions CI CD pipelines for REST and GraphQL microservices to automate container builds, tests, and deployments, increasing deployment success rates to 95% and shortening release cycles.
- Implemented enterprise observability with Prometheus, Grafana, Splunk, and CloudWatch to centralize metrics and logs, reducing incident detection time by 30% through proactive alerting and dashboards.
- Containerized microservices with Docker and deployed workloads to Kubernetes clusters using autoscaling, rolling updates, and readiness probes to ensure high availability and faster recoveries.
- Automated operational workflows and runbooks using Python, Bash, and PowerShell for patching, configuration validation, and environment remediation, cutting manual operational effort by 35%.
- Led cross functional SRE and development efforts to define SLIs and SLOs and integrated Kafka, ServiceNow, and xMatters for real time incident communication, improving coordination and response workflows during production incidents.
- Designed and automated AWS infrastructure using Terraform, CloudFormation, and Ansible to provision EC2, S3, and RDS resources, reducing environment deployment setup time by 45%.
- Developed Jenkins and GitLab CI pipelines with automated builds, security scans, blue green deployments, and infrastructure validation, raising release success rates to 95%.
- Managed Dockerized applications on Kubernetes clusters with horizontal autoscaling, rolling updates, and health checks to improve production uptime and workload scalability.
- Implemented centralized observability using Prometheus, Grafana, Splunk, ELK, and CloudWatch to increase monitoring coverage and accelerate root cause analysis for production incidents.
- Automated configuration management, patching, and compliance validation using Ansible, Terraform, Chef, and Puppet to eliminate configuration drift and improve operational consistency across hybrid environments.
- Administered Kafka clusters using AWS MSK and integrated producers and consumers into CI CD workflows, improving message reliability and deployment traceability for distributed applications.
- Automated cloud infrastructure provisioning on GCP and hybrid environments with Terraform, developing reusable modules for Compute Engine, BigQuery, and Cloud Storage to improve configuration consistency and reduce setup time.
- Built and optimized CI CD pipelines using CircleCI and Bitbucket Pipelines to automate builds, tests, and deployments, accelerating software delivery cycles by 40%.
- Containerized backend applications with Docker and managed Kubernetes deployments with autoscaling and health checks to increase application reliability and deployment stability.
- Implemented ArgoCD GitOps workflows and authored reusable Helm charts to enforce declarative deployments and eliminate configuration drift across environments.
- Integrated Datadog and New Relic for observability to monitor latency and throughput, improving proactive incident detection and reducing response times by 30%.
- Optimized distributed application performance by adding Redis and Memcached caching layers, reducing database load and improving API response times by 25%.
- Implemented registration, scheduling, real time testing, and automated grading features using Java and Spring Boot, reducing administrative workload by 50% and improving grading accuracy by 30%.
- Designed and developed REST APIs with robust error handling and validation to ensure reliable communication between frontend and backend systems.
- Provisioned and maintained AWS infrastructure including EC2, S3, IAM, and VPC configurations to improve scalability and environment reliability for hosted applications.
- Automated infrastructure operations and deployment tasks using Bash scripting and Terraform, reducing configuration errors and streamlining deployments.
- Operated in Linux environments to execute shell scripts, manage file systems, and troubleshoot system issues to support smooth development and testing cycles.
- Performed system debugging, deployment validation, and issue resolution across cloud hosted applications to improve production stability and user experience.
Projects
- Built scalable backend services using Java and Spring Boot with REST APIs and AWS integrations to support QR generation, message routing, and reliable data processing.
- Designed a QR based communication platform to enable secure, privacy first interactions without exposing personal contact details, improving user safety and accessibility in real world scenarios.
Education
Certifications
Powered by Drivetube · Create your own profile at drivetube.ai
Explore Drivetube
- Drivetube Profile — your free digital resume — at drivetube.ai/in/your-name: one true standard resume with a Hiring Snapshot (visa status, expected salary, notice period, work preference, relocation), an ATS-ready PDF download and a single shareable link. Free forever; interview requests come from verified employers and your contact details stay masked until you accept. Documentation.
- Free Job Board — verified openings crawled ATS-by-ATS from 100,000+ real company career pages across 35 ATS platforms. Shows the true posting date from the source ATS — not when a listing was indexed — and deletes every general listing 3 days after it was actually posted. No ghost jobs, no ad-sponsored listings, no staffing reposts, no account needed. Documentation.
- Job Hunt Program — managed job hunting, a one-time purchase from $199.99. JobScout matches verified roles to your real experience band, Blend AI writes a uniquely tailored resume and cover letter for every application, and the Autofill extension fills the form — or Let Us Apply submits it for you. Documentation.
- Resume Writing Services — human-written, ATS-optimised resumes by senior career writers, from ₹499.99 / $25.99. Available in every country, written to the destination country's own standard — a US resume, UK CV, German Lebenslauf and Indian resume are genuinely different documents. A paid service, separate from the free Drivetube Profile. Documentation.
- Community Membership — from $4.99/month (₹1,999/year in India). Unlocks the gated job-board filters — visa sponsorship, security clearance, workplace and application time — plus Job-Scout AI matching, Resume Report AI, Interview AI prep sheets, Recruiter Outreach AI sent from your own Gmail, Apply or Skip triage, a daily market feed and a $10,000+ library including 23 ATS-validated resume templates. Documentation.
- Drivetube Hire — for employers — hiring with no job postings and no applications. Paste your real job description and AI matches it against candidates' true standard resumes, returning ranked candidates with a match %, matched and missing skills and written reasoning. Free tier included; employers pay, candidates never do. Documentation.
Full product documentation — every product explained, with feature-by-feature comparisons against the job boards, AI apply tools, resume services and hiring platforms people actually use.
The job board covers the United States, India, United Kingdom, Canada, Europe and Australia, and Resume Writing Services are available in every country.