Katarla Bharat
System Operations Analyst / Major Incident Manager • k***************@gmail.com • +91*******258 • linkedin.com/••••• • drivetube.ai/•••••
Professional Summary
Major Incident Manager with 3+ years of experience leading P1/P2 incident response, RCA, and ITIL-based compliance to protect critical production services and meet SLA targets. Proven single-point-of-contact on major incident bridges, author of PIR/RCA, and driver of change governance, MTTR reduction, and operational improvements.
Technical Skills
Data and Analytics: Data analysis
Tools and Methodologies: ServiceNow,BMC Control-M,Change window tracking,Production deployments
Incident Management & ITSM: Major Incident Management MIM,ITIL framework,Problem Management,Change Management,Emergency Change Management,Outage & Crisis Management,Post-Incident Review PIR
Monitoring & Observability: Splunk,BigPanda,Monitoring tools,Service Management SLM
Reporting & Analysis: SLA & KPI reporting,Mean Time to Recovery MTTR analysis,Trend analysis,Root Cause Analysis RCA
Operations & Resilience: Disaster recovery DR,Incident runbooks & playbooks,Incident audits,ITSM knowledge base management
Communication & Collaboration: Stakeholder communication,Executive incident summaries,Cross-functional coordination,Escalation handling
Work Experience
ICE Data Services
System Operations Analyst
April 2025 – March 2026
Supported operations for a financial market data and analytics provider, focusing on platform availability and major incident response for data delivery systems.
Tech Stack: ServiceNow, BigPanda, Splunk, BMC Control-M, Service Management, Monitoring tools
- Served as single point of contact for P1/P2 major incidents, driving rapid triage and coordinated remediation using ServiceNow and BigPanda to ensure incidents were handled within SLA thresholds.
- Orchestrated bridge calls and war-room coordination across engineering, network, and support teams, leveraging Splunk for diagnostics to accelerate root-cause identification and recovery.
- Authored Post-Incident Reviews and Root Cause Analysis reports and delivered executive summaries to leadership to improve transparency, compliance, and corrective action tracking.
- Governed emergency change processes in ServiceNow and coordinated emergency deployments, ensuring changes were reviewed, approved, and communicated to minimize production risk.
- Produced SLA, KPI, and MTTR reports using Service Level Management practices and monitoring data to identify trends and recommend operational improvements.
- Conducted incident audits and updated runbooks and knowledge-base articles to standardize response playbooks, reducing repeat escalations and improving team readiness.
Colruyt Group
IT Operations
May 2022 – October 2024
Worked in retail IT operations supporting availability of point-of-sale, supply chain, and backend systems; owned major incident response and operational resiliency activities.
Tech Stack: ServiceNow, Splunk, BigPanda, BMC Control-M, Service Management, Monitoring tools
- Led Major Incident Management as the primary escalation point, achieving 30% faster response times, reducing downtime by 20%, and maintaining 95% SLA compliance through structured processes.
- Directed end-to-end incident resolution across IT Operations, Network Security, and Infrastructure teams, using ServiceNow and monitoring telemetry to restore services and limit business impact.
- Performed Root Cause Analysis and Post-Incident Reviews that identified corrective actions and reduced recurring incidents by 25%, strengthening overall operational resilience.
- Supported change management, disaster recovery planning, and continuous improvement initiatives; implemented process changes that improved team efficiency by 20%.
- Implemented structured escalation workflows and proactive monitoring that resulted in zero SLA breaches over six consecutive months for critical services.
- Led critical P1 incident response and coordinated cross-team remediation that prevented substantial business loss, and standardized a PIR framework adopted across operational teams.
Education
Siddhartha Institute of Engineering & Technology
Bachelor of Science in Computer Science • 2020
Achievements
- Star Performer — Critical P1 Resolution: Recognized at Colruyt Group for leading a critical P1 resolution that prevented an estimated $200K+ business loss while maintaining SLA thresholds.
- SLA Compliance Streak: Achieved zero SLA breaches over six consecutive months by implementing structured escalation workflows and proactive monitoring strategies.
- Post-Incident Review Framework: Spearheaded a PIR framework that reduced recurring incidents by 25% and was adopted as standard practice across multiple operational teams.
Powered by Drivetube · Create your own profile at drivetube.ai