Disrupted operation of cloud platform sites in Azure (EU) region

Disrupted operation of cloud platform sites in Azure (EU) region

Date

Monday, February 17 2020 11:02 (GMT+2) - Monday, February 17 2020 11:24 (GMT+2)

 

Status

Complete, some action items in progress

Summary

Access to cloud platform sites hosted in the Azure (EU) region is severely impacted. Slow connections or page load failures are being experienced by end-users.

Impact

Cloud Platform sites (BSS and Storefront web applications)

Users of the service may have experienced service disruption, slow connections, http timeouts or web page load failures regarding the interworks.cloud platform web sites. This applies to both BSS and Storefront (v3 and v4) applications

Root Causes

A significant and sudden increase in batch requests occurred in the cloud platform databases that host interworks.cloud platform data within a very brief time frame (less than a minute). This in turn caused a proportional increase of blocked processes within the databases and extremely higher than usual wait time. User requests were not served within acceptable times, which lead to web page timeouts, slow responses and in some cases failure to successfully load the contents of the page.

Trigger

The situation was triggered during a heavy data operation involving collection of usage data from Microsoft Azure.

Resolution

interworks.cloud senior engineering team was able to troubleshoot the situation by resetting existing database connections and restarting the affected web applications. Further monitoring of the situation as well as constant feedback from affected users verified final resolution of the incident.

Detection

The issue was reported by various customers who experienced loss of connectivity, slow performance or malfunction of their cloud platform services. It was also spotted by interworks.cloud platform monitoring systems that alerted the responsible response team.

Action Items

  • Adjustment of monitoring platform metrics and parameters to cover similar situations where sudden high spikes may affect cloud platform database performance         

  • Internal escalation of issue so that further investigation can take place regarding the usage data retrieval process from Microsoft Azure

Timeline


Monday, February 17 2020 11:02 (GMT+2)

Azure usage data collection was automatically initiated by the responsible cloud platform service.

Monday, February 17 2020 11:12 (GM+2) - 11:14 (GM+2) 

System alerts were received via the monitoring system. A number of reports from affected users were also communicated through our Support dept.

Monday, February 17 2020 11:12 (GM+2) - 11:21 (GMT+2)

The engineering team was actively engaged in incident troubleshooting. Root cause was revealed and immediate actions were undertaken in order to remedy the situation. In addition, an internally escalated issue was created in order to pursue further investigation as to the reason of the service behavior.

Monday, February 17 2020 11:24 (GMT+2)

Cloud platform sites operation was restored. Engineering team maintained close monitoring of the previously affected resources.

Monday, February 17 2020 11:54 (GMT+2)

All services verified as up and running within established parameters. Incident was marked as resolved.