As fellow pet owners who have built a company focused on better care through understanding, we know how important it is to be able to rely on the products that help care for our pets. On Tuesday, August 11, our Petlibro app experienced an outage that affected a significant portion of our customer base. This issue was fully resolved as of Wednesday evening, August 12. We sincerely apologize for the disruption and understand the frustration and concern it caused. We want to be transparent about what happened, but more importantly, share the steps we’re moving forward with. We are taking immediate extra measures to strengthen our system and help avoid a similar issue from occurring in the future.
Looking ahead
We will continue to hold ourselves to a high standard and are committed to learning from this situation.
-
We have investigated the issue, solved the underlying problem, and increased demand capacity.
-
Our crisis task force is preparing an audit protocol to search our backend and work to resolve any similar inconsistencies.
-
We will continue investigating secondary issues caused by the outage and update our systems with appropriate solutions.
-
We will continue reviewing reports that a device did not carry out a routine as expected, so we fully understand what occurred.
-
We will be making the
service status page a permanent part of our website, so users can always check the status of our app and servers.
-
We will be scaling up our customer support team to expedite the backlog of support tickets resulting from the outage.
What happened
The outage was due to cloud server issues that caused a backlog of data requests, overloading the system memory and bringing down device service in the process. After successfully restarting the cloud service, the backlog reoccurred and the system went down again. At that point, the team conducted systematic troubleshooting of the cloud service system, identified the underlying issue, patched a solution, and then closely monitored the slow upscale of traffic towards stable recovery.
Because of this outage, the app was no longer able to connect to devices and as a result device communication was not possible. Services related to the app, including on-demand actions, operation history, records, notifications, and videos, were unavailable.
A more detailed timeline can be found below. Note, throughout this time our team was working 24/7 to find a resolution.
-
August 11th:
-
5:20AM PT: A critical error occurred affecting app login and control.
-
6:30AM PT: App functionality was restored and the system began to process a large backlog of data requests resulting from the outage.
-
6:53AM PT: A flood of incoming requests overwhelmed connection capacity and the system failed again.
-
8:04AM PT: Basic functionality was restored and service was added incrementally to monitor for any outstanding issues.
-
9:07AM PT: As more functions came online the app's loading time slowed and we made the choice to scale back functionality.
-
8:04PM PT: Throughout the day, service would restore before eventually crashing again. The development team identified a cache issue and disabled service to remedy the issue.
-
11:43PM PT: A solution to the cache issue was implemented and service was reactivated.
-
August 12th:
-
11:05AM PT: Over the course of the night device service was slowly restored until we reached normal connectivity.
-
4:51PM PT: Demand loads returned to pre-failure levels with normal stability.
-
7:39PM PT: After monitoring the app's stable operation through a peak-usage cycle, the team moved the status of the failure to resolved. Additional patches would be released during off-peak hours that would not affect user experience.
Our Response
-
Upon identifying the failure we immediately activated our crisis response team. Our first action was to work to restore service to the app as quickly as possible to minimize inconvenience.
-
Once service was initially restored we updated our communications announcing resolution but had to rescind this when the app failed again. We made every attempt to restore functionality without halting service but had to make the difficult decision to pause the app after we exhausted all other alternatives.
-
After our team successfully identified the underpinning issue and implemented a solution, we restored service piece by piece to monitor stability of the patched solution.
-
Over the course of the night, we continued to observe performance as more devices reconnected before returning to pre-failure levels the next morning. While much of the service was restored by this point, we continued to operate in a crisis setting until we were able to experience a full cycle of demand to guarantee consistent performance.
-
Simultaneously, our team was working to keep customers and other stakeholders aware of the unfolding process. Over the course of the outage, our team sent 5 emails to active customers with connected devices, posted four updates to both Reddit and Instagram, and activated a status update page on our website with a direct link from our homepage. When the app was in service we also activated popups with relevant information in order for you to always have a way to check on what was happening.
As the recovery continues, we are aware of secondary issues that the outage may have triggered. In many cases, device connectivity has been inconsistent, and we are monitoring these instances. In most cases, a manual restart of the device has resolved the problem. Our products are designed to carry out pre-scheduled routines like feeding and self-cleaning on their own, independent of the app, and this functioned for the vast majority of devices. We are actively reviewing isolated reports where a routine may not have performed as expected to determine if any secondary factors were involved.
Again, we deeply apologize for the inconvenience and frustration this caused and we are taking every effort to improve our products based on this event. We will continue to be in contact with our ongoing response to the outage and next steps. Thank you for your patience this week and for trusting us with your pets. We take this responsibility seriously.
This statement is provided to keep you informed of our progress. It does not affect your rights as a valued customer.