- Title
- Improving service availability of cloud systems by predicting disk error
- Creator
- Xu, Yong; Sui, Kaixin; Chintalapati, Murali; Zhang, Dongmei; Yao, Randolph; Zhang, Hongyu; Lin, Qingwei; Dang, Yingnong; Li, Peng; Jiang, Keceng; Zhang, Wenchi; Lou, Jian-Guang
- Relation
- 2018 USENIX Annual Technical Conference (USENIX ATC '18). 2018 USENIX Annual Technical Conference (Boston, MA 11-13 July, 2018) p. 481-494
- Relation
- https://www.usenix.org/conference/atc18/presentation/xu-yong
- Publisher
- USENIX Association
- Resource Type
- conference paper
- Date
- 2018
- Description
- High service availability is crucial for cloud systems. A typical cloud system uses a large number of physical hard disk drives. Disk errors are one of the most important reasons that lead to service unavailability. Disk error (such as sector error and latency error) can be seen as a form of gray failure, which are fairly subtle failures that are hard to be detected, even when applications are afflicted by them. In this paper, we propose to predict disk errors proactively before they cause more severe damage to the cloud system. The ability to predict faulty disks enables the live migration of existing virtual machines and allocation of new virtual machines to the healthy disks, therefore improving service availability. To build an accurate online prediction model, we utilize both disk-level sensor (SMART) data as well as systemlevel signals. We develop a cost-sensitive ranking-based machine learning model that can learn the characteristics of faulty disks in the past and rank the disks based on their error-proneness in the near future. We evaluate our approach using real-world data collected from a production cloud system. The results confirm that the proposed approach is effective and outperforms related methods. Furthermore, we have successfully applied the proposed approach to improve service availability of Microsoft Azure.
- Subject
- service availability; cloud systems; disk error; hard disk drives
- Identifier
- http://hdl.handle.net/1959.13/1402671
- Identifier
- uon:35051
- Identifier
- ISBN:9781931971447
- Rights
- © 2018 The Authors.
- Language
- eng
- Full Text
- Reviewed
- Hits: 8906
- Visitors: 9111
- Downloads: 245
Thumbnail | File | Description | Size | Format | |||
---|---|---|---|---|---|---|---|
View Details Download | ATTACHMENT02 | Publisher version (open access) | 2 MB | Adobe Acrobat PDF | View Details Download |