In this paper we compare the soft-error sensitivity of parallel applications on modern Graphics Processing Units (GPUs) obtained through architectural-level fault injections and high-energy particle beam radiation experiments. Fault-injection and beam experiments provide different information and uses different transient-fault sensitivity metrics, which are hard to combine. In this paper we show how correlating beam and fault-injection data can provide a deeper understanding of the behavior of GPUs in the occurrence of transient faults. In particular, we demonstrate that commonly used architecture-level fault models (and fast injection tools) can be used to identify critical kernels and to associate some experimentally observed output errors with their causes. Additionally, we show how register file and instruction-level injections can be used to evaluate ECC efficiency in reducing the radiation-induced error rate.
Analyzing the criticality of transient faults-induced SDCs on GPU applications / Dos Santos, F. F.; Rech, P.. - (2017), pp. 1-7. ((Intervento presentato al convegno 8th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems, ScalA 2017 - Held in conjunction with the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2017 tenutosi a usa nel 2017 [10.1145/3148226.3148228].
Scheda prodotto non validato
I dati visualizzati non sono stati ancora sottoposti a validazione formale da parte dello Staff di IRIS, ma sono stati ugualmente trasmessi al Sito Docente Cineca (Loginmiur).
Titolo: | Analyzing the criticality of transient faults-induced SDCs on GPU applications | |
Autori: | Dos Santos, F. F.; Rech, P. | |
Autori Unitn: | ||
Titolo del volume contenente il saggio: | Proceedings of ScalA 2017: 8th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems - Held in conjunction with SC 2017: The International Conference for High Performance Computing, Networking, Storage and Analysis | |
Luogo di edizione: | usa | |
Casa editrice: | Association for Computing Machinery, Inc | |
Anno di pubblicazione: | 2017 | |
Codice identificativo Scopus: | 2-s2.0-85054836672 | |
ISBN: | 9781450351256 | |
Handle: | http://hdl.handle.net/11572/346657 | |
Citazione: | Analyzing the criticality of transient faults-induced SDCs on GPU applications / Dos Santos, F. F.; Rech, P.. - (2017), pp. 1-7. ((Intervento presentato al convegno 8th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems, ScalA 2017 - Held in conjunction with the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2017 tenutosi a usa nel 2017 [10.1145/3148226.3148228]. | |
Appare nelle tipologie: | 04.1 Saggio in atti di convegno (Paper in Proceedings) |