Skip to main navigation Skip to search Skip to main content

EvalDNN: A Toolbox for Evaluating Deep Neural Network Models

  • Yongqiang Tian
  • , Zhihua Zeng
  • , Ming Wen
  • , Yepang Liu
  • , Tzu Yang Kuo
  • , Shing Chi Cheung

Research output: Chapter in Book/Report/Conference proceedingConference PaperResearchpeer-review

Abstract

Recent studies have shown that the performance of deep learning models should be evaluated using various important metrics such as robustness and neuron coverage, besides the widely-used prediction accuracy metric. However, major deep learning frameworks currently only provide APIs to evaluate a model's accuracy. In order to comprehensively assess a deep learning model, framework users and researchers often need to implement new metrics by themselves, which is a tedious job. What is worse, due to the large number of hyper-parameters and inadequate documentation, evaluation results of some deep learning models are hard to reproduce, especially when the models and metrics are both new.To ease the model evaluation in deep learning systems, we have developed EvalDNN, a user-friendly and extensible toolbox supporting multiple frameworks and metrics with a set of carefully designed APIs. Using EvalDNN, evaluation of a pre-trained model with respect to different metrics can be done with a few lines of code. We have evaluated EvalDNN on 79 models from TensorFlow, Keras, GluonCV, and PyTorch. As a result of our effort made to reproduce the evaluation results of existing work, we release a performance benchmark of popular models, which can be a useful reference to facilitate future research. The tool and benchmark are available at https://github.com/yqtianust/EvalDNN and https://yqtianust.github.io/EvalDNN-benchmark/, respectively. A demo video of EvalDNN is available at: Https://youtu.be/v69bNJN2bJc.

Original languageEnglish
Title of host publicationProceedings - 2020 ACM/IEEE 42nd International Conference on Software Engineering: Companion Proceedings, ICSE-Companion 2020
EditorsGregg Rothermel, Doo-Hwan Bae
Place of PublicationPiscataway NJ USA
PublisherIEEE, Institute of Electrical and Electronics Engineers
Pages45-48
Number of pages4
ISBN (Electronic)9781450371223
DOIs
Publication statusPublished - 2020
Externally publishedYes
EventInternational Conference on Software Engineering 2020 - Online, Seoul, Korea, South
Duration: 27 Jun 202019 Jul 2020
Conference number: 42nd
https://dl.acm.org/doi/proceedings/10.1145/3377811 (Proceedings)
https://conf.researchr.org/home/icse-2020 (Website)
https://dl.acm.org/doi/proceedings/10.1145/3377812 (Proceedings - Companion Proceedings)

Conference

ConferenceInternational Conference on Software Engineering 2020
Abbreviated titleICSE 2020
Country/TerritoryKorea, South
CitySeoul
Period27/06/2019/07/20
Internet address

Keywords

  • Deep Learning Model
  • Evaluation

Cite this