Ecological validity and the evaluation of speech summarization quality

Par Conseil national de recherches du Canada

DOI	Trouver le DOI : https://doi.org/10.1109/SLT.2012.6424269
Auteur	Rechercher : McCallum, Anthony; Rechercher : Penn, Gerald; Rechercher : Munteanu, Cosmin¹; Rechercher : Zhu, Xiaodan¹
Affiliation	Conseil national de recherches du Canada. Technologies de l'information et des communications
Format	Texte, Article
Conférence	2012 IEEE Workshop on Spoken Language Technology, SLT 2012, December 2-5, 2012, Miami, FL, USA
Sujet	Automatically generated; Ecological validity; Evaluation criteria; Evaluation measures; Evaluation protocol; Goal orientations; Gold standards; Human-centric; NAtural language processing; Natural languages; Reference data; Speech summarization; Summarization systems; Gold; Human computer interaction; Natural language processing systems; Speech analysis; Speech recognition; Quality control
Résumé	There is little evidence of widespread adoption of speech summarization systems. This may be due in part to the fact that the natural language heuristics used to generate summaries are often optimized with respect to a class of evaluation measures that, while computationally and experimentally inexpensive, rely on subjectively selected gold standards against which automatically generated summaries are scored. This evaluation protocol does not take into account the usefulness of a summary in assisting the listener in achieving his or her goal. In this paper we study how current measures and methods for evaluating summarization systems compare to human-centric evaluation criteria. For this, we have designed and conducted an ecologically valid evaluation that determines the value of a summary when embedded in a task, rather than how closely a summary resembles a gold standard. The results of our evaluation demonstrate that in the domain of lecture summarization, the well-known baseline of maximal marginal relevance [1] is statistically significantly worse than human-generated extractive summaries, and even worse than having no summary at all in a simple quiz-taking task. Priming seems to have no statistically significant effect on the usefulness of the human summaries. This is interesting because priming had been proposed as a technique for increasing kappa scores and/or maintaining goal orientation among summary authors. In addition, our results suggest that ROUGE scores, regardless of whether they are derived from numerically-ranked reference data or ecologically valid human-extracted summaries, may not always be reliable as inexpensive proxies for task-embedded evaluations. In fact, under some conditions, relying exclusively on ROUGE may lead to scoring human-generated summaries very favourably even when a task-embedded score calls their usefulness into question relative to using no summaries at all. © 2012 IEEE.
Date de publication	2012
Dans	Spoken Language Technology Workshop (SLT), 2012 IEEE, 6424269 (2012) : 467–472.
Langue	anglais
Publications évaluées par des pairs	Oui
Numéro NPARC	21269996
Exporter la notice	Exporter en format RIS
Signaler une correction	Signaler une correction (s'ouvre dans un nouvel onglet)
Identificateur de l’enregistrement	e02c6836-327e-4030-8d89-96d3ef98cd39
Enregistrement créé	2013-12-13
Enregistrement modifié	2020-04-21

Date de modification :: 2024-04-19